This paper discusses methods for the harmonization and combination of large-scale patent and trademarkdatasets with each other and other sources of data. Dictionary- and rule-based approaches to the consolidationof applicant names in patent data are presented and shown to have both benefits and drawbacks inisolation. We combine the two methods and develop a set of rules and dictionaries to consolidate European,Patent Cooperation Treaty (PCT) and US patent data with firm accounting data. The resulting dataencompass about 131,000 patent applicant names from 46 countries, covering 58.8 percent of EPOapplications and 50.6 percent of PCT applications by business organizations during the time periodfrom 1979 to 2008. For US data, the resulting dataset includes around 54,000 assignee names and51.3 percent of US granted patents during approximately the same time period.

Harmonizing and combining large datasets – An application to firm-level patent and accounting data

THOMA, Grid;
2010-01-01

Abstract

This paper discusses methods for the harmonization and combination of large-scale patent and trademarkdatasets with each other and other sources of data. Dictionary- and rule-based approaches to the consolidationof applicant names in patent data are presented and shown to have both benefits and drawbacks inisolation. We combine the two methods and develop a set of rules and dictionaries to consolidate European,Patent Cooperation Treaty (PCT) and US patent data with firm accounting data. The resulting dataencompass about 131,000 patent applicant names from 46 countries, covering 58.8 percent of EPOapplications and 50.6 percent of PCT applications by business organizations during the time periodfrom 1979 to 2008. For US data, the resulting dataset includes around 54,000 assignee names and51.3 percent of US granted patents during approximately the same time period.
2010
9780010732481
268
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11581/226461
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact