Institutions-first or Technology first Development Models: China vs the USA
Description
The dataset used in this study consists of 279,231 USPTO-granted patents, covering firms from 26 countries across 347 cities and classified into 137 NACE industrial codes. Each patent is treated as the unit of analysis and is mapped into 13 major technological sectors, spanning high-, medium-, and low-technology domains based on OECD and Eurostat technology-intensity classifications. The data are drawn exclusively from granted patents at the United States Patent and Trademark Office (USPTO) to ensure a consistent institutional examination standard for cross-country comparison, with both Chinese and U.S. patents evaluated within the same legal framework. The dataset further includes patent-level attributes such as B1/B2 patent type, firm age, and city-level geographic concentration measures, along with national population controls. This structure enables a sectoral-level analysis of relative technological strength and supports comparative assessment of innovation trajectories between China and the United States within a unified patent system.
Files
Steps to reproduce
Patent Data Collection Retrieve granted patent records from the United States Patent and Trademark Office (USPTO) database for the selected study period. Ensure inclusion of both B1 and B2 granted patents. Country Attribution Identify applicant country based on assignee information. Classify patents into three groups: China, United States, and Rest of World (reference category). Data Cleaning and Standardization Remove duplicate patent records Standardize assignee names and country labels Ensure consistent formatting of patent classification fields Retain only granted patents Technological Classification Map each patent to NACE industrial codes (137 categories) and aggregate them into 13 technological sectors, further grouped into high-, medium-, and low-technology categories using OECD/Eurostat classification rules. Variable Construction Dependent variables: binary indicators for each of the 13 technological sectors Key independent variables: China dummy, USA dummy Control variables: patent type (B1/B2), firm age, city-level patent concentration, city-level firm concentration, and national population Dataset Structuring Construct a patent-level panel where each observation represents a granted patent with linked country, sector, firm, and geographic attributes. Descriptive Analysis Generate sectoral distributions and correlation matrices across: technological sectors (13 categories) countries (top patenting countries) Econometric Estimation Estimate logistic regression models separately for each technological sector: Outcome: sector membership (binary) Estimation: logit models Output: odds ratios for China and USA relative to Rest of World Robustness Checks Re-estimate models with: full control variables alternative sector aggregations pooled multi-sector models Technological Intensity Index (Optional Extension) Construct a 100-point high-technology index to evaluate relative positioning across technology intensity gradients. Visualization and Reporting Produce: sectoral comparison figures (high/medium/low tech) correlation matrices regression tables (odds ratios) robustness tables