A dataset of firm-level geopolitical risk perception for Chinese listed companies

Published: 10 February 2026| Version 1 | DOI: 10.17632/7pp7jy5zmf.1
Contributors:
,
, Chong Qian

Description

This dataset provides firm-level measures of corporate geopolitical risk perception for Chinese A-share listed companies (2010–2024), constructed using NLP techniques applied to Management Discussion and Analysis (MD&A) sections of annual reports. Methodology: The index employs TF-IDF weighting based on a hierarchical "Four Pillars, Twelve Dimensions" lexicon containing 449 core terms, expanded via Word2Vec semantic modeling (vector size=300, similarity threshold=0.6). Text preprocessing includes jieba segmentation and stop-word filtering. Scores represent normalized keyword frequency per 100 words, winsorized at 1% and 99% levels. Theoretical Framework: Economic & Trade (Econ): Trade Barriers (tariffs, anti-dumping), Green Barriers (CBAM, carbon footprints), Investment & Financial (CFIUS, SWIFT sanctions) Technology & Innovation (Tech): Sanction Blacklists (Entity List, EAR), Bio/Hard Tech Decoupling (CHIPS Act, biosecurity), Defensive Substitution (IT innovation, domestic substitution) Supply Chain (Chain): Ethical Compliance (UFLPA, forced labor), Strategic Reconstruction (China Plus One, de-risking), Resource & Logistics (rare earths, energy security) Macro-Political & Data (Macro_Data): Data Sovereignty (GDPR, cybersecurity), Ideological & Political (systemic rivalry), Regional Stability (geopolitical conflicts) Data Content: Primary file: GRPI_China_Firm_Level_2010_2024.csv (54,524 firm-year observations) Key variables: Stock code (Stkcd), Year, Total_Words, GRPI_Total (aggregate index), four pillar scores (Econ, Tech, Chain, Macro_Data), twelve sub-dimension scores, industry classification (6 sectors), ownership type (Private/SOE)

Files

Steps to reproduce

1. Environment Preparation Install Python 3.7+ and libraries: jieba, gensim, pandas, numpy, and scipy. Secure access to the CSMAR Database, and gather a specialized Chinese financial stop word dictionary plus metadata (stock codes, industry, ownership types) from CSMAR. 2. Data Acquisition and Filtering Download standardized .txt MD&A sections of Chinese A-share listed companies (2010–2024) from CSMAR. Exclude observations with cleaned MD&A text shorter than 100 words for score reliability. 3. Text Preprocessing Remove non-Chinese characters, formatting artifacts, and irrelevant content via regular expressions. Segment cleaned texts into terms with jieba. Filter non-informative high-frequency words using the specialized stop word dictionary. 4. Lexicon Construction and Expansion Manually curate 449 core seed words covering the "Four Pillars, Twelve Dimensions" framework. Train a Word2Vec model (vector size=300) on the preprocessed corpus with gensim to expand semantically similar terms. Manually audit expanded terms to refine the final hierarchical risk lexicon. 5. Index Calculation Count the frequency of lexicon terms in each firm-year MD&A text. Compute normalized scores using a TF-IDF-based approach, adjusting for text length and cross-firm term distribution. Sum sub-dimension scores to get GRPI_Total and four pillar indices (Econ, Tech, Chain, Macro_Data). 6. Quality Control and Validation Apply winsorization at the 1% and 99% levels to mitigate outliers. Calculate skewness and kurtosis, and conduct a mean difference test (2010–2017 vs. 2018–2024) to validate sensitivity to geopolitical shocks. Merge CSMAR metadata for cross-sectional analysis. 7. Data Integration and Output Compile the final .csv dataset with columns: Stkcd, Year, Total_Words, GRPI_Total, pillar scores, and sub-dimension scores. Verify sample size (target: 54,524 observations) and summary statistics. Upload to Mendeley Data with accession number 10.17632/7pp7jy5zmf.1 8. Reproducibility Validation Replicate key findings: 2018 structural break (10.3% post-2018 increase) and industry/ownership heterogeneities. Retrain the Word2Vec model to confirm lexicon and index consistency.

Institutions

Categories

Economics

Funders

  • the 2025 National Social Science Fund of China
    Grant ID: 25BGJ078

Licence