Harmonization Crosswalks and Code for India's ASUSE/UNAE Non-Farm Unincorporated Enterprise Survey Series (2010-11 to 2025)
Description
This dataset provides the derived harmonization infrastructure used to construct a unified establishment-level analysis file from six rounds of India's national non-farm unincorporated enterprise surveys: the Unincorporated Non-Agricultural Enterprises Survey (UNAE, NSS rounds 67 and 73, 2010-11 and 2015-16) and the Annual Survey of Unincorporated Sector Enterprises (ASUSE, 2021-22, 2022-23, 2023-24, and 2025), both conducted by India's Ministry of Statistics and Programme Implementation (MoSPI). The deposit contains four components. First, a concordance mapping the three-digit National Sample Survey (NSS) region code to state and union territory, covering all 88 officially published regions and three empirically identified corrections not present in the published concordance, validated to zero unmatched establishments across all six survey rounds. Second, a classification of all 72 two-digit National Industrial Classification (NIC 2008) divisions observed in the ASUSE 2021-22 sample into contact-intensity tiers (High, Medium, Low), together with a five-digit sub-classification of the 57 retail-trade product codes within Division 47. Third, a validated mapping of the numeric worker-category codes used in each survey round's Employment Particulars block to standardized categories (working owner, formal hired worker, informal hired worker, other worker, and self-help-group member), including identification of subtotal rows requiring exclusion, a source of measurement error not previously documented for this survey series. Fourth, an R script implementing the full harmonization pipeline, including establishment-key construction for each round's distinct naming convention, state derivation, sector classification, employment-quality and registration-status variable construction, and merging with the Oxford COVID-19 Government Response Tracker. This deposit does not include the underlying survey microdata, which is collected and licensed by MoSPI and must be obtained directly from the National Statistical Office. The crosswalks and code provided here are intended to let researchers with independent access to this microdata reconstruct a harmonized, analysis-ready dataset without re-deriving the state concordance, sector classification, or worker-category mapping from scratch, each of which required substantial validation against official published aggregates during the construction of the accompanying research paper.
Files
Steps to reproduce
Steps to reproduce 1. Obtain the relevant round(s) of raw microdata directly from MoSPI/NSO (https://www.mospi.gov.in). Not included in this deposit. 2. Place each round's downloaded microdata in separate folders, preserving the original block structure (Block 1, 2, 8 for UNAE; Level 01-16 for ASUSE). 3. Open asuse_unae_harmonization_pipeline.R in R (4.4+) and update the folder paths in Section 0 to your microdata and the four crosswalk CSVs in this deposit. 4. Install required packages: dplyr, tidyr, fixest. 5. Run the script. It builds a harmonized establishment ID per round, derives state from the NSS region code, classifies activity into contact-intensity tiers, constructs employment-quality and registration variables, excluding documented subtotal rows. 6. For the identified tier (ASUSE 2021-22, in-coverage 2022-23), the script merges Oxford COVID-19 Government Response Tracker stringency data; download command and path are included. 7. The script prints a validation summary (establishment counts, unmatched-state counts, registration rates). Compare against Tables 1, 8, and 12 of the accompanying paper to confirm reconstruction. 8. Use the resulting data frames (analysis_2010, analysis_2015, analysis_2122, analysis_2223, analysis_2324, analysis_2025) with run_cluster_bootstrap_fast() and run_by_tier() to reproduce Sections 4-5.
Institutions
- Gokhale Institute of Politics and EconomicsMaharashtra, Pune