Aggregate Shocks and Foundational Schooling: Evidence from India

Published: 23 July 2026| Version 1 | DOI: 10.17632/fkwpzdkxs3.1
Contributors:
,

Description

Description This repository contains the complete replication package for the manuscript "Aggregate Shocks and Foundational Schooling: Evidence from India." The study examines how pandemic–related schooling disruption affected school non-attendance among children aged 6–14 years in India by exploiting the staggered timing of National Family Health Survey (NFHS-5) fieldwork relative to the nationwide school-closure announcement. The central hypothesis is that pandemic-era schooling disruptions altered children's educational participation, with heterogeneous effects across age groups and gender. The empirical analysis combines nationally representative household survey data from the National Family Health Survey (NFHS-4, 2015–16; NFHS-5, 2019–21), the Time Use Survey (TUS 2019 and 2024), and the Periodic Labour Force Survey (PLFS 2019–20 and 2023–24), together with harmonized district crosswalk files. The analysis uses variation in interview timing to estimate the effects of pandemic-era schooling disruption on school non-attendance and examines age-specific and gender-specific heterogeneity, reasons for school non-attendance, and complementary evidence from children's time allocation and labour-force participation. Robustness analyses include event-study models, falsification tests, propensity score methods, Oster sensitivity analysis, difference-in-differences estimators, alternative treatment definitions, alternative clustering schemes, and wild-cluster bootstrap inference. The repository includes all Stata source code, documentation, district harmonization files, intermediate datasets, final analytical datasets, and the original TUS and PLFS datasets required to reproduce the complete analysis. The NFHS raw datasets are not included because they are distributed under the Demographic and Health Surveys (DHS) Program data-use agreement. Researchers should obtain the NFHS-4 and NFHS-5 Household Member (PR) files directly from the DHS Program and place them in the prescribed directory structure before executing the replication package. The complete workflow is executed through a single master replication script (00_Master_Replication.do), which automatically prepares all analytical datasets, harmonizes district identifiers across surveys, estimates every empirical specification, reproduces all tables and figures reported in the manuscript and appendix, and generates replication log files. The replication package was developed and tested using Stata 17. The accompanying README provides detailed instructions on software requirements, folder organization, raw data sources, and replication procedures to facilitate full reproducibility and reuse.

Files

Steps to reproduce

1. Download the NFHS-4 (2015–16) and NFHS-5 (2019–21) Household Member (PR) datasets from the Demographic and Health Surveys (DHS) Program after obtaining the necessary access permissions. 2. Place the downloaded NFHS datasets in the appropriate directories within the replication package: o 02_Raw_Data/NFHS 4/Household Member Recode/ o 02_Raw_Data/NFHS 5/Household Member/ 3. Ensure that the original filenames of all raw datasets are preserved. The replication package already includes the required Time Use Survey (TUS) datasets, Periodic Labour Force Survey (PLFS) datasets, and district harmonization crosswalk files. 4. Install Stata 17 (or a later compatible version). The required user-written packages (ftools, reghdfe, boottest, psmatch2, estout, and coefplot) are installed automatically by 01_Setup.do if they are not already available. 5. Set Stata's working directory to the replication package folder. 6. Open Stata and execute: 7. do "00_Master_Replication.do" 8. The master replication script automatically initializes the project environment, prepares all analytical datasets, harmonizes district identifiers across NFHS, TUS, and PLFS, estimates all empirical models, reproduces every table and figure reported in the manuscript and appendix, and generates replication log files. 9. All replicated outputs, including intermediate datasets, final analytical datasets, tables, figures, and log files, are saved automatically in their respective output directories described in the accompanying README.

Institutions

Categories

Economics, Education, Development Studies

Licence