Dataset on Bidirectional Relationships Between Government Budgetary Expenditure and Marriage and Fertility Rates in China (2010–2024)
Description
This dataset supports the study titled “Bidirectional Relationships Between Government Budgetary Expenditure and Marriage and Fertility Rates in China.” It contains panel data used to examine the dynamic and bidirectional associations between fiscal expenditure and demographic outcomes in China. The dataset covers the period from 2010 to 2024 and integrates data from multiple authoritative sources, including the Ministry of Finance of China, the National Bureau of Statistics of China, policy documents from the CPC Central Committee and the State Council, and the World Bank. The dataset includes the following categories of variables: Dependent variables: crude marriage rate and total fertility rate Core explanatory variables: government budgetary expenditures, classified into direct and indirect categories (e.g., family planning affairs, healthcare, social security, education, housing, and science and technology) Instrumental variables: public budget revenue, central government expenditure, historical marriage rate, and policy dummy variables (including single-child policy, partial two-child policy, full two-child policy, and three-child policy) Control variables: aging rate, GDP per capita, female labor force participation rate, and crude divorce rate To ensure robustness and comparability, missing values were removed, and stationarity tests were conducted. First-order differencing was applied to selected variables where necessary. Variable definitions, transformations, and naming conventions are provided in the accompanying documentation. This dataset was used to perform Granger causality tests and instrumental variable regressions to explore both forward (fiscal expenditure → marriage and fertility rates) and reverse (marriage and fertility rates → fiscal expenditure) relationships. The dataset can be reused for research on demographic economics, public finance, and policy evaluation, particularly in the context of low fertility and population aging. All data have been anonymized and are provided for research and academic purposes only.
Files
Steps to reproduce
Data collection Collect raw data from the following sources: Ministry of Finance of China (budgetary expenditure and revenue data) National Bureau of Statistics of China (marriage rate, fertility rate, population, GDP, divorce rate, and central government expenditure) Policy documents from the CPC Central Committee and the State Council (policy timing variables) World Bank (female labor force participation rate) Data integration and cleaning Merge all variables into a panel dataset covering 2010–2024 (with extended years for selected variables where applicable) Remove observations with missing values using listwise deletion Ensure consistent variable naming and units Variable construction Calculate crude marriage rate and aging rate using provided formulas Construct budget execution indicator: budget execution = (actual expenditure / budgeted amount) × 100 Generate fertility policy dummy variables (single-child, partial two-child, universal two-child, three-child) based on policy years Variable classification Classify budgetary expenditure variables into direct and indirect categories based on expert scoring Direct: family planning, healthcare, social security Indirect: education, science and technology, agriculture, housing Preprocessing for time-series analysis Conduct unit root tests (ADF test) for all variables Apply first-order differencing to non-stationary variables where necessary Retain original variables if differencing is not applicable Descriptive and exploratory analysis Plot long-term trends between budgetary expenditure variables and marriage/fertility rates Apply logarithmic transformation to expenditure variables Calculate Pearson correlation coefficients Granger causality analysis Perform pairwise Granger causality tests Use a lag length of 3 Test both directions: expenditure → MFR MFR → expenditure Instrumental variable estimation Conduct two-stage least squares (2SLS) regression Use heteroskedasticity-robust standard errors (HC1) Include control variables in all models Select models based on significant Granger relationships Software and tools Perform all analyses in R Main packages used: ggplot2 (visualization) vars (Granger causality tests) AER (instrumental variable estimation)
Institutions
- Peking UniversityBeijing, Beijing