Data for: Mismatch between Urban Eco-sanitation Improvement and Health Service Supply Capacity: Evidence from Chinese City Panel Data and Interpretable Machine Learning

Published: 22 June 2026| Version 1 | DOI: 10.17632/hw4j6xv99v.1
Contributor:
Zhen Peng

Description

This dataset is linked to the manuscript "Mismatch between Urban Eco-sanitation Improvement and Health Service Supply Capacity: Evidence from Chinese City Panel Data and Interpretable Machine Learning" (BMC Health Services Research). It contains a panel of 289 Chinese prefecture-level and above cities covering 2002–2024, with 29 worksheets including: (1) the core panel with the Health Service Supply Capacity index, Eco-sanitation Environment index, Green Infrastructure Index, Municipal Sanitation Index, Pollution Pressure Index, and socioeconomic controls; (2) city-level energy consumption and CEADs sectoral CO₂ emission data from 2002–2019; (3) entropy-weight index weights and a full variable dictionary; (4) descriptive statistics, missing data summaries, and annual aggregates; (5) complete regression output tables (baseline two-way fixed effects, dimension decomposition, mechanism tests, moderation, heterogeneity, and robustness checks); and (6) random forest and gradient boosting performance metrics, feature importance, and full SHAP value matrices. Key data sources are the China City Statistical Yearbook, China Urban Construction Statistical Yearbook, China Energy Statistical Yearbook, and the CEADs CO₂ emission inventory.

Files

Steps to reproduce

The dataset package contains the processed city-level panel data, variable definitions, index weights, model specifications, statistical results, and figure-related data used in the associated study. To reproduce the analyses: (1) Download and open the Excel workbook. Review the sheets “README_submission”, “variable_dictionary”, “source_manifest”, “Results_index”, and “Model_specs” before conducting the analysis. (2) Use the sheet “main_2002_2024” for the baseline analysis. Retain observations with main_model_sample = 1. The sheets “energy_2006_2024”, “main_ceads_2002_2019”, and “energy_ceads_2006_2019” are used for the energy- and carbon-emission-extended analyses. (3) The health service supply capacity index (HSC), eco-sanitation environment index (ESE), green infrastructure index (GII), municipal sanitation index (MSI), and pollution pressure index (PPI) are provided in the analytical data sheets. To reconstruct these indices, winsorize continuous component variables at the 1st and 99th percentiles, apply min-max normalization using the pooled city-year sample, and calculate entropy weights according to the values reported in the “index_weights” sheet. Lagged variables should be generated separately within each city. (4) Estimate the baseline two-way fixed effects model with HSC as the dependent variable and the one-year lagged ESE as the core explanatory variable. Include city and year fixed effects and cluster standard errors at the city level. Add the control variables progressively following the specifications reported in “Model_specs” and Table 6. (5) For dimension-specific analysis, jointly include the one-year lagged GII, MSI, and PPI. Conduct mechanism, moderation, heterogeneity, placebo, alternative-index, alternative-sample, and robustness analyses according to the corresponding specifications in “Model_specs”. The reported estimates can be checked against Tables 7–11. (6) For the machine-learning analysis, use observations from 2018 and earlier as the training sample and observations from 2019 onward as the test sample. Fit random forest and gradient boosting models using the hyperparameters reported in the manuscript. Evaluate predictive performance using R², RMSE, and MAE, and calculate SHAP values to assess feature importance and nonlinear dependence patterns. (7) Compare the reproduced results with the sheets “Table5_Descriptive” to “Table13_ML_importance” and “SHAP_importance_full”. The “yearly_summary” and relevant model-result sheets can be used to reproduce the figures. The workbook contains processed analytical data rather than all original source files. Researchers wishing to reconstruct the dataset from the original public sources should consult the “source_manifest” and “variable_dictionary” sheets for data-source information, variable definitions, and harmonization procedures.

Categories

Public Health, Urban Studies, Sustainable Development, Health Services Research

Licence