Risk-informed Conservation of Mountain Railway Heritage Using GIS and Explainable Machine Learning: Landslide Susceptibility and Exposure along the Late Qing Sichuan-Hankou Railway, China
Description
This dataset provides an integrated spatial and analytical record for the corridor-scale landslide-susceptibility and heritage-exposure assessment of 90 Late Qing Sichuan-Hankou Railway heritage sites in Huanghua, Shuiyuesi, Wuduhe, and Xiakou townships, Yichang, China. The dataset contains 145 landslide-positive samples, 145 non-landslide-negative samples, 28 candidate conditioning factors and their screening results, a 35-feature modelling matrix, repeated five-fold cross-validation outputs for five machine-learning models, Random Forest probability and five-class susceptibility rasters, SHAP values and factor-importance tables, and exposure classifications for all 90 heritage sites. Random Forest achieved an out-of-fold AUC of 0.8101 and a PR-AUC of 0.7779. The exposure assessment identifies 76 sites, accounting for 84.44% of the heritage inventory, within High or Very High susceptibility zones. The archive provides analysis-ready CSV tables, 30 m GeoTIFF rasters in EPSG:4526, GeoPackage spatial databases, a bilingual site-identifier and coordinate crosswalk, and a self-contained offline WebGIS. It also includes variable definitions, model configurations, figure- and table-level source mappings, a complete file manifest, and Python scripts for data export, Random Forest validation and fitting, and summary visualization. Derived factor values and full provenance information are included, while third-party source rasters remain with the providers identified in the variable dictionary. The dataset supports reproducibility, spatial verification, conservation-priority assessment, digital heritage management, and comparative GIS-based research on mountainous linear transport heritage. Data, documentation, derived figures, and the WebGIS are distributed under the Creative Commons Attribution 4.0 International License, while the Python code is distributed under the MIT License.
Files
Steps to reproduce
1. Extract the archive while preserving the directory structure and read README.md before using the data. 2. Use the files in 01_samples and 02_factor_screening to examine the balanced inventory of 145 landslide and 145 non-landslide samples and the screening of 28 candidate conditioning factors. 3. Use the 35-feature matrices and model outputs in 03_model_cv to evaluate Logistic Regression, Support Vector Machine, Random Forest, Gradient Boosting, and XGBoost under repeated stratified five-fold cross-validation with ten repeats. 4. Consult 07_metadata/model_configuration.csv for the model settings and reported performance values. 5. Use the probability raster, five-class susceptibility raster, Jenks breaks, and area statistics in 04_susceptibility_raster to reproduce the susceptibility maps and class summaries. 6. Use the SHAP values and factor-importance tables in 05_SHAP to reproduce the model-interpretation results. 7. Overlay the heritage-site data in 06_heritage_risk with the five-class susceptibility raster to reproduce the exposure classification of the 90 heritage sites. 8. Variable definitions, units, coordinate systems, and source information are provided in 07_metadata/factor_variable_dictionary.csv.
Institutions
- Huazhong University of Science and TechnologyHubei, Wuhan