Source–path–endpoint geohazard-chain runout prediction data for southeastern Tibet
Description
This dataset supports the manuscript “A geomorphology-informed terrain-matching framework for regional geohazard-chain runout endpoint prediction”. It contains processed tabular data used for event-level model development, path-conditioned endpoint prediction, uncertainty analysis, residual diagnostics, and regional application in southeastern Tibet. The release includes 591 interpreted geohazard-chain events, 267,045 path-step records under the path-extended Protocol2 representation, predictions from five Stage-1 candidate models, Stage-2 residual-analysis outputs, attributes for 8,369 regional candidate sources, 8,355 successfully generated regional prediction records, and Monte Carlo dropout and random-seed uncertainty outputs. Raw optical satellite imagery is not redistributed. The repository contains derived interpretation products, model inputs, and model outputs used in the associated study.
Files
Steps to reproduce
The repository contains the processed tabular data used to reproduce the statistical summaries, model comparisons, uncertainty diagnostics, residual analyses, and regional prediction results reported in the associated manuscript. Import the CSV files using software capable of reading tabular data, such as Python, R, MATLAB, or spreadsheet software. Preserve the original field names and treat blank or null cells as missing values rather than zeros. Use Dataset_S1_stage1_event_index.csv to identify the 591 interpreted events and the event-level split consisting of 415 training events, 93 validation events, and 83 held-out test events. Records belonging to the same event should remain within a single split. Use Dataset_S2_protocol1_path_level_predictions_all_models.csv and Dataset_S3_protocol2_path_level_predictions.csv to reproduce the Protocol1–Protocol2 comparison and the held-out Stage-1 model metrics. Aggregate path-level predictions by event before calculating runout MAE, RMSE, median absolute error, endpoint MAE, the percentage of predictions within 10% runout error, and the percentage of endpoints within 250 m. Use Dataset_S4_stage2_path_level_predictions.csv and Dataset_S5_stage1_stage2_final_master_table.csv to compare the retained Stage-1 predictions with the Stage-2 residual-diagnostic outputs. Event-level absolute errors and endpoint errors can be used to identify the subset of test events improved by Stage-2. Use Dataset_S8_mc_dropout_predictions.csv and Dataset_S9_seed_ensemble_predictions.csv to reproduce the Monte Carlo dropout and random-seed uncertainty summaries. Repeated predictions should be grouped by event before calculating prediction means and standard deviations. Use Dataset_S10_runout_distribution_long_table.csv and Dataset_S11_stage2_residual_training_table.csv to reproduce runout-distribution, residual, subgroup, and explanatory-variable analyses. Use Dataset_S6_regional_source_attribute_table.csv and Dataset_S7_regional_hybrid_prediction_master.csv to reproduce the regional candidate-source counts, source-class summaries, predicted runout distributions, uncertainty statistics, and Stage-2-supported diagnostic summaries. Spatial coordinates use WGS 84 / UTM zone 46N (EPSG:32646), unless otherwise specified. The repository provides processed inputs and outputs for reproducing the reported analyses. Raw optical satellite imagery is not redistributed. Complete model retraining additionally requires the implementation and configuration described in the associated manuscript.
Institutions
- Zhejiang UniversityZhejiang, Hangzhou