Supporting data for explainable machine learning and deep learning modeling of soil erosion susceptibility in the Mandakini River basin

Published: 10 August 2026| Version 1 | DOI: 10.17632/pynm8527hk.1
Contributors:
,
,
,

Description

This dataset supports a study that evaluates soil erosion susceptibility and subwatershed management priority in the Mandakini River basin, India, using an integrated RUSLE, machine learning, deep learning, and SHAP explainability framework. The underlying hypothesis is that soil erosion susceptibility varies systematically with geological, topographic, hydrological, spectral, and land-cover conditions, and that these relationships can be learned and interpreted using data-driven models. The dataset contains the processed modeling inventory derived from RUSLE-based soil-loss classes, ten erosion conditioning factors, model evaluation results, SHAP outputs, and subwatershed priority results for CatBoost, Extra-Trees, MLP-ANN, and TabNet. The modeling inventory was generated from spatial datasets processed at 30 m resolution and sampled across 23 subwatersheds in the basin. Supporting outputs include model performance metrics, predictor influence results, and subwatershed-level rankings. The data show that high and very high erosion susceptibility is concentrated mainly in the upper and north-central parts of the basin. Across the four models, approximately one-third of the basin was classified within the high to very high susceptibility classes. SHAP analysis identified the Bare Soil Index, Modified Normalized Difference Water Index, and Sediment Transport Index as the dominant predictors of model output. The subwatershed priority results consistently identified Markanda Ganga (SW13), Mandakini River (SW11), Bantoli Gad (SW14), Vasuki Ganga (SW10), Sina Gad (SW09), and Madhani River (SW12) as the highest-priority areas for erosion management. The dataset can be used to reproduce the reported analyses, compare model behavior, examine predictor influence, assess erosion susceptibility patterns, and support future studies on explainable remote-sensing-based erosion assessment and subwatershed prioritization in mountainous environments.

Files

Steps to reproduce

The data were generated through an integrated remote sensing, GIS, RUSLE, machine learning/deep learning, and explainable AI workflow for the Mandakini River basin, India. Multi-source geospatial datasets were processed mainly in Google Earth Engine and Google Colab at a common 30 m spatial resolution. Rainfall information was obtained from CHIRPS, elevation and terrain information from SRTM and MERIT Hydro, soil information from publicly available soil datasets, and spectral and land-cover information from Sentinel-2 imagery. These datasets were used to derive the RUSLE rainfall erosivity (R), soil erodibility (K), slope length and steepness (LS), cover-management (C), and support-practice (P) factors and to estimate spatial soil-loss rates. The resulting RUSLE soil-loss values were grouped into erosion classes and used to construct the modeling inventory. Ten erosion conditioning factors including lithology, lineament density, drainage density, terrain ruggedness index, topographic wetness index, sediment transport index, bare soil index, normalized burn ratio, modified normalized difference water index, and land use/land cover were extracted for sampled locations across 23 subwatersheds. The final inventory contained 20,700 samples, with 6,900 observations for each of the three modeling target classes. The data were divided into training, validation, and test subsets using a 70:15:15 stratified split. CatBoost, Extra-Trees, multilayer perceptron artificial neural network (MLP-ANN), and TabNet models were developed and evaluated using accuracy, precision, recall, F1-score, ROC-AUC, Cohen’s kappa, confusion matrices, and stratified five-fold cross-validation. SHAP analysis was subsequently used to quantify and interpret the contribution of the erosion conditioning factors to model predictions. Finally, model-specific susceptibility outputs were summarized at the subwatershed level and combined with RUSLE-derived soil-loss information to calculate a Subwatershed Priority Index and consensus management ranking. The deposited files contain the processed modeling data and supporting outputs required to understand, reproduce, and extend the analytical workflow.

Categories

Geomorphic Hazard

Licence