Influence of index contribution rate and machine learning models on the susceptibility of Benggang under slope unit

Published: 6 May 2025| Version 1 | DOI: 10.17632/8gznt7wx8k.1
Contributor:
Fei Guo

Description

The dataset contains the research data supporting the manuscript titled "Influence of index contribution rate and machine learning models on the susceptibility of Benggang under slope unit". The data includes the following key components: Evaluation Indicators Data: This comprises 22 evaluation indicators used to assess the susceptibility of Benggang. These indicators include rainfall erosivity, formation lithology, forest height, vegetation coverage, leaf area index, elevation, slope, modified soil adjusted vegetation index, normalized backscatter coefficients of the VV and VH channels, sand content, soil erosion modulus, clay content, erodibility, slope length and slope factor, topographic humidity index, brightness index, coloring index, slope aspect, plane curvature, profile curvature, and hydrodynamic index. The data is sourced from various remote sensing satellites (e.g., Sentinel-1, Sentinel-2), geological surveys, and other environmental datasets. Slope Unit Division Data: The study area (Huichang County, Jiangxi Province) was divided into slope units using a multi-scale segmentation algorithm. This dataset includes the spatial distribution and characteristics of the slope units, which serve as the basic spatial units for the susceptibility assessment. Machine Learning Model Results: The dataset contains the prediction results of three machine learning models (logistic regression, XGBoost, and random forest) applied to assess the susceptibility of Benggang. The results include the AUC values, accuracy statistics, and susceptibility zoning maps under different index contribution rates (100%, 95%, 90%, 85%, 80%, 75%, and 70%). Key Findings: The data supports the key findings of the study, which indicate that the highest prediction accuracy for Benggang susceptibility is achieved when the index contribution rate is 90%. The random forest model demonstrates superior performance with the highest AUC value of 0.849. The dataset also highlights the spatial distribution of high - susceptibility areas in Huichang County, providing valuable insights for disaster prevention and spatial planning in the region.

Files

Steps to reproduce

To reproduce the research data and results presented in the manuscript "Influence of index contribution rate and machine learning models on the susceptibility of Benggang under slope unit", follow these steps: Data Collection: Obtain the 22 evaluation indicators data from the respective sources mentioned in the manuscript, including remote sensing satellites (Sentinel-1, Sentinel-2), geological surveys, and other environmental datasets. Collect the geographical and geological information of the study area (Huichang County, Jiangxi Province), such as elevation, slope, and lithology. Data Preprocessing: Preprocess the remote sensing data using tools like Google Earth Engine to extract relevant indicators such as vegetation coverage and leaf area index. Process the geological and topographical data to generate datasets suitable for analysis. Slope Unit Division: Apply the multi - scale segmentation algorithm to the study area's DEM data to extract slope aspect and mountain shadow as the basic input image. Use the trial - and - error method, combined with the morphological and scale characteristics of historical Benggang in the study area, to determine the parameter values including scale, shape feature weight, and compactness weight. Merge the processed results and optimize them through segmentation, merging, and smoothing to divide the study area into slope units. GeoDetector Method: Use the GeoDetector method to analyze the 22 basic evaluation indicators and calculate the q - value of each indicator to determine their explanatory power. Classify the indicators into different levels based on the cumulative contribution rates to examine how varying combinations influence Benggang susceptibility. Machine Learning Models: Prepare the dataset for the three machine learning models (logistic regression, XGBoost, and random forest) by combining the selected evaluation indicators and the divided slope units. Train and validate the models using the dataset, and adjust the model parameters as needed to optimize their performance. Prediction and Analysis: Use the trained models to predict the susceptibility of Benggang and generate susceptibility zoning maps. Evaluate the prediction accuracy of the models through precision statistical analysis and ROC curve, and compare the results under different index contribution rates and models. Result Interpretation: Analyze the results to identify the optimal evaluation index contribution rate combination and the best prediction model. Interpret the spatial distribution of high - susceptibility areas and their implications for disaster prevention and spatial planning in the region. By following these steps and using the provided data and methods, researchers can reproduce the results of this study and further explore the influence of index contribution rate and machine learning models on the susceptibility of Benggang under slope unit.

Institutions

  • China Three Gorges University

Categories

Geomorphology, Spatial Analysis, Soil Conservation

Licence