Dataset for A Machine Learning Framework to Examine Score Sensitivity in Potentially Polluting Shipwreck Risk Assessment
Description
This dataset contains the data and calculation flow for Operationalizing the Precautionary Approach in Marine Pollution Control: A Machine Learning Framework to Examine Score Sensitivity in Potentially Polluting Shipwreck Risk Assessment
Files
Steps to reproduce
- S1 (README): Data description and metadata (this sheet). - S2: Overview of the nine baseline evaluation criteria and their maximum point allocations within the current risk assessment framework. - S3: Data collection sources, parameters, and baseline inventory profile for the evaluated national cohort of sunken vessels (n= 1,196, as of Q4 2024). - S4: Baseline score distributions for the managed (GM, IM) and EX groups. - S5: Frequency distribution of total cumulative risk scores across the entire evaluated national sunken vessel inventory (n = 1,196) under the existing framework. - S6: Descriptive baseline statistics of the evaluated fleet. - S7: Correlation matrix heatmap illustrating pairwise Pearson correlation coefficients (r) across individual evaluation criteria and final cumulative scores for the entire vessel population (n= 1,196). - S8: PCA diagnostics for the original entire-fleet training subset (n = 956). - S9: Score-tertile reconstruction metrics on the entire-fleet test subset (n = 240). - S10: Confusion matrices on the entire-fleet test subset (n = 240). - S11: Test-subset ROC coordinates and mean class-specific AUCs (n = 240). - S12: Fishing-cohort correlations (n = 926) and training-subset PCA (n = 740). - S13: Ranking of evaluation criteria based on supervised machine-learning feature importance: an isolated case study of the fishing vessel cohort (n = 926). - S14: Fishing-cohort test-subset reconstruction metrics (n = 186). - S15: Fishing-cohort test-subset confusion matrices (n = 186). - S16: Single-criterion sensitivity; all 1,196 records rescored against 247 initially managed vessels. - S17: Paired-criterion sensitivity for C4/C5, C4/C6, and C5/C6. No three-axis analysis. - S18: Nine minimum-adjustment scenarios S1 to S9, exact weights, rounded display allocations, and transitions. - S19 Inputs: Complete vessel-level analysis inputs, including all nine criterion scores. - S20 Code: Analysis source and pinned analysis requirements. Execute the supplied Python files to regenerate results. - S21 Audit: Diagnostic outputs, scenario selection, and calculation definitions. - S22 Partitions: Original training/test membership and labels, with unstandardized criterion scores.
Institutions
- Pukyong National UniversityBusan, Busan