Dataset for A Machine Learning Framework to Resolve Scale-Dependency in Potentially Polluting Shipwreck Risk Assessment

Published: 18 August 2026| Version 1 | DOI: 10.17632/3b8mgddfgd.1
Contributors:
,

Description

This dataset contains the data and calculation flow for Operationalizing the Precautionary Approach in Marine Pollution Control: A Machine Learning Framework to Resolve Scale-Dependency in Potentially Polluting Shipwreck Risk Assessment.

Files

Steps to reproduce

- S1 (README): Data description and metadata (this sheet) - S2: Overview of the nine baseline evaluation criteria and their maximum point allocations within the current risk assessment framework. - S3: Data collection sources, parameters, and baseline inventory profile for the evaluated national cohort of sunken vessels (n= 1,196, as of Q4 2024). - S4: Density and frequency distribution of baseline risk scores for general management (GM) and Exempt (EX) vessels cohorts evaluated across three representative criteria - S5: Frequency distribution of total cumulative risk scores across the entire evaluated national sunken vessel inventory (n = 1,196) under the existing framework - S6: Descriptive baseline statistics of the evaluated fleet - S7: Correlation matrix heatmap illustrating pairwise Pearson correlation coefficients (r) across individual evaluation criteria and final cumulative scores for the entire vessel population (n= 1,196). - S8: Principal Component Analysis (PCA) diagnostics for the entire national fleet (n = 1,196) - S9: Complete predictive performance metrics of the Decision Tree (DT), Random Forest (RF), and Logistic Regression (LR) models evaluated across the entire national vessel fleet (n\ = 1,196). - S10: Comparison of predictive confusion matrices across the three supervised machine-learning models for the entire national fleet (n = 1,196): (a) Decision Tree (DT), (b) Random Forest (RF), and (c) Logistic Regression (LR). - S11: Comparison of Receiver Operating Characteristic (ROC) curves across the three evaluated machine-learning models for the entire national fleet (n = 1,196), illustrating true positive rates against false positive rates with corresponding Area Under the Curve (AUC) values. - S12: Targeted sub-cohort diagnostics for the isolated fishing vessel population (n = 926) - S13: Ranking of evaluation criteria based on supervised machine-learning feature importance: an isolated case study of the fishing vessel cohort (n = 926) - S14: Complete predictive performance metrics of the Decision Tree (DT), Random Forest (RF), and Logistic Regression (LR) models evaluated across the isolated fishing vessel cohort (n= 926). - S15: Comparison of predictive confusion matrices for the isolated fishing vessel cohort (n= 926) across the three supervised machine-learning models: - S16: One-way sensitivity analysis profiles illustrating the independent impact of individual criteria variations on target vessel allocations: - S17: Bivariate sensitivity analysis response surfaces illustrating operational metric trade-offs across paired criteria modifications - S18: Refinement scenarios for framework enhancement

Institutions

Categories

Ocean Engineering, Machine Learning, Precautionary Principle, Shipwreck

Licence