Data and code for explainable machine-learning screening of diffuse nutrient pollution across climate and land-use gradients
Description
Version 2 provides processed data, fitted-model artifacts, evaluation outputs, interpretation products, integrity records, and analysis programs supporting the study “Explainable machine learning for screening diffuse nutrient pollution across climate and land-use gradients to inform adaptive watershed management.” It focuses on six surface-water nutrient endpoints: nitrate-N, nitrate+nitrite-N, nitrite-N, orthophosphate-P, total phosphorus (as P), and total nitrogen-N. Built around 474,622 station-year-endpoint observations from 106,400 United States Water Quality Portal monitoring locations, the core dataset pairs every response with a documented 60-predictor matrix. Predictors represent spatial and hydrologic context, CHIRPS precipitation, HYDE land use and population, and FAOSTAT fertilizer inputs. Results cover temporal validation, HUC8-grouped spatial validation, source-specific external evaluation independent of WQP fitting using EEA Waterbase, GEMStat, and GLORICH records, calibration, SHAP-based interpretation, uncertainty diagnostics, and support-filtered HUC8 watershed screening. Reproducibility resources include fitted models and predictions, processed tables, data dictionaries, source documentation, checksums, and a path-independent verifier for integrity checks, model loading, and held-out prediction reproduction. Large upstream source collections remain with their original providers. Retrieval details, provenance records, and processing metadata document construction of the analysis-ready dataset. HUC8 products support prioritization of follow-up monitoring within represented domains, and the drinking-water benchmark tables provide an illustrative research comparison.
Files
Steps to reproduce
Download and extract the archive. From the extracted archive root, verify file integrity by running: shasum -a 256 -c SHA256SUMS Create a Python 3.12 environment and install the dependencies listed in code/requirements.txt. Run the archive-wide verification program with: python code/verify_release.py This program verifies checksums, scientific row counts, endpoint coverage, manifests, and the complete fitted-model inventory. Model loading and held-out prediction reproduction can be checked with: python code/verify_release.py --load-models --prediction-smoke 10 Use FILE_INDEX.tsv, DATA_DICTIONARY.md, DATA_DICTIONARY.tsv, and metadata/environment.md to identify the processed nutrient datasets, predictors, model artifacts, evaluation outputs, interpretation products, and analysis programs. Each program under code/ provides its own command-line guidance through the --help option. The archive supports verification of the six-endpoint nutrient analysis, including temporal testing, HUC8-grouped spatial validation, source-specific external evaluation, calibration, SHAP-based interpretation, uncertainty diagnostics, support-filtered HUC8 summaries, and an illustrative drinking-water benchmark comparison. Large upstream WQP, CHIRPS, HYDE, FAOSTAT, EEA Waterbase, GEMStat, and GLORICH collections are not redistributed. SOURCES_AND_LICENSES.md and the accompanying provenance metadata identify the required sources, retrieval information, processing contracts, and applicable terms for users undertaking a full reconstruction from upstream data.
Institutions
- Wuhan University of Science and TechnologyHubei, Wuhan