Data and Code for Transfer-Aware Multi-Contaminant Groundwater-Quality Screening and National-Scale Monitoring Prioritization

Published: 16 June 2026| Version 1 | DOI: 10.17632/9mnk4gkm6t.1
Contributor:
QIANGQIANG TAO

Description

This dataset provides the curated data products, model-ready matrices, derived results, and reproducible code supporting a transfer-aware, multi-contaminant groundwater-quality screening framework. The study integrates public groundwater-quality observations, benchmark exceedance labels, hydroclimatic variables, soil properties, land-cover indicators, and spatial validation outputs to evaluate contaminant-specific and multi-contaminant risks across monitored groundwater systems. The archive is organized into three main components. The raw_data folder provides source links, access notes, and source inventories for the public datasets used in the study; full raw source files are not redistributed because they remain governed by their original providers. The result_data folder contains processed analytical tables, model inputs, trained-model outputs, validation summaries, attribution results, and monitoring-priority support files. The core_code folder contains the main scripts, environment specifications, and reproduction notes needed to rebuild the processed datasets and model outputs from the documented source data. These materials are intended to support transparency, reproducibility, and reuse of the analytical workflow. They can be used to reproduce the main model-development, validation, interpretation, and monitoring-prioritization steps, or to adapt the workflow to related groundwater-quality screening and surveillance applications.

Files

Steps to reproduce

1. Download the archive and extract all files while preserving the folder structure. 2. Review README.md, README_MENDELEY_DATASET.md, DATA_SOURCES.tsv, and raw_data/RAW_DATA_LINKS.tsv to identify the public source datasets and access notes. 3. Obtain the original public source data from the provider links listed in raw_data/RAW_DATA_LINKS.tsv. Place the downloaded source files in the expected locations described in DATA_SOURCES.tsv and core_code/REPRODUCE.md. 4. Create the computational environment using core_code/environment.yml or core_code/requirements.txt. 5. Run the pipeline scripts in core_code/pipeline following the order described in core_code/REPRODUCE.md. The workflow rebuilds the processed groundwater-quality matrices, benchmark exceedance labels, model-development datasets, validation summaries, attribution outputs, and monitoring-priority support files. 6. Compare regenerated outputs with the files provided in result_data. File integrity can be checked using CHECKSUMS_SHA256.tsv, and package contents can be checked using MANIFEST.tsv.

Categories

Earth Sciences, Environmental Science, Environmental Chemistry, Environmental Monitoring, Water Science, Data Science, Machine Learning

Licence