Cross-reservoir Sentinel-1/Sentinel-2 shoreline delineation benchmark: manual labels, model predictions, and evaluation outputs

Published: 3 September 2026| Version 1 | DOI: 10.17632/67yd767ydy.1
Contributor:

Description

This dataset supports a leakage-controlled cross-reservoir benchmark of shoreline delineation using Sentinel-1, Sentinel-2, and optical–SAR fusion data. It contains manual shoreline reference masks for 15 reservoir states across five Mexican reservoirs—Cointzio, Queréndaro, Soledad, Esperanza, and Mata—together with scene-pairing registries, derived feature metadata, model predictions, evaluation metrics, multi-seed results, common-state definitions, global and reservoir-adaptive rankings, canonical method-selection records, and quality-control outputs. The benchmark compares 43 configurations spanning spectral indices, Sentinel-2 Scene Classification Layer baselines, SAR thresholding, Random Forest, XGBoost, U-Net, SegFormer, Mask2Former, and SAM 2. Fourteen states were eligible for optical evaluation, while 10 temporally matched states—two per reservoir—formed the strict common subset used for controlled optical, SAR, and fusion comparisons. The full benchmark comprises 505 scene–method evaluations. The dataset also includes screening-level transfer outputs for Little Rock Reservoir, Randy Poynter Lake, and Croton Falls Reservoir in the United States. These external sites lack temporally matched manual shoreline labels and therefore must not be interpreted as supervised external validation. The materials are intended to support reproducibility, independent inspection of the benchmark protocol, comparison of shoreline-delineation approaches, and development of leakage-controlled cross-reservoir evaluation strategies. Bathymetric reconstruction, reservoir-storage estimation, hydrometric validation, and downstream uncertainty analysis are outside the scope of this dataset. Raw Sentinel satellite imagery is not redistributed; scene identifiers and processing registries are provided to support data retrieval from the original public services.

Files

Steps to reproduce

The dataset can be reproduced using the accompanying Python Jupyter Notebook and configuration files. Raw Sentinel imagery is not redistributed and must be retrieved from the Microsoft Planetary Computer using the scene identifiers and reservoir definitions provided in the dataset. First, install the Python dependencies listed in the environment or requirements file, including NumPy, pandas, Rasterio, GeoPandas, SciPy, scikit-learn, XGBoost, PyTorch, Transformers, and the official SAM 2 package. Download the SAM 2.1 Hiera Base Plus checkpoint and update its local path in the notebook configuration. Configure the reservoir definitions and analysis periods in reservoirs_master.csv. Then execute the notebook sequentially to: (1) retrieve and preprocess Sentinel-2 L2A and Sentinel-1 RTC observations; (2) generate optical, SAR, and optical–SAR feature stacks; (3) construct label-anchored benchmark states; (4) train or execute the spectral-index, machine-learning, deep-learning, transformer, and SAM 2 methods; (5) apply leave-one-reservoir-out validation and multi-seed experiments; and (6) calculate per-state metrics, strict common-state summaries, global and reservoir-adaptive rankings, canonical selections, and external screening diagnostics. Sections must be executed in their numbered order because downstream analyses consume registries and files generated by previous sections. Random seeds, temporal tolerances, quality-control thresholds, model hyperparameters, probability thresholds, selection coefficients, and tie-breaking rules are declared explicitly in the notebook. The expected benchmark contains 15 manual shoreline states, 14 optical-eligible states, 10 strict common states, 43 method configurations, and 505 scene–method evaluations.

Categories

Computer Science, Earth Sciences, Water Science, Remote Sensing, Environmental Challenge

Funders

Licence