Replication data and code for “Separating drainage topology from spatial aggregation: A matched-network benchmark using Brazilian hydropower”

Published: 10 August 2026| Version 1 | DOI: 10.17632/3b3g53ph65.1
Contributor:

Description

This dataset contains the model-stage replication package for the manuscript “Separating drainage topology from spatial aggregation: A matched-network benchmark using Brazilian hydropower.” The archive includes frozen derived analytical inputs, reduced monthly ERA5-Land runoff data, HydroBASINS level-6 target and network mappings, six archived false-network families, Python code, locked software environments, provenance and checksum manifests, a data dictionary, and machine-readable expected outputs. It reproduces the central Poisson pseudo-maximum-likelihood estimate, the 999-draw matched false-network benchmarks, the deterministic false-network comparison, the contiguity sensitivity analyses, the external natural energy inflow (ENA) validation, and the principal coefficient-distribution figures. The study covers 84 Brazilian hydropower entities linked to 52 initial targets across 17 river networks from 2000 to 2025. The primary model begins with 21,692 entity-month observations. The documented fixed-effect separation rule removes 91 observations, leaving 21,601 effective observations, 82 entities, 50 targets, and 16 networks. The first-order contiguity comparison uses a 39-target common sample. The package is self-contained for the reported model stage. It does not reconstruct every raw ONS, ERA5-Land, or HydroBASINS source product from the original portals. Instead, it provides the exact derived analytical inputs used for estimation, while the attribution and licensing conditions of the original providers continue to apply. Accumulated runoff represents a spatially aggregated hydrological exposure, not routed discharge or observed plant inflow. The reported generation coefficients are associational and should not be interpreted as causal effects.

Files

Steps to reproduce

1. Download the complete dataset while preserving the original folder structure. If Mendeley delivers the files as a compressed archive, extract it before running the workflow. 2. Open a terminal in the package root directory containing README.md, environment.yml, and the 01_data/ through 05_outputs/ folders. 3. Create the locked Conda environment: conda env create -f environment.yml 4. Activate the environment: conda activate brazil-hydro-replication 5. Run the verified quick reproduction workflow: python 03_code/run_all.py This command verifies archived checksums and schemas, re-estimates the primary PPML model, reconstructs the first primary false-network draw, reproduces the deterministic-network and ENA analyses, reconstructs Supplementary Tables S5–S8 and the 999-draw ENA benchmark in Figure S6, and regenerates the principal summaries and figures. 6. Inspect the newly generated files in 05_outputs/. The mapped-network PPML coefficient should be approximately 0.1110261009 for an effective sample of 21,601 observations. Figure S6 should report a mapped-network correlation of approximately 0.666142, a placebo median of approximately 0.559554, and an empirical p-value of 0.001. 7. To re-estimate all randomized draws in the six archived families, run: python 03_code/run_all.py --full-randomization The full workflow is computationally intensive and is expected to require approximately 20–30 minutes, depending on the hardware.

Institutions

Categories

Hydrology, Spatial Analysis, Hydroelectric Power Generation

Funders

Licence