Processed optical-flow descriptors and reproducibility materials for benchmarking RAFT, FastFlowNet, and Farnebäck on the River Dart dataset
Description
This dataset contains processed optical-flow descriptors, machine-learning inputs, statistical outputs, and reproducibility materials supporting a comparative evaluation of RAFT, FastFlowNet, and Farnebäck for image-based river surface-velocity estimation using the open River Dart dataset. For each optical-flow method, separate calibration and independent temporal-validation descriptor tables are provided. The common descriptors used for Random Forest regression include mean optical-flow magnitude, standard deviation of magnitude, 75th and 99th percentiles of magnitude, and mean horizontal and vertical flow components. The repository also includes temporal cross-validation definitions, model-selection outputs, out-of-fold and independent validation predictions, performance metrics, bootstrap confidence intervals, pairwise statistical comparisons, permutation feature importance, feature-ablation results, velocity-range analyses, computational-demand summaries, and code to reproduce the downstream analysis and manuscript figures. The calibration set contains 4,241 observations and the independent validation set 5,910 observations. Validation data were not used for hyperparameter selection. Raw River Dart videos are not redistributed. They remain available from the original Newcastle University dataset by Perks, DOI: 10.25405/data.ncl.19762027. This repository begins at the processed descriptor stage; the complete video-level optical-flow extraction implementation is outside its reproducibility boundary.
Files
Steps to reproduce
1. Download and unzip the dataset. Do not rename the folders. 2. Create a Python 3.13 environment and install the pinned packages: python -m pip install -r requirements.txt 3. Verify that the download is complete: python code/verify_repository.py 4. Reproduce the primary analysis (writes to results/): python code/reproduce_analysis.py 5. Reproduce the temporal-dependence sensitivity analysis (writes to results/temporal_sensitivity/): python code/temporal_sensitivity_analysis_v2.py 6. Reproduce the post hoc diagnostics, files 14 to 18: python code/07_scope_mask_and_autocorrelation_diagnostics.py 7. Regenerate the figures: python code/generate_publication_figures.py Optional integrity check (macOS or Linux): shasum -a 256 -c CHECKSUMS_SHA256.txt NOTES The optical-flow stage is not rerun. All scripts start from the archived descriptor tables, so the original River Dart videos are not required. Run every script from the repository root, not from inside code/. Exact reproduction of every digit requires scikit-learn 1.7.2. Random-forest split selection resolves near-ties differently across library versions and hardware platforms; this affects the fourth decimal of the RAFT validation R2 only. MAE and RMSE are unaffected at the precision reported in the manuscript. The environment used for the deposited outputs is recorded in results/17_READ_ME_DIAGNOSTICS.txt and results/analysis_environment.json.
Institutions
- Centro de Investigación en Materiales AvanzadosChihuahua, Chihuahua City
Categories
Funders
- Secretaría de Ciencia, Humanidades, Tecnología e InnovaciónMexico City, Mexico CityGrant ID: Ph.D. scholarship No. 930739