Data and model outputs for “Freeze-Thaw Effects on RCA: A Comparative Study of Semi-Empirical and Machine Learning Models”

Published: 27 July 2026| Version 1 | DOI: 10.17632/8cndwjw4y4.1
Contributors:
,

Description

This dataset supports the manuscript “Freeze-Thaw Effects on RCA: A Comparative Study of Semi-Empirical and Machine Learning Models.” It contains 90 sequence-level resilient-modulus observations from a previously reported experimental programme involving six independently prepared recycled concrete aggregate base specimens assigned to 0, 1, 3, 5, 10, and 20 freeze–thaw cycles. Fifteen measurement stress sequences were applied to each specimen, and one sequence-level resilient-modulus value was calculated as the arithmetic mean of cycles 96–100. No new freeze–thaw or resilient-modulus experiments were conducted specifically for the present manuscript. The deposit includes the fixed 63-observation training and 27-observation test partition used for all four models, final predictions and observation-level errors for the semi-empirical model, Linear Regression, Random Forest Regressor, and Gradient Boosting Regressor, model parameters, common test metrics, condition-specific errors, error distributions, semi-empirical optimisation diagnostics, start-point stability, and 1,000 observation-level bootstrap fits. The semi-empirical and linear-regression predictions can be recalculated from the deposited variables and parameters. RFR and GBR predictions are provided as exported final outputs. Candidate-level cross-validation scores and fold assignments are independent reconstructions based on the fixed training subset and stated software settings, not original serialized GridSearchCV objects. Because the 15 stress-sequence observations within each freeze–thaw condition were repeated measurements from one physical specimen, the results represent internal sequence-level hold-out performance rather than independent-specimen or external validation.

Files

Steps to reproduce

1. Download all files into the same directory. 2. Use RCA_FT_sequence_level_data_and_predictions.csv as the machine-readable source table. 3. Run verify_reported_metrics.py with Python 3. The script uses only the Python standard library. 4. The script verifies the 90-row dataset size, the 27-row fixed test subset, the semi-empirical and linear-regression formulas, and the R², MAE, RMSE, and MAPE values reported for all four models. 5. Use the two Excel workbooks for formula-linked calculations, semi-empirical calibration diagnostics, bootstrap results, reconstructed cross-validation documentation, and detailed interpretation notes.

Categories

Civil Engineering, Machine Learning, Pavement, Geotechnics, Semi-Empirical Model

Licence