Data and Code for “Using international public data at data-scarce ready-mixed concrete plants: selecting a source-defined predictor with one 28-day result”

Published: 18 August 2026| Version 1 | DOI: 10.17632/pnndrz3zwh.1
Contributors:
Yuyue Zhang,

Description

This dataset contains the reproducibility materials associated with the manuscript “Using international public data at data-scarce ready-mixed concrete plants: selecting a source-defined predictor with one 28-day result.” The study evaluates whether international public concrete data can support predictor selection for data-scarce ready-mixed concrete plants. Ten Public source blocks were used to construct 1,023 fixed source-defined predictors, and External source units were used to develop and evaluate a one-result predictor-selection workflow. A planned mixture set is used to nominate an X-centroid Sentinel before strength is known; after one 28-day Sentinel result becomes available, the workflow selects an existing predictor without local refitting or calibration. The repository contains the Public and External source registry, processed data that are permitted for redistribution, source data underlying figures and tables, analysis code, and fixed model artifacts required to reproduce the reported analyses. Original third-party datasets that cannot be redistributed should be obtained from their cited sources as identified in the source registry. Historical Local production records are not publicly included because they contain plant-level operational information and are subject to authorization by the data owner. The archived materials correspond to the version of the analysis reported in the associated manuscript.

Files

Steps to reproduce

The repository contains the data, code, configurations, and fixed model artifacts used for the analyses reported in the associated manuscript. To reproduce the analyses: 1. Obtain any original third-party datasets that are not redistributed in this repository from the sources identified in the source registry. 2. Use the supplied preprocessing scripts to construct the harmonized Public and External analysis datasets. 3. Run the predictor-library scripts to construct the 1,023 fixed source-defined predictors from the 10 Public source blocks. 4. Run the External evaluation and selector-development scripts to reproduce the source-utility analyses, engineering baselines, nested leave-one-target-out evaluation, and Fixed One-Sentinel results. 5. Run the figure and table scripts to reproduce the corresponding manuscript figures and numerical tables. The repository README and configuration files specify the software environment, execution order, random seeds, and file dependencies. Local historical production records are not included because they contain plant-level operational information and are subject to authorization by the data owner. Consequently, analyses requiring those restricted records cannot be independently rerun from the public repository alone.

Institutions

Categories

Civil Engineering

Licence