Ltpp rutting analysis dataset
Description
Analysis-ready records underlying the article "Leakage-free multi-horizon forecasting of asphalt pavement rutting from LTPP field data: interpretable machine learning ensembles with conformal prediction intervals". 10,549 transverse-profile rutting surveys from 1,354 in-service flexible pavement section-construction cycles, monitored between 1989 and 2023 across 58 states and provinces, drawn from six Long-Term Pavement Performance (LTPP) modules via InfoPave. Each record carries the maximum rut depth of the survey, the prediction target, and fifteen physical predictors spanning traffic, climate, structure, material, mechanical response and elapsed time. The unit of analysis is the section-construction cycle, not the physical pavement: a rehabilitated section contributes one unit per construction cycle, each with its own age origin. Every time-varying quantity was accumulated only up to the survey date, so no information postdating a survey enters its record. Five binary flags mark every imputed cell, allowing imputed values to be restored to missing before any leakage-free re-analysis. Included: the dataset in XLSX with an embedded variable dictionary, the same records in CSV, and a README documenting provenance, units, construction rules and caveats. Section identifiers are text because leading zeros are significant; the dynamic modulus is in MPa. Source: LTPP program, Federal Highway Administration (https://infopave.fhwa.dot.gov).
Files
Steps to reproduce
A. How the deposited file was produced 1. Extraction. Six Long-Term Pavement Performance (LTPP) modules were downloaded from InfoPave: the rutting transverse-profile survey module supplying the target, and the traffic, climate, structure, material and mechanical-response modules supplying the predictors. 2. Merge with temporal alignment. Modules were joined on the LTPP section identifier and the survey date. Cumulative quantities, above all cumulative equivalent single-axle loads and cumulative freeze-thaw cycles, were accumulated only up to each survey date, so that no value postdating a survey enters that survey's record. 3. Duplicate aggregation. Records sharing a section and a date were averaged, reducing 10,555 raw records to the 10,549 deposited here. 4. Series delimitation. Monitoring series were split by LTPP construction number: a rehabilitation closes one series and opens another, and the pavement-age origin resets with it. Rut decreases occurring inside a cycle were retained as recorded, not excised. 5. Bounding and imputation. Base-layer thickness was winsorized at its 1500 mm physical bound. Missing material and mechanical attributes were filled with the median of the pavement family (asphalt concrete over asphalt-treated base, over unbound base, over treated base, over Portland-cement concrete). Missing climatic and traffic attributes, all below 1 percent, were filled section-wise forward then backward, then with the global median. Every imputed cell is marked by its companion flag column. 6. Units. The dynamic modulus was converted from kPa to MPa so that the values match the article. B. How to reproduce the analysis reported in the article from this file 1. Set each material or mechanical value back to missing wherever its flag column equals 1, so that imputation can be re-estimated inside the resampling loop rather than globally. 2. Form horizon pairs. For each horizon of one to five years, pair every survey with the survey nearest that horizon later within the same section_id, accepting the pair only inside a tolerance window of plus or minus 0.6 years for horizons of one to three years and plus or minus 0.8 years for four and five. The article reports 5,947, 5,942, 5,042, 4,734 and 3,887 usable pairs. 3. Set the learning target to the rut-depth increment between the two surveys, not the absolute depth, and withhold the origin depth from the predictors. 4. Partition with ten-fold cross-validation grouped on section_id, so that every survey of a section falls in the same fold. Estimate imputation medians and standardization constants on the training folds alone and apply them unchanged to the held-out fold. 5. Fit the five learners, reconstruct the absolute forecast by adding the origin depth to the predicted increment, and score on that reconstructed depth.
Institutions
- Al-Qasim Green UniversityBābil, al-Qasim