Longitudinal eGFR Trajectory and Survival Benchmark Dataset for Chronic Kidney Disease (SMART-C Cohort)
Description
1. Scientific Overview & Objective This benchmark dataset, developed under the Reproducible Longitudinal Dynamical Patterns (RLDP) framework, provides high-density, multi-year clinical time-series for evaluating continuous-time trajectory forecasting, state-space estimation, and out-of-sample prognostic risk models in Chronic Kidney Disease (CKD). Longitudinal modeling of renal function requires overcoming significant observational challenges, including sparse and irregularly spaced visits, analytical measurement noise, and non-linear trajectory decay. This dataset establishes a standardized, fully reproducible benchmark for machine learning and dynamical systems in nephrology. 2. Cohort Specification & Clinical Design The dataset comprises N = 5,000 adult CKD patients prospectively tracked across a 3-year timeline (Days 1 to 1,096), generating 41,800 longitudinal laboratory evaluations (mean: 8.36 visits per patient; range: 2 to 14 visits). The primary continuous biomarker is Estimated Glomerular Filtration Rate (eGFR, mL/min/1.73m2). Baseline eGFR across the cohort is 55.70 +/- 17.97 mL/min/1.73m2 (range: 6.0 to 115.0 mL/min/1.73m2). Randomization is structured 1:1 into active intervention (n = 2,501) and control placebo (n = 2,499), stratified at screening by baseline disease severity into three clinical strata: - Screening eGFR 30 to <45 mL/min/1.73m2 (n = 1,497, 29.9%) - Screening eGFR 45 to <60 mL/min/1.73m2 (n = 1,466, 29.3%) - Screening eGFR 60 to <90 mL/min/1.73m2 (n = 2,037, 40.7%) Baseline concurrent GLP-1 receptor agonist therapy is documented in 24.9% of participants. 3. File Structure & Data Dictionary The archive contains two harmonized CSV files: - baseline.csv (5,000 records): usubjid (patient ID), trt01pn (arm: 1=active, 0=placebo), randfl (randomized), ittfl (intent-to-treat), trtfl (treated), blgfr (baseline eGFR), blglp1 (baseline GLP-1 status), strata (screening renal category). - followup.csv (41,800 records): usubjid, randfl, ittfl, trtfl, anl01fl, aval (observed visit eGFR), base (baseline eGFR), paramcd (GFRBSCRT), param, ady (analysis day: 1 to 1096), avisitn, avisit (protocol visit name: BASELINE, WEEK 3, WEEK 13, etc.). 4. Data Quality & Harmonization Audit Automated audit confirms: (a) 0.0% single-visit static records (all multi-year trajectories up to 3 years); (b) 0.0% acute ICU stays (<= 14 days); (c) 0.0% pediatric records (age >= 18); (d) hyperfiltration capped at eGFR <= 140; and (e) 0.0% missingness on protocol eGFR. 5. Intended Benchmark Applications Designed for benchmarking: (1) Sparse Functional Data Analysis (PACE / FPCA) via BLUP conditional expectation; (2) Continuous-Time State-Space Kalman Filtering (CT-SSM) with Maximum Likelihood estimation and closed-form uncertainty bounds; (3) Empirical Bayes Linear Mixed-Effects Models (EB-LMM); (4) Neural Velocity ODEs; and (5) Out-of-sample survival models for severe renal failure progression (eGFR < 30 mL/min/1.73m2).
Files
Steps to reproduce
1. Dataset Extraction & Environment Setup: - Download and extract the dataset archive (SMART_C_Longitudinal_CKD_Benchmark_Dataset.zip). - Ensure Python 3.12+ is installed with standard scientific dependencies: numpy (>=2.5.0), pandas (>=2.3.0), scipy (>=1.14.0), scikit-learn (>=1.5.0), lifelines (>=0.30.0), and pytorch (>=2.13.0). 2. Data Loading & Stratified Patient-Level Partitioning: - Load 'baseline.csv' (N = 5,000) and 'followup.csv' (41,800 records). - Execute a patient-level stratified holdout split based on screening eGFR categories (strata): * Training Cohort: 60% (n = 3,000 patients, 25,103 visits) * Validation Cohort: 20% (n = 1,000 patients, 8,362 visits, 75 events) * Locked Test Cohort: 20% (n = 1,000 patients, 8,335 visits, 87 events) - Use a deterministic random seed (SEED = 42) for exact reproducibility. 3. Sparse Functional Data Analysis (PACE) Pipeline: - Step 3.1: Estimate the population mean trajectory mu(t) by fitting a quadratic polynomial across all pooled training visits. - Step 3.2: Compute pairwise residual cross-products, exclude diagonal measurement error (sigma_eps^2 = 78.52), and smooth the 2D covariance surface G(s,t) using a 2D Gaussian kernel (bandwidth = 0.5 years). - Step 3.3: Perform eigen-decomposition on G(s,t) to derive orthogonal eigenfunctions phi_k(t) and eigenvalues lambda_k. - Step 3.4: Derive individual BLUP functional scores (xi_ik) via conditional expectation for sparse visit schedules (<= 90 days). 4. Continuous-Time State-Space Modeling (CT-SSM): - Formulate a 2-state continuous linear SDE: dx/dt = A*x(t) + w(t), y(t) = C*x(t) + v(t). - Optimize parameters (damping gamma, process noise Q, measurement noise R) in unconstrained log-space via L-BFGS-B on training trajectories. - Project forward filtered state estimates and propagated covariance P(t) to generate multi-horizon forecasts with +/- 1.96*sigma uncertainty bands. 5. Multi-Horizon & Survival Validation: - Condition models on early visits (t <= 90 days) and evaluate at 180, 365, and 730 days on the locked test set. - Fit multivariable Cox Proportional Hazards on the validation split (penalizer = 0.01) and evaluate out-of-sample Harrell's C-index with 1,000 paired bootstrap resamples.
Institutions
- Pabna University of Science and TechnologyRajshahi Division, Pābna