Data for: Prediction Without Intervention? A Systematic Review and Empirical Benchmark of Learning-Analytics Early-Warning Systems for Student Retention in Higher Education

Published: 23 June 2026| Version 1 | DOI: 10.17632/zjtd55y9b5.1
Contributor:

Description

This dataset contains the underlying data for a study that pairs a PRISMA 2020 systematic review with an original, reproducible empirical benchmark on learning-analytics early-warning systems for student retention and dropout in higher education. The review component comprises the bibliographic search log, the corpus of records retrieved from the OpenAlex database (2015-2026), the record-level screening decisions (include/exclude, analytic-and-decision category, study design, outcome-reporting flag, relevance), the list of 176 included studies, the prioritised synthesis set, the subset of studies reporting a retention outcome, the PRISMA flow counts, and the computed descriptive bibliometric results. The benchmark component contains the public student dataset used for the empirical analysis (Realinho et al., 2022; 4,424 students, 36 features) together with all computed results: cross-validated predictive performance by information tier and model, calibration, subgroup fairness with bootstrap confidence intervals, capacity-limited targeting, and a robustness check. Four figures summarising the review funnel and benchmark results are included. All bibliographic metadata derive from the open OpenAlex database (CC0). The benchmark dataset is redistributed under its original CC BY 4.0 licence with attribution. No personal or confidential data are included.

Files

Steps to reproduce

1. Systematic search: query the OpenAlex works API over titles and abstracts with the twelve search strings in search_log.json, restricted to journal and conference articles in English, 2015-2026, retrieving up to 120 records per query. De-duplicate by work identifier (642 records -> 427 unique). 2. Screening: set aside records without a retrievable abstract (n = 32); screen the remaining 395 on title and abstract against the inclusion criteria, recording for each record an include/exclude decision, the analytic-and-decision category, the study design, and whether a retention/persistence/completion outcome is reported. This yields 176 included studies (screening_decisions.json, included_studies.json). 3. Bibliometrics: compute descriptive statistics over the included set (year, category, design, outcome-reporting share, analytic methods, data sources, geography, venues, citations, explainability and fairness coverage) -> stats.json. 4. Benchmark: using dataset.csv, define the binary outcome (Dropout vs non-dropout); build three nested feature sets (at enrolment, after first semester, after first year); train logistic regression, random forest, and gradient boosting; evaluate with stratified five-fold cross-validation and a 30 percent held-out test set, computing AUC, calibration, subgroup error rates with bootstrap confidence intervals, and capacity-limited targeting -> empirical_results.json. 5. Analysis scripts that regenerate every file are available from the author on request.

Institutions

Categories

Educational Technology, Machine Learning

Licence