A Linked Dataset for Analyzing Educational Trajectories from Secondary School to Higher Education in Colombia

Published: 30 July 2026| Version 1 | DOI: 10.17632/jg6kkv8mg6.1
Contributors:
Enrique De La Hoz,
, Diego Polanco

Description

This dataset provides a curated, student-level linkage of Colombia's two national standardized examinations: SABER 11 (end of secondary education) and SABER PRO (end of undergraduate studies), both administered by ICFES with census coverage. It contains 332,990 linked student records and 32 variables covering demographics, high school of origin, SABER 11 baseline scores (five subjects, 0–100; global 0–500), higher-education context (361 institutions and 4,249 SNIES programs, with official codes as authoritative keys), and SABER PRO outcome scores (five generic modules, 0–300). SABER 11 cohorts span 2014–2021 and SABER PRO cohorts 2016–2023, with a median interval of 6 years between examinations. The release follows a documented curation pipeline (identifier removal, derived top-coded ages, recoding of non-scored essays, label harmonization) and is accompanied by a bilingual Spanish/English machine-readable codebook. The dataset supports research on academic value added, institutional effectiveness, equity, and educational data mining, and has already underpinned peer-reviewed recommender systems (Applied Sciences 2024, 14, 8311; Electronics 2025, 14, 4121). Derived from ICFES open microdata; distributed under CC BY 4.0 with attribution to ICFES.

Files

Steps to reproduce

The dataset was constructed from the official ICFES open-data platform (DataIcfes), which distributes anonymized microdata per administration period. (1) Period files were downloaded for SABER 11 (2014-2 to 2021-1) and SABER PRO (2016 to 2023). (2) Records were deterministically linked at the student level using the SABER 11 registration reference included in SABER PRO records, yielding 332,993 records and 324 columns. (3) A documented curation pipeline reduced the schema to 32 variables: removal of direct identifiers and raw birth dates (replaced by derived, top-coded age variables), exclusion of variables with structural missingness above 90%, resolution of merge duplicates, and exclusion of cohort-specific modules. (4) Three records with column-shift artifacts were removed (final n = 332,990), non-scored Written Communication essays (score = 0) were recoded to missing, and categorical labels were harmonized. (5) Range, temporal-coherence, and key-integrity checks were run before export. The full pipeline is implemented in the R script prepare_release_v2.R, included in this repository; the accompanying bilingual codebook documents every variable, its coding, and known data-quality flags.

Categories

Education, Academic Assessment, Academic Performance

Licence