LLM-based proficiency assignment for learner corpora

Published: 1 September 2026| Version 1 | DOI: 10.17632/dw4ksr2yt2.1
Contributors:
Tony Berber Sardinha,

Description

This dataset accompanies the study on LLM-based proficiency assignment for learner corpora. The study addresses the lack of comprehensive proficiency assessment data in learner corpus research and tests whether large language models (LLMs) can provide useful proficiency information for large collections of learner writing. Two approaches were investigated: direct assignment of Common European Framework of Reference for Languages (CEFR) levels and comparative judgment (CJ), in which essays are evaluated relative to calibrated proficiency reference texts. The analyses were conducted on texts from the International Corpus of Learner English (ICLE). The ICLE texts themselves are not included in this repository because they are copyrighted material belonging to the Catholic University of Louvain and cannot be redistributed. The repository therefore contains only derived data and analysis outputs associated with the proficiency assessment procedures.

Files

Categories

Corpus Linguistics

Funders

Licence