KurdishMCQ: A multiple-choice question dataset for the Kurdish language (Sorani dialect)

Published: 8 July 2026| Version 1 | DOI: 10.17632/z8z28jhszp.1
Contributors:
,

Description

This dataset is the first large-scale, publicly available collection of multiple-choice questions in the Kurdish language (Sorani dialect), helping fill an important gap in evaluating large language models (LLMs) for low-resource languages. It includes 17,233 questions gathered from Kurdish middle and high school materials, such as textbooks and official exams, all written in Sorani using the Perso-Arabic script. The dataset was carefully created and reviewed by native Kurdish speakers to ensure accuracy and consistency. Each entry follows a clear and structured format, including fields such as the question text, four answer options (A–D), the correct answer label, and metadata like subject, grade level, and group (e.g., STEM). The dataset covers a wide range of subjects, with the largest portions coming from Biology (2636 questions), Chemistry (2326), Kurdish grammar (1591), Natural Science (1537), and Kurdish literature (1488), along with other areas like Physics, Math, History, and Economics. In terms of grade distribution, most questions are from grade 12 (9453), followed by grades 9, 10, and 11, with a smaller portion categorized as “other”. This resource can be used for tasks like question answering, knowledge evaluation, and reasoning, and is well-suited for training, fine-tuning, and benchmarking language models, especially in multilingual and low-resource settings.

Files

Institutions

Categories

Natural Language Processing, Large Language Model

Licence