Generative AI competency of Arabic language and literature students: a four-dimension questionnaire dataset with open-ended responses

Published: 21 July 2026| Version 1 | DOI: 10.17632/xd27t4g547.1
Contributor:
rabie ramadan

Description

This dataset contains survey responses from 145 university students on their competency in using generative AI for the study of Arabic language and literature. Participants were students [in the Arabic Language Department at Sultan Qaboos University, Oman] and completed a structured questionnaire [administered online/in print] in [month, year]. The sample comprises 132 Bachelor's, 10 Master's, and 3 Doctorate students. The questionnaire captures five demographic variables — study level, grade-point-average band, prior AI training (yes/no), and weekly frequency of AI use for study — followed by 28 closed items rated on a five-point Likert scale (5 = strongly agree to 1 = strongly disagree). These items are organized into four competency dimensions: Awareness & Availability (6 items), Functional Application (8 items), Prompt Engineering & Critical Appraisal (6 items), and Aspiration & Vision (8 items). The Functional Application and Prompt Engineering items are domain-specific, covering tasks such as phonetic and morphological analysis, syntactic parsing (iʿrab), rhetorical and literary text analysis, prompt writing, and source verification for Arabic content. The instrument concludes with seven open-ended questions, answered in Arabic, on desired tools, perceived challenges and solutions, dependency versus innovation, most useful tools, negative experiences, and skills needed to use AI effectively. The data are provided as a single CSV file (145 rows × 40 columns) accompanied by a codebook (codebook.csv) that maps every variable to its dimension, response type, coding scheme, and the original Arabic item text with an English translation. Records are fully anonymized: participants are identified only by sequential codes (P001–P145), and no names or other identifying information are included. Blank cells denote missing responses. The dataset supports analyses of students' awareness, use, critical appraisal, and aspirations regarding generative AI in Arabic-language education, including relationships with academic performance, prior training, and usage frequency, as well as qualitative analysis of the open-ended responses. It may be reused for comparative studies across institutions or languages, instrument validation, and research on AI adoption in humanities education.

Files

Steps to reproduce

Instrument design. A structured questionnaire was developed to measure students' competency in using generative AI for Arabic language and literature study. It comprised five demographic items (study level, GPA band, prior AI training, weekly AI-use frequency) and 28 closed statements grouped into four dimensions — Awareness & Availability (6 items), Functional Application (8 items), Prompt Engineering & Critical Appraisal (6 items), and Aspiration & Vision (8 items) — each rated on a five-point Likert scale (5 = strongly agree to 1 = strongly disagree), plus seven open-ended questions. Full item wording (Arabic with English translation) is provided in the accompanying codebook. Instrument validation. [Describe how content validity was established — e.g. the items were reviewed by N expert reviewers in Arabic linguistics / educational technology, and revised accordingly; and/or a pilot test was run with N students to check clarity and internal consistency.] Population and sampling. The target population was [students in the Arabic Language Department at Sultan Qaboos University, Oman]. A [convenience/random/stratified] sample was recruited, yielding 145 respondents (132 Bachelor's, 10 Master's, 3 Doctorate). Ethics and consent. [State ethics approval — e.g. approval was obtained from the relevant institutional review board, reference number XXX — and that participants gave informed consent before taking part and participation was voluntary and anonymous.] Data collection. The questionnaire was administered [online via Google Forms / Microsoft Forms — or on paper] during [month, year]. Respondents completed all sections in Arabic. Data extraction and coding. Responses were exported from [the survey platform] to CSV. Likert responses were coded 1–5 as above; categorical demographics were retained as labels; the seven open-ended questions were kept as verbatim Arabic text. Blank cells denote missing responses. Anonymization. All direct identifiers were removed. Each respondent was assigned a sequential code (P001–P145); no names, emails, or contact details are included in the released file. Cleaning and verification. [Describe any cleaning — e.g. duplicate submissions were removed, out-of-range values checked, and the final file verified to contain 145 rows × 40 columns.] Files released. The final dataset is provided as arabic_genai_competency_data.csv (145 × 40) together with codebook.csv, which maps each variable to its dimension, response type, coding scheme, and Arabic/English item text.

Institutions

Categories

Student, Generative Artificial Intelligence

Licence