Emotion Vocabulary Dataset from Beginner-Level Italian Textbooks: Valence, Arousal, and Pedagogical Annotation
Description
This dataset contains the coded lexical and contextual data used in a corpus-informed analysis of emotion-related vocabulary in Italian as a foreign language (IFL) textbooks at the beginner level (CEFR A1) used in Polish lower secondary education. The dataset is based on four textbooks from two series: Progetto Italiano Junior (Marin, 2017; 2018) and Va bene! (Kaliska & Kostecka-Szewc, 2021). All materials were analysed in their printed form. The dataset includes all lexical items identified as emotion-related based on their presence in the ANEW-IT database (Montefinese et al., 2014), which provides normative ratings of valence and arousal for Italian words. Each entry in the dataset corresponds to a single token occurrence of an emotion-related lexical item. The dataset includes both surface forms as they appear in the textbooks and their corresponding base forms (lemmas) as listed in ANEW-IT. For each token, the dataset provides contextual, linguistic, and affective information. The variables included are: • textbook and series identification • unit/chapter, page number, and exercise reference • material type (e.g., dialogue, reading text, exercise, review) • pedagogical focus (e.g., grammar, vocabulary, comprehension, mixed) • word form (surface form) and lemma • part of speech (POS) • contextual sentence or description of occurrence • frequency measures (FreqColfis, Ln_Colfis) • affective ratings (valence and arousal, scale 1–9) The dataset enables replication of the quantitative analyses reported in the study, including token counts, type–token ratios, valence and arousal distributions, and comparisons across textbook series and pedagogical contexts.
Files
Steps to reproduce
1. Four Italian as a foreign language (IFL) textbooks at CEFR A1 level were selected based on their official approval for use in Polish lower secondary education. 2. The printed textbooks were analysed manually, unit by unit, to identify all lexical items occurring in the instructional materials. 3. Each lexical item was cross-checked against the ANEW-IT database (Montefinese et al., 2014), which provides normative ratings of valence and arousal for Italian words. 4. Only items present in ANEW-IT were retained. Each occurrence of such an item in the textbooks was recorded as a separate token. 5. For each token, contextual and linguistic information was annotated, including textbook source, unit, page, material type, pedagogical focus, surface form, and corresponding lemma. 6. Affective (valence, arousal) and frequency-related variables (FreqColfis, Ln_Colfis) were retrieved from the ANEW-IT database and assigned to each lemma. 7. Inclusion and exclusion criteria were applied systematically: items occurring only in instructions, listening-only materials, or non-standard formats were excluded; inflected forms were included if the base form appeared in ANEW-IT; homonymous items were retained regardless of contextual interpretation. 8. The resulting dataset consists of all identified emotion-related tokens with their associated contextual, linguistic, and affective annotations.
Institutions
- University of WarsawMazovia, Warsaw