AI-Assisted Syllabification and Tone Annotation Datasets for South-South Nigerian Languages
Description
This dataset includes lexical, syllabic, and tonal data from 25 South-South Nigerian languages. To achieve the data, phonotactic rules and tone patterns were semantically interpreted by GPT-4 to build syllabification and tone-labelling engines. The process followed a human-in-the-loop framework ensuring accurate rule execution across languages. Contents: 1) Language_Profile_Wordlist.xlsx: Original language profile/template and wordlists 2) Language_Profile_Wordlist_Updated.xlsx: Updated (corrected/edited) profile/template and wordlists 3) wordlist.xlsx: Swadesh lexical list (108 English words) used as the primary input for documenting the 25 languages. 4) syllabified_output_result.zip: Syllabified outputs (without corrected/edited profile and wordlists) annotated without human-in-the-loop validation labels across 25 languages. 5) syllabified_output_result_with_HILV.zip: Syllabified outputs (without corrected/edited profile and wordlists) annotated with human-in-the-loop validation labels across 25 languages. 6) syllabified_output_result_updated.zip: Updated syllabified outputs (with corrected/edited profile and wordlists) per language. 7) structure_transformed_updated.zip: Syllabified outputs further transformed into syllable structure representations (e.g., V-CV, N-CVV) with a transformed syllable structure column (N_CV Structure) 8) tone_labelling_output_result.zip: Wordlists with auto-generated tone labels without human-in-the-loop validation. 9) tone_labelling_output_result_updated.zip: Wordlists with auto-generated tone labels with human-in-the-loop validation (edited/corrected tone labels). Applications: 1) Phonological typology: Contents 3, 4, 5, 6, and contents 7, 8, provide syllabified and tone-annotated wordlists across 25 Nigerian languages, enabling comparative analysis of phonotactic constraints, syllable structures (e.g., CV, CVV, N-CVC), and tonal systems. 2) Cross-linguistic NLP development: By offering consistent syllabification and tone annotations, the dataset would facilitate the design and evaluation of NLP tools that can generalize across typologically diverse languages. Use cases include multilingual grapheme-to-phoneme converters, syllable-aware machine translation models, and pronunciation-aware text-to-speech systems. 3) Low-resource language modeling: The dataset could serve as a valuable resource for training and evaluating AI models in low-resource settings. 4) Tonal morphology and suprasegmental analysis: The tone-labelled column in the dataset allow for the study of tonal morphology, such as tone sandhi, tone disambiguation, and morphotonemic alternation. Format: All files are provided in .xlsx and .zip formats for ease of integration into language modeling pipelines.
Files
Steps to reproduce
Overview This dataset was developed through a structured, linguistically-informed and computationally scalable process targeting twenty-five under-resourced South-South Nigerian languages. The data acquisition workflow combined traditional linguistic methods with automated processing pipelines powered by a human-in-the-loop framework integrating GPT-4 for rule interpretation and engine logic generation. 1. Language Selection Validated languages were selected from Ekpenyong et. al. (2025) 2. Language Profile Construction Each language profile was documented in a structured spreadsheet format and included: a) Speech and sound inventories (Vowels, Consonants, Syllabic Nasals) b) Valid syllable structures (e.g., CV, V, CVC) c) NLP syllabification rules and constraints d) Tone pattern specifications and syllabic alignment logic 3. Rule Interpretation and Engines Design GPT-4 was employed to semantically parse and transform the human-written NLP rules into executable logic. Rather than using static templates, the engines adaptively interpreted each language’s profile to enforce phonotactic constraints, detect reduplications, and apply rule-by-rule syllable parsing. Rules were executed in ordered passes, restarting from Rule 1 after each successful syllable break. 4. Automated Workflow and Output Generation Two core Python engines were developed: a) A syllabification engine, which hyphenates transcribed words based on language-specific constraints b) A tone labelling engine, which aligns tone marks with syllables using defined tone templates These engines were deployed via Google Colab notebooks for transparency and reproducibility. Processing preserved the fidelity of the language profile and ensured no rule logic was hardcoded. 5. Output Format and Validation Generated outputs support iterative correction, data refinement and further computational processing. 6. Reproducibility and Reuse All workflows are reproducible using the provided Excel language profiles, source code, and wordlists. The system is fully language-specific and extensible to other under-resourced language contexts. Reference Ekpenyong, M. et al. (2025). Language Ecology and Endangerment Dataset for South-South Nigeria. Mendeley Data, V7, https://doi.org/10.17632/sp23gbtvsj.7