Language Ecology and Endangerment Dataset for South-South Nigeria
Description
This dataset was obtained from a field survey conducted on proper households (those with at least one parent and one child) in the South-South Geopolitical Zone of Nigeria. The dataset, collected between July 2023 and April 2024 using a purposeful sampling method, includes 543 validated responses captured in real-time through an online, electronic survey (e-survey) instrument developed with Google Forms. The survey instrument was synthesised from the United Nations Educational, Scientific, and Cultural Organisation (UNESCO) 2003 Language Vitality and Endangerment (LVE) framework/questionnaire to capture personalised views from households (five per Local Government Area (LGA)) within the language communities. The synthesised instrument makes the dataset suitable for identifying the causal LVE factors, group(s) or agent(s), thereby supporting efficient knowledge extraction and localisation. Also included are data on the speech systems of the languages spoken in these communities–consisting of recorded speech of 108-item Swadesh wordlist, with textual documentation of the gloss, syllable, and tone patterns of each word. The dataset is useful for mining insights into LVE patterns, providing a basis for understanding linguistic trends, changes, interactions over time, and other factors impacting language sustainability. It can also support the analysis of complex linguistic behaviours across language communities.
Files
Steps to reproduce
The LEE Framework synthesises the UNESCO’s 2003 LVE Framework indicators/questions, providing a systematic approach for assessing LVE with the goal of supporting language preservation efforts. The process begins by understanding the different sections of the UNESCO framework, which includes key indicators such as intergenerational language transmission, community attitudes, number of speakers, and the availability of educational materials. Drawing from this, a household-specific data collection instrument is developed to gather standardised linguistic and sociocultural data from speakers of the language. This instrument is then digitised using Google Survey Forms, enabling real-time data collection across diverse regions, such as the South-South geopolitical zone of Nigeria. The field data collection process involves purposive sampling, capturing responses from household-level units about demographic information, language use patterns, and community attitudes toward language vitality. Complementing this textual data, audio documentation is integrated by recording lexical items from the Swadesh wordlist to capture pronunciation, tone patterns, and phonological features essential for further analysis. All collected data, including survey responses and audio files, are curated into a centralised LVE Knowledge Base, a repository that stores structured datasets, annotated speech data, and metadata crucial for analysis. This knowledge base is then used to develop a LVE App, an AI-powered tool that allows users, including linguists, policymakers, and educators, to explore language vitality data interactively. The app generates visual dashboards, statistical summaries, and geo-linguistic maps that enable users to interpret language trends, speaker demographics, and endangerment levels. Using these insights, stakeholders such as language communities and policy-makers can make informed decisions about language revitalisation and preservation programmes. The framework’s final stage is public engagement and policy integration, where data products are disseminated to relevant stakeholders, promoting awareness and supporting decision-making for language revitalisation efforts. The present system combines traditional linguistic methods with modern digital tools such as Google Forms, audio documentation, and app development, creating a scalable and reusable model for assessing language endangerment. This approach enhances transparency, reproducibility, and community involvement in language documentation while offering a robust, data-driven model for preserving linguistic diversity in endangered regions. Through the integration of field-based data collection, digital processing, and interactive analytical tools, the framework aims to support the sustainable preservation of languages and empower local communities in their language revitalisation efforts, ultimately contributing to the broader goal of maintaining global linguistic diversity.
Institutions
- University of Uyo