RIYE Textual and Geospatial Corpus: An Aggregated Multidialectal Dataset for Low-Resource Language Processing

Published: 10 August 2026| Version 1 | DOI: 10.17632/v9gdz3k7zg.1
Contributors:
,
,
,

Description

This dataset integrates textual and spatial metadata collected under the Digiculture RIYE Project framework to support research in computational linguistics, spatial humanities, and low-resource dialect analysis. The corpus comprises three complementary datasets: RIYE_CSC-HTM_Aggregate.csv, which aggregates raw, qualitative text entries across cultural categories (such as food, festivals, music, history, and folklore) collected by the Computer Science (CSC) and Hospitality and Tourism Management (HTM) groups from the Federal University of Agriculture, Abeokuta; RIYE_CSC_Dataset.csv, which provides a structured matrix of thematic cultural elements (including dialect, clothing, religion, leadership, and performance arts) mapped across Local Government Areas (LGAs), towns/wards, and GIS location coordinates; and ogun_state_lgas_wards_coordinates.csv containing constitutionally recognised LGAs, wards, and coordinates extracted from the publicly available database compiled by Independent National Electoral Commission (INEC) before the 2023 General Elections. By combining unstructured descriptive text with spatially anchored domain attributes across regional zones, this collection establishes a rich baseline for requirement analysis, computational text processing, and geospatial modelling of dialectal and cultural elements.

Files

Steps to reproduce

To reproduce the results or utilise this dataset for text and geospatial analysis, the process begins with the systematic extraction and organisation of the CSV files. Once the dataset archive is downloaded, RIYE_CSC-HTM_Aggregate.csv, RIYE_CSC_Dataset.csv and ogun_state_lgas_wards_coordinates.csv should be placed into a dedicated local working directory. The architecture of the dataset is intentionally designed using a relational tabular paradigm, where raw text records collected by the Computer Science (CSC) and Hospitality and Tourism Management (HTM) groups are provided in RIYE_CSC-HTM_Aggregate.csv, while structured category extractions and geographic markers are catalogued in RIYE_CSC_Dataset.csv and are being complemented with official location-based data in ogun_state_lgas_wards_coordinates.csv for further insights. This structure simplifies the preprocessing pipeline, as it allows researchers to programmatically parse text entries, thematic categories, Local Government Areas (LGAs), and GIS locations directly from tabular files without needing a complex external database. The next phase involves automating the data ingestion through a script that reads these files. A standard script should load RIYE_CSC-HTM_Aggregate.csv to process the aggregated plain-text descriptions, while simultaneously reading RIYE_CSC_Dataset.csv to map extracted cultural categories (such as food, clothing, festivals, and dialect) alongside their corresponding LGA and GIS coordinates. As the script iterates through each entry, it pairs the extracted linguistic and cultural elements with their respective geographic parameters. This approach is highly compatible with modern data processing frameworks, where data loaders can be configured to parse and map attributes automatically, ensuring that the integrity of the spatial and category mapping remains consistent across different analytical environments. Once the files are loaded, the data should be standardised to ensure uniform input for analysis. This typically involves parsing text fields for consistent formatting and validating coordinate data for spatial accuracy. Because the geographic locations and category fields are explicitly organised within RIYE_CSC_Dataset.csv, it is easy to perform regional queries or subset the data across specific LGAs and cultural tags. This straightforward workflow ensures that the transition from raw tabular aggregations to a structured, analysis-ready dataset is both efficient and error-free.

Institutions

Categories

Linguistics, Computer Science, Artificial Intelligence, Requirement Engineering, Natural Language Processing, GIS Database, Spatial Analysis, Textual Analysis, Applied Machine Learning

Funders

  • Tertiary Education Trust Fund [TETFUND]
    Grant ID: 2023 TETFund/NRF/HSS/HST_0040

Licence