Text Detoxification in isiXhosa and Yorùbá Dataset
Published: 31 October 2025| Version 1 | DOI: 10.17632/jz8mpwdmgr.1
Contributor:
Abayomi AgbeyangiDescription
A parallel dataset of toxic and detoxified sentence pairs (178 pairs each) was manually generated for isiXhosa and Yorùbá, which included a wide range of linguistic and communicative forms such as direct insults, implicit hostility, sarcasm, emotional outbursts, and culturally specific slurs.
Files
Steps to reproduce
The data for isiXhosa and Yorùbá were manually compiled. To ensure linguistic authenticity, special care was taken to preserve diacritic usage and orthographic idiosyncrasies in Yorùbá, as well as agglutinative word patterns in isiXhosa.
Institutions
- Walter Sisulu University - Buffalo City Campus
Categories
Computational Linguistics, Natural Language Processing, Machine Learning, Subsaharan Africa, South Africa, Nigeria