BOISHOMMO: Multi Label Hate Speech Dataset

Published: 17 December 2025| Version 2 | DOI: 10.17632/v9jws6zyxz.2
Contributor:
Md Abdullah Al Kafi

Description

The dataset addresses the lack of comprehensive hate speech (HS) datasets in low-resource languages like Bangla. It introduces BOISHOMMO, a multi-label Bangla HS dataset with over 2,000 annotated examples across categories like race, gender, religion, and politics. The dataset captures the complex, multi-dimensional nature of HS and supports better detection in Bangla through algorithmic evaluation and analysis of language-specific challenges.

Files

Institutions

  • Daffodil International University

Categories

Computational Linguistics, Natural Language Processing

Licence