BanglaCounter: A Benchmark Dataset for Bangla Counterspeech Generation
Description
BanglaCounter is a benchmark dataset for Bangla counterspeech generation and offensive-language research. The dataset contains 3,011 Bangla offensive text–counterspeech pairs. The offensive texts were collected from publicly accessible Facebook posts and YouTube comment sections and were paired with human-written Bangla counterspeech responses. The offensive texts were prepared for research use through a data-cleaning and quality-control process. Duplicate entries were removed, and usernames, profile URLs, profile mentions, and other information that could identify individual users were not included in the released dataset. Each record contains the original Bangla offensive text, a human written Bangla counterspeech response, English translations of both texts, three independent annotations for the counterspeech strategy, three independent annotations for the offensive content category, and final labels. The counterspeech responses are grouped into five strategies: empathy, fact-based, moral appeal, discouraging abuse, and warning of consequences. The offensive texts are grouped into seven categories: abusive language, personal insult, family insult, political abuse, religious hate, gender abuse, and threat. Three annotators independently assigned labels for both the counterspeech strategy and the offensive-content category. When at least two annotators selected the same label, the majority label was used. When all three annotators selected different labels, the sample was reviewed according to the annotation criteria and a final label was assigned. The released dataset retains both the individual annotations and the final adjudicated labels. Inter annotator agreement was measured using Fleiss kappa. The agreement score was 0.601 for counterspeech strategy and 0.616 for offensive content category. The dataset also went through several quality control steps, including duplicate removal, formatting and white-space normalization, annotation consistency checks, and review of the English translations. The final release contains 3,011 offensive texts and counterspeech responses, with no missing values. BanglaCounter can support research on Bangla counterspeech generation, offensive-language and hate-speech analysis, counterspeech-strategy classification, offensive-content classification, strategy-aware text generation, multilingual NLP, and evaluation of language models. The current version focuses on Bangla-script text. Romanized Bangla and Bangla–English code-mixed text are outside the scope of this release.
Files
Institutions
- Bangladesh University of ProfessionalsDhaka Division, Dhaka
- Jahangirnagar UniversityDhaka Division, Dhaka