BanglaCounter: A Benchmark Dataset for Bangla Counterspeech Generation

Published: 5 August 2026| Version 2 | DOI: 10.17632/6ysbv3mg3s.2
Contributors:
,
,
,
,

Description

This dataset provides a collection of Bangla offensive text–counterspeech pairs compiled from publicly available Facebook and YouTube comments. All offensive texts were obtained from publicly accessible sources, and personally identifiable information, including usernames and personal names, was removed or anonymized during data preparation. In total, the dataset contains 3,011 offensive text–counterspeech pairs covering eleven offensive content categories, including abusive language, hate speech, threat, gender abuse, sexual harassment and others. Unlike existing Bangla offensive language datasets, This dataset contains offensive Bangla texts along with their corresponding counterspeech responses. Each offensive text is associated with its offensive category, a counterspeech response, one of five counterspeech strategies (Empathy, Fact-based Response, Moral Appeal, Discouraging Abuse, and Warning of Consequences), and English translations of both the offensive text and the counterspeech response. The counterspeech responses were initially generated using ChatGPT and subsequently reviewed and corrected where necessary before inclusion in the final dataset. Regarding the counterspeech strategy distribution, the dataset contains responses generated under five different counterspeech strategies to encourage constructive engagement with offensive online content. For example, an offensive text states: "এই বাটপারটার জন্যই জণগন শান্তিতে নাই।" ("It's because of this swindler that the people have no peace."). A corresponding counterspeech response is: "গালি না দিয়ে সমস্যার সমাধান নিয়ে আলোচনা করা উচিত।" ("Instead of hurling abuse, we should discuss how to solve the problem.").

Files

Institutions

Categories

Natural Language Processing, Machine Learning, Supervised Learning, Social Media Analytics

Licence