BanglaHarassNet: A Multi-Platform Annotated Dataset for Online Harassment and Cyberbullying Detection in Bangladeshi Social Media
Description
This dataset contains over 5500 text-based instances collected from various social media platforms for research in natural language processing (NLP), text classification, and automated content moderation. The dataset includes content written in Bangla, English, and Banglish (Bangla written using the Roman alphabet), reflecting diverse online communication patterns and user interactions. The collected texts represent multiple forms of abusive and non-abusive language commonly observed in digital environments. The dataset is organized into five separate spreadsheet files: "Hate Speech.xlsx", "Cyberbullying.xlsx", "Sexual Harassment.xlsx", "Trolling.xlsx", and "Neutral.xlsx". Each entry has been manually reviewed, categorized, and structured with consistent annotations to maintain data quality and reliability. The dataset is designed to support the development, training, validation, and evaluation of machine learning and deep learning models for detecting harmful online content. Due to the nature of social media communication, the data may include informal language, spelling variations, code-switching, and contextual expressions. Categories for this data: Neutral: Non-abusive, safe, and ordinary conversational content representing regular online interactions without harmful intent (Severity: 0) found inside "Neutral.xlsx". Trolling: Content containing provocative, sarcastic, mocking, or disruptive comments intended to annoy, provoke, or derail discussions (Moderate Abuse Level; Severity: 2) found inside "Trolling.xlsx". Cyberbullying: Content involving personal attacks, harassment, humiliation, threats, or repeated harmful behavior targeting specific individuals (Moderate Abuse Level; Severity: 2) found inside "Cyberbullying.xlsx". dividuals or groups (Extreme Abuse Level; Severity: 4) found inside "Sexual Harassment.xlsx". Hate Speech: Content containing attacks, discrimination, insults, or hostile expressions directed toward groups, communities, religions, ethnicities, or political identities (Severe Abuse Level; Severity: 3) found inside "Hate Speech.xlsx".. Sexual Harassment: Content containing sexually explicit, suggestive, objectifying, or predatory language directed toward individuals or groups (Extreme Abuse Level; Severity: 4) found inside "Sexual Harassment.xlsx".
Files
Institutions
- Southeast UniversityDhaka Division, Dhaka