BenCyber: A Multi-Aspect Bengali Cyberbullying Corpus with Demographics and Social Engagement Metrics

Published: 8 July 2026| Version 1 | DOI: 10.17632/w727tnt32v.1
Contributors:
,

Description

This dataset features 44,001 meticulously annotated Bangla social media comment records designed for granular cyberbullying, hate speech, and online harassment detection. The corpus captures contextual linguistic features by tracking five distinct target profiles (Actors, Social Figures, Singers, Politicians, and Sports Personalities) along with their gender identification. Each comment record includes metadata on social engagement through public reaction numbers and is manually categorized into one of five operational classes: Neutral, Harassment, Sexual Aggression, Hate Speech, and Violent Extremism. This structured contextual layout is highly optimized for developing explainable machine learning architectures, deep learning classifiers (such as Banglish/Bangla BERT, RoBERTa, and BiLSTM), and multi-task learning frameworks aimed at protecting vulnerable digital demographics in low-resource language environments.

Files

Steps to reproduce

To replicate or utilize this Bangla cyberbullying research dataset, follow these sequential data engineering and quality control steps: 1. Data Sourcing and Meta-Variable Harvesting: - Extract raw textual comment threads written in Bangla from diverse social media pages and public profiles. - For every extracted comment, strictly record auxiliary context: the target entity's profession ('Category'), the target's 'Gender', and the social validation weight ('Comment React Number'). 2. Human Annotation and Label Alignment: - Establish a comprehensive cross-validation pipeline with native Bangla annotators. - Categorize the text rows based on explicit linguistic intent into 5 target classes: 'Neutral' (safe comments), 'Harassment' (targeted insults/trolling), 'Sexual Aggression' (explicit or non-consensual sexualized context), 'Hate Speech' (attacks on race, religion, or community), and 'Violent Extremism' (radicalization or physical threats). 3. Data Hygiene and Structural Normalization: - Clean structural features by rectifying typographical case irregularities in categorical features (e.g., merging minor string mismatches in gender formatting like 'male' and 'Male' into a single uniform string). - Drop rows containing corrupt text or complete null values using data-handling toolkits (e.g., Python Pandas). 4. Preservation of Text Register: - Retain the raw textual records in their native script without aggressive stemming or pruning to keep punctuation markers, colloquial Bangla dialects, and structural emotional markers completely intact for transformer-based tokenizers (e.g., Bangla-BERT). 5. Final Integration & Export: - Format the consolidated dataframe to match a precise structural configuration of 5 features ('Text', 'Category', 'Gender', 'Comment React Number', 'Label') across exactly 44,001 unique rows. - Export the verified asset block as a comma-separated artifact titled "Cyberbulling Bangla Dataset.csv".

Institutions

Categories

Linguistics, Computer Science, Artificial Intelligence, Cybersecurity, Computational Linguistics, Social Media, Mental Health, Natural Language Processing, Statistical Natural Language Processing, Toxicity, Speech Analysis, Classification System, Bengali Language, Absorption-Distribution-Metabolism-Excretion Toxicity, Behavioral Toxicity, Bullying, Sexual Harassment, Deep Learning, Harassment, Social Media Analytics, Sentiment Analysis

Licence