Public Sentiment Dataset of Bangladeshi People Towards Government in Social Media
Description
This dataset contains 6,000 manually annotated Bangla Facebook comments representing public sentiment toward the Interim Government of Bangladesh, collected between August 2024 and 2025. The collection period falls immediately following the July 2024 political uprising that led to a major governmental transition in the country. The central research hypothesis is that Bangla political sentiment expressed on social media can be reliably classified into three categories (Positive, Negative, and Neutral), and that language-specific transformer models can capture the nuanced, context-dependent nature of such discourse more effectively than traditional machine learning approaches. The dataset is perfectly balanced, with exactly 2,000 comments per sentiment class. Comments were collected from major public Facebook pages including Chief Advisor GOB, Jamuna Television, Somoy News, Daily Prothom Alo, Kalbela Digital, and Dhaka Tribune, ensuring diversity across both news outlets and political discussion spaces. "Positive" comments tend to express approval, trust, or satisfaction toward government actions; "Negative" comments reflect criticism, dissatisfaction, or distrust; and "Neutral" comments include factual statements, questions, or ambiguous opinions without clear polarity. Comments were collected using a hybrid approach combining manual selection, browser extension-based extraction, and automated scraping via Apify. Only publicly accessible posts and comments were included; private profiles, closed groups, and personally identifiable information were excluded. Three annotators with Bengali language proficiency and social media political awareness labeled each comment independently into one of three sentiment classes (Positive, Negative, Neutral) following a detailed annotation guideline. An expert reviewer with 15 years of experience resolved all disagreements and validated borderline cases. Inter-annotator reliability was measured using Fleiss' Kappa, yielding a score of κ = 0.91, indicating almost perfect agreement and confirming the reliability of the annotations. The dataset is provided in two files. "raw_data_bangla_public_sentiment_toward_government.xlsx" contains the original annotated comments with two columns: Sentence (raw Bangla text) and Sentiment (label: Positive, Negative, or Neutral). "preprocessed_data_bangla_public_sentiment_toward_government.xlsx" contains three columns: Sentence (original text), Sentiment (label), and Clean Sentence (preprocessed text after Unicode normalization, removal of HTML tags, URLs, mentions, hashtags, digits, emojis, non-Bangla characters, and extra whitespace). Researchers may use the raw file for custom preprocessing pipelines or the preprocessed file directly for feature extraction and model training. The dataset is suitable for Bangla sentiment analysis, NLP benchmarking, political discourse analysis, and training or fine-tuning transformer-based models for low-resource Bangla text classification.
Files
Institutions
- Daffodil International UniversityDhaka Division, Dhaka