BanglaFakeNews2025: A High-Quality Benchmark Dataset for Bangla Fake News Detection

Published: 5 August 2026| Version 3 | DOI: 10.17632/c6hf7g5f3t.3
Contributors:
,
,
,

Description

This dataset presents a balanced, manually annotated collection of 4,000 Bangla news articles for fake news detection research, comprising 2,000 Real and 2,000 Fake news samples published between 2021 and 2025. Fake news articles were collected from established fact-checking organizations operating in Bangladesh, including Rumor Scanner, DhakaFactCheck, FactWatch, Jachai, Boom BD, Aaj Tak Bangla Fact Check, and eArki, also from the social media posts whose published verification reports served as reliable sources of labeled fake content. Real articles were collected from credible mainstream Bangla news outlets, including Prothom Alo, BBC News Bangla, Jamuna TV, Samakal, News24, Dhaka Tribune, Somoy TV, Independent TV, Ittefaq, and Bangladesh Pratidin. All data were obtained from publicly available sources for academic research purposes. Each article is stored with ten structured fields: Article ID (unique identifier), Domain (publisher website), Date (publication date), Category (one of twelve topical categories: Politics, National, International, Sports, Entertainment, Crime, Education, Technology, Finance, Lifestyle, Editorial and Miscellaneous), Headline, Content (full article body), Label (1 = Real, 0 = Fake), Source (verifying source for the claim), Relation (Related if the headline accurately reflects the content's claim, otherwise Unrelated), and F-type (fine-grained fake news type: Fabricated, Clickbait, Satire, or Altered; Not Applicable for real articles). Every article was manually reviewed and labeled, and the Relation field additionally captures a common form of soft misinformation in which a factual article body is paired with a misleading headline. The dataset is provided in four files: two Bangla files containing the original real and fake news articles, and two English files in which the Headline and Content fields were translated from Bangla to English using Google Translate, enabling multilingual misinformation research. Unlike earlier Bangla fake news resources, this dataset is perfectly class-balanced, covers contemporary news from 2021–2025 and includes rich metadata supporting binary fake news detection, fake news type classification, headline–content consistency analysis, topical category prediction, and temporal generalization studies. It is intended as a benchmark resource for researchers working on misinformation detection in Bangla.

Files

Institutions

Categories

Natural Language Processing, Supervised Learning, Bengali Language

Licence