Bangla Human & Ai Text Classification

Published: 3 August 2026| Version 1 | DOI: 10.17632/tvtm9tkvdn.1
Contributor:

Description

The dataset consists of 6,000 Bangla text samples designed for binary classification of human-written and AI-generated content. It is evenly balanced, containing 3,000 human-written texts and 3,000 AI-generated texts, ensuring an unbiased distribution for training and evaluating machine learning and deep learning models. Each record in the dataset contains three attributes. The source_text field stores the Bangla text that serves as the primary input for classification. The Source column indicates the class label, identifying whether the text is Human or AI. The Source Name column specifies the origin of the text. For human-written samples, this field records the original content source, while for AI-generated samples, it identifies the large language model (LLM) used to generate the text, such as Gemini, GPT, Qwen, or Grok. The dataset covers a diverse range of writing styles and topics to improve model generalization and reduce source-specific bias. It is suitable for tasks including AI-generated text detection, Bangla natural language processing (NLP), stylometric analysis, explainable AI (XAI), and comparative evaluation of traditional machine learning, deep learning, and transformer-based language models. The balanced class distribution and inclusion of source metadata make the dataset valuable for benchmarking text classification algorithms and analyzing linguistic differences between human-written and AI-generated Bangla content.

Files

Institutions

Categories

Natural Language Processing, Binary Classification

Licence