A Benchmark Dataset for Human Vs. AI-Rewritten Bengali Text Detection
Description
This dataset provides a benchmark resource for AI-rewritten Bengali text detection. It contains 4,000 labeled Bengali text samples, comprising 2,000 original human-written texts and 2,000 corresponding AI-rewritten texts. The human-written texts were collected from authentic Bengali sources, including National Curriculum and Textbook Board (NCTB) textbooks, Bengali literary works available through Wikisource, and Bengali news articles published before 2020. Each original text was rewritten using one of four widely used large language models such as ChatGPT, Claude, Gemini, or DeepSeek, through a standardized prompt designed to preserve the original semantic meaning, writing style, and approximate text length. Each human-written text is paired with its corresponding AI-rewritten version through a unique Pair_ID, enabling direct comparison between original and rewritten texts. The dataset includes metadata such as the text source or rewriting model, content category, word count, and class label. Covering multiple domains, including Education, Bengali Literature, and News, this dataset is intended to support research on AI-rewritten text detection, stylometric analysis, authorship attribution, paraphrase identification, and other Bengali natural language processing tasks.
Files
Institutions
- Chittagong University of Engineering & TechnologyChittagong, Chittagong
- Jahangirnagar UniversityDhaka Division, Dhaka