A Balanced Six-Region Dataset for Bangla Dialect Identification

Published: 5 August 2026| Version 3 | DOI: 10.17632/4gx4y95mn3.3
Contributors:
,
,
,
,

Description

The dataset is provided as a single tabular file (.xlsx/.csv) containing four columns: Standard English, Standard Bangla, Regional Dialect, and Label. The Standard English and Standard Bangla columns contain the original elicitation prompt sentence used to collect the corresponding regional-dialect translation. The Regional Dialect column contains the sentence as rendered by a native speaker in one of the six target regional dialects. The Label column contains the corresponding region name (Rangpur, Bogura, Sylhet, Chittagong, Pabna, or Cumilla), indicating the dialect class of that sentence. The dataset contains exactly 3,600 sentences in total, with 600 sentences per region. The dataset is provided in randomly shuffled order to avoid any ordering bias. All text is UTF-8 encoded to preserve the Bangla script and dialect-specific spelling conventions.

Files

Steps to reproduce

1. Prepare a fixed list of Standard Bangla sentences covering everyday conversational topics, along with their English translations. 2. Create separate Google Sheets for each target region, with the Standard Bangla and English prompt sentences placed in fixed columns. 3. Recruit at least three native-speaking volunteer contributors per region and share the corresponding sheet with them. 4. Ask each contributor to translate every prompt sentence into their own regional dialect, based on their natural, everyday usage. 5. Recruit one additional native speaker per region to validate a random sample (50-100 sentences) of the collected translations for dialectal naturalness. Revise any flagged sentences in consultation with the original contributor. 6. Consolidate all region-wise sheets into a single master file containing four columns: Standard English, Standard Bangla, Regional Dialect, and Region Label. 7. Remove duplicate entries and rows with missing values. Balance the dataset to an equal number of sentences per region (600 in this study). 8. Randomly shuffle the final dataset to avoid ordering bias before distribution.

Institutions

Categories

Linguistics, Computer Science, Artificial Intelligence, Natural Language Processing

Licence