Bangla Regional Dialects Speech Dataset
Description
The Bangla Regional Dialects Speech Dataset is a curated speech corpus containing regional Bangla dialect audio recordings collected from multiple areas of Bangladesh. The dataset is designed to support research and development in speech recognition, natural language processing (NLP), dialect classification, speaker analysis, and other AI-based language technologies. The dataset contains speech recordings from 5 different regions of Bangladesh, representing diverse regional accents and pronunciations of the Bangla language. A total of 19 speech categories are included in the dataset. Each region contains 1330 unique Bangla sentences, resulting in a total of 6650 recorded speech samples across the complete dataset. All audio files are provided in: .wav audio format Mono channel 16 kHz sampling rate The dataset also includes transcription metadata files for efficient training and evaluation of machine learning and deep learning models. This dataset can be used for: Automatic Speech Recognition (ASR) Bangla dialect identification Speech-based AI systems NLP research Linguistic analysis Audio classification Deep learning research The dataset aims to contribute to the advancement of Bangla language technology and regional dialect research by providing a structured and high-quality regional speech resource for researchers and developers. Keywords: Bangla speech, Bengali speech recognition, ASR, speech dataset, dialect recognition, Bangla NLP, audio dataset, 16kHz speech, mono audio, regional dialects, machine learning dataset Researchers who wish to access the complete dataset (including the .wav audio files and full metadata) may: 1. Complete the Access Request Form: https://docs.google.com/forms/d/e/1FAIpQLSd68-BZtHXPwb1RB5KkN5fYUPZIIXDPLCEbPFZ7y-EsDNmEOA/viewform 2. Or contact the corresponding author directly via email: M. Sakib Rahman Email: 22201240@uap-bd.edu with your name, institution, and a brief description of your research goal. Requests will be reviewed, and access will be provided for eligible academic and non-commercial research.
Files
Steps to reproduce
1. Recruit native Bangla speakers from the five target dialect regions (Barishal, Chattogram, Noakhali, Rangpur, and Sylhet) and obtain informed consent. 2. Record the predefined scripted Bangla sentences using smartphones or external microphones in quiet or moderately noisy environments. 3. Convert all recordings to a standardized format (16 kHz, 16-bit, mono, .WAV). 4. Apply preprocessing steps including noise reduction, amplitude normalization, silence trimming, and sentence segmentation. 5. Manually verify each recording and its corresponding Standard Bangla transcription for quality and accuracy. 6. Organize the recordings into regional folders and generate the accompanying metadata file containing anonymized speaker information, transcripts, and recording attributes. 7. Use the provided metadata and documentation to train and evaluate Automatic Speech Recognition (ASR), dialect classification, or other speech processing models.
Institutions
- University of Asia PacificDhaka Division, Dhaka