BanglaAccent: Annotated Bangla Dialects with Demographic Metadata

Published: 13 November 2025| Version 3 | DOI: 10.17632/6fv9dbh5kr.3
Contributors:
Md Abdullah Al Kafi, Fuad Hasan, Minhazur Rahman Shuvo, Md. Dewan Mamun Raza

Description

Desher Vasha is a curated speech dataset capturing dialectal variation across more than 20 distinct regions of Bangladesh. Designed to support linguistic research, educational tools, and low-resource NLP development, the corpus includes voice recordings annotated with rich metadata such as speaker location, age, gender, and dialectal features. Use Cases • Linguistic research: Dialect classification, phoneme variation, sociolinguistic mapping • Educational tools: Role-based apps for dialect awareness and pronunciation training • NLP & ASR: Benchmarking for Bangla speech recognition, especially in low-resource and dialect-sensitive contexts • Cultural preservation: Documenting endangered or underrepresented dialects

Files

Institutions

  • Daffodil International University

Categories

Linguistics, Speech Processing, Computational Linguistics, Phonetics, Audio Recording, Sociolinguistics Variation, Bengali Language, Bangladesh, Low-Resource LLM

Licence