A speech dataset of three ethnic languages of Bangladesh: Chakma, Garo, and Marma.

Published: 19 November 2025| Version 1 | DOI: 10.17632/yjhybztwf4.1
Contributors:
,
,

Description

This is a managed speech corpus of three ethnic languages of Bangladesh, specifically, Chakma, Marma, and Garo. It comprises 2321 short audio recordings (between 1 and 4 seconds long) of 11 native speakers (20-26 years old) who were reading 211 predefined sentences in the Bengali language. Smartphones were used to record in various acoustic environments and standardized with further metadata information such as the speaker ID, age, gender, ethnicity, recording environment, recording device and file length. The data can be used in automatic speech recognition, speaker/language recognition, acoustic modeling as well as other low-resource speech processing problems.

Files

Institutions

  • Daffodil International University

Categories

Ethnic Group, Natural Language Processing, Informatics, Underdevelopment, Language

Licence