ILC-BD: Indigenous Languages Corpus of Bangladesh (Chakma-English and Marma-English parallel sentences)
Description
ILC-BD contains 10,000 short everyday English sentences with their translations into Chakma and Marma, two indigenous languages of Bangladesh (10,000 Chakma-English and 10,000 Marma-English sentence pairs). The English sentences come from the Tatoeba Project. Five native speakers of each language translated them by hand, writing Chakma and Marma in English (Latin) letters. Five typists digitized the translations with cross-checking, and two independent language experts per language reviewed every pair. The data are provided as an Excel file and a CSV file with the columns ID, English, Chakma and Marma. The dataset can be used for machine translation, transfer learning, and linguistic studies of Chakma and Marma. English sentences: Tatoeba Project (https://tatoeba.org), licensed CC BY 2.0 FR. An earlier version of this corpus is listed on IEEE DataPort (https://doi.org/10.21227/e6jn-fz88).
Files
Institutions
- Jahangirnagar UniversityDhaka Division, Dhaka