BANTEN: A Parallel Banglish–English Dataset for Machine Translation
Published: 6 August 2026| Version 2 | DOI: 10.17632/j3y2y9nfck.2
Contributor:
Description
BANTEN is a manually curated parallel Banglish–English dataset comprising 14,000 sentence pairs collected from publicly available online sources, including newspapers, Facebook posts and comments, YouTube comments, daily conversations, and blogs. Each instance contains a Banglish sentence written in Roman script and its corresponding human-translated English sentence. The dataset was developed through data collection, filtering, cleaning, manual translation, and expert validation. It is intended to support research in Banglish-to-English machine translation, code-mixed language processing, transliteration, and low-resource natural language processing.
Files
Categories
Computer Science, Natural Language Processing, Machine Translation