Bangla Text Paraphrase Corpus for Natural Language Processing
Published: 12 August 2025| Version 1 | DOI: 10.17632/ffkm5rk2yg.1
Contributors:
, , Description
This dataset is Bangla Paraphrase Sentence Pair Dataset (BPDS), contains pairs of Bangla sentences labeled as paraphrase (same meaning) or non-paraphrase (different meaning). The data has been collected from diverse Bangla sources including books, newspapers and literature articles, covering a wide range of topics and writing styles. It is designed for research in natural language processing tasks such as paraphrase detection, semantic textual similarity, text generation and plagiarism detection in Bangla. The dataset is provided in .xlsx format with three columns: Sentence1, Sentence2.
Files
Institutions
- Daffodil International University
Categories
Linguistics, Computational Linguistics, Semantics, Natural Language Processing, Text Mining