Bangla Text Paraphrase Corpus for Natural Language Processing

Published: 12 August 2025| Version 1 | DOI: 10.17632/ffkm5rk2yg.1
Contributors:
,
,

Description

This dataset is Bangla Paraphrase Sentence Pair Dataset (BPDS), contains pairs of Bangla sentences labeled as paraphrase (same meaning) or non-paraphrase (different meaning). The data has been collected from diverse Bangla sources including books, newspapers and literature articles, covering a wide range of topics and writing styles. It is designed for research in natural language processing tasks such as paraphrase detection, semantic textual similarity, text generation and plagiarism detection in Bangla. The dataset is provided in .xlsx format with three columns: Sentence1, Sentence2.

Files

Institutions

  • Daffodil International University

Categories

Linguistics, Computational Linguistics, Semantics, Natural Language Processing, Text Mining

Licence