BanglaPlagPair-1K: A controlled Bangla text-pair dataset for plagiarism-oriented transformations and semantic non-equivalence
Description
BanglaPlagPair-1K is a controlled Bangla text-pair dataset created for research on plagiarism detection, paraphrase recognition, semantic similarity, semantic-equivalence classification, passage retrieval, and text-pair analysis. It contains 1,000 ordered Bangla text pairs, with an equal balance of 500 meaning-preserving pairs and 500 semantically non-equivalent Mismatch pairs. The meaning-preserving pairs are divided equally into five transformation categories: Morphological, Order, Synonymy, Lexical, and Mixed, with 100 pairs in each category. The first 600 records form 100 six-record construction blocks. In each block, the same reference passage is paired with one example from each of the five transformation categories and one Shared-reference Mismatch example. This structure allows controlled comparison of different rewriting mechanisms while keeping the reference passage fixed. The remaining 400 Mismatch pairs include 200 Same-source pairs, where both passages come from the same source page, and 200 Cross-source pairs, where the passages come from different source pages. These negative examples provide different levels of surface and lexical overlap. Reference passages were collected from publicly accessible Bangla Wikipedia articles. Selected classical Bengali literary texts were also used for some Cross-source Mismatch examples. Provenance URLs are recorded for all externally sourced passages so that the source of each passage can be traced. The dataset was manually constructed, annotated, and reviewed through multiple quality-control stages. A stratified subset of 100 pairs was independently evaluated by three non-author annotators proficient in Bangla using a blinded binary semantic-equivalence task. Their agreement with the released binary labels was 100%, 99%, and 98%, and Fleiss' kappa across the three annotators was 0.960. Programmatic checks were also performed to verify dataset structure, category and label distributions, construction blocks, required fields, Unicode normalization, whitespace consistency, source-link coverage, and the absence of duplicate pairs, repeated text2 values, and self-pairs. The released dataset contains seven fields: pair_id, text1, pair_type, binary_label, text2, source_link_text1, and source_link_text2. BanglaPlagPair-1K is intended as a controlled research resource rather than a collection of verified real-world plagiarism cases. It can support benchmarking, threshold analysis, transformation-specific evaluation, qualitative error analysis, and the development of Bangla plagiarism-detection and paraphrase-recognition systems.
Files
Institutions
- Jahangirnagar UniversityDhaka Division, Dhaka