Arabic Paraphrasing Corpus for NLP Applications

Published: 13 October 2025| Version 1 | DOI: 10.17632/x8wj68wrfj.1
Contributors:
, Abeer Shdefat, Nailah Al-Madi, Bassam Hammo

Description

The Arabic Paraphrasing dataset consists of 143 short-text records collected from journal articles. It is designed to support research in Arabic paraphrasing and text similarity detection. The dataset includes 100 paraphrased texts and 43 non-paraphrased texts, providing a balanced set of examples for evaluating paraphrase identification and generation models.

Files

Institutions

  • Princess Sumaya University for Technology

Categories

Arabic Language

Licence