QSU: Quran Semantic Units Dataset

Published: 22 June 2026| Version 1 | DOI: 10.17632/v7hyhk7krd.1
Contributors:
,
,
,

Description

The Quran Semantic Units (QSU) Dataset, a novel segmentation of the Quranic text that achieves fine-grained granularity by applying the classical rules of Waqf wa Ibtida' (stopping and resuming), which is considered an essential tool for understanding and correctly interpreting the Quranic text. Unlike standard verse-by-verse divisions, which often contain multiple ideas within a single ayah, our dataset partitions the text into self-contained, semantically complete units. Built to advance NLP applications such as semantic search, question-answering, contextual embeddings, and automated tafsir, our methodology synthesizes established stopping conventions with rigorous linguistic analysis to accurately delineate semantic boundaries.

Files

Steps to reproduce

The raw Quranic text in QSU was obtained from the Tanzil Project, a verified digital resource that underwent automatic, rule-based, and manual verification against the Medina Mushaf. Based on the 1924 Cairo edition endorsed by Al-Azhar University, the text standardizes the Hafs 'an 'Asim reading. The script used is Imla'ei (modern orthographic) . The verse-end Waqf marks were sourced exclusively from the printed Taj Mushaf, produced by Taj Company (South Asia) . This Mushaf follows the standard Uthmani rasm and adheres to the South Asian (Indo-Pak) script tradition. Only the verse-end boundaries and their associated Waqf marks were extracted from this source, while all intra-verse semantic boundaries were derived from the Madina Mushaf. The segmentation of units is achieved by precisely identifying the points of pause and continuation. The dataset was constructed through a multi-stage experimental pipeline involving data acquisition, segmentation, validation, and quality assurance. The QSU dataset provides (11,242 units), while QAU dataset contains a number of units that corresponds to the number of verses in the Quran in the Medina Mushaf (6,236 units). The WBS.csv file provides word-level boundary state annotation for all (77,800) words of the Quran. Each row contains a word's location, its Arabic text, and two boundary state attributes (unit_begins and unit_ends) that indicate whether a semantic unit starts or ends at that word. All data files are provided in CSV format with UTF-8 encoding .

Institutions

Categories

Artificial Intelligence, Computational Linguistics, Natural Language Processing, Semantic Processing

Funders

Licence