Requirement Classification: Hierarchical Multi-Label Transformer Pipeline

Published: 5 June 2026| Version 1 | DOI: 10.17632/2k7fkxxwbt.1
Contributor:
Nam Nguyen

Description

This dataset accompanies the article “Evaluating Robustness of Hierarchical and Multi-Label Requirements Classification: Effects of Data Partitioning and Supervision Design”, published in Information and Software Technology (2026), DOI: 10.1016/j.infsof.2026.108212. The dataset contains 14,739 unique software requirement statements collected and consolidated from multiple public and industry-oriented sources, including MSHD-IAU, FR_NFR_Mendeley, PROMISE_exp, and CCHIT. Duplicate records and unclassifiable entries were removed during data preparation. Each requirement statement was manually annotated according to a hybrid ISO-based taxonomy derived from ISO/IEC/IEEE 29148:2018 and ISO/IEC 25010. The taxonomy consists of three hierarchical levels covering both Functional Requirements (FRs) and Non-Functional Requirements (NFRs), including 2 top-level categories, 13 intermediate categories, and 44 leaf-node classes. Multi-label annotation is supported, and approximately 79% of the requirements contain multiple labels. The dataset includes requirement texts, unique identifiers, hierarchical labels, preprocessing outputs, and metadata used for machine learning experiments. The preprocessing pipeline includes text normalization, acronym expansion, lemmatization, stop-word filtering, terminology standardization, and word-count analysis. This resource was developed to support research in requirements engineering, natural language processing, hierarchical text classification, multi-label learning, software quality analysis, and transformer-based machine learning models. The dataset was used to evaluate hierarchical classification architectures based on Sentence-BERT embeddings and staged versus joint supervision strategies across multiple taxonomy levels. Researchers may use this dataset for benchmarking, reproducibility studies, taxonomy evaluation, class imbalance research, data augmentation experiments, and automated software requirement classification tasks.

Files

Categories

Requirement Engineering, Binary Classification

Licence