Mabast Dataset

Published: 18 June 2026| Version 2 | DOI: 10.17632/msx42n3kfm.2
Contributors:
,

Description

Mabast is an annotated semantic text classification dataset developed for the Sorani Kurdish dialect. The dataset aims to support Kurdish Natural Language Processing (NLP), especially text classification and harmful-content detection in a low-resource language. The dataset contains 19,922 text samples in CSV format with two columns: text and label. Each sample is assigned to one of five categories: Normal (0), Romantic (1), Advice (2), Threat (3), and Hate Speech (4). The label distribution is 6,000 Normal, 4,513 Romantic, 5,079 Advice, 2,093 Threat, and 2,237 Hate Speech samples. The data was prepared with the contribution of volunteer participants and refined through cleaning, normalization, spelling review, duplicate removal, missing-value checking, and label consistency verification. The final dataset contains no missing values and no duplicate text entries. Mabast can be used for training and evaluating machine learning and deep learning models for Sorani Kurdish semantic text classification, hate speech detection, threat detection, and other low-resource Kurdish NLP tasks.

Files

Institutions

Categories

Artificial Intelligence, Natural Language Processing

Licence