Kurdish Social Media Opinions

Published: 19 February 2026| Version 3 | DOI: 10.17632/495h8779p6.3
Contributors:
Fanar Rofoo,
,

Description

The corpus was compiled through a systematic data collection process from authentic digital sources, including Kurdish-language social media, public forums, and news commentary sections. This strategy ensures the dataset reflects the dynamic and colloquial nature of the language as used by native speakers in organic, real-world settings. To ensure thematic relevance and linguistic diversity, the data collection was stratified across a range of topics of significant interest to Kurdish-speaking communities, such as politics, culture, social affairs, sports, and local current events. This domain-specific approach enhances the practical applicability of the resulting model. A rigorous multi-stage data preprocessing pipeline was implemented to ensure corpus integrity. Initially, raw user comments were subjected to a filtering process where non-Kurdish text, unintelligible statements, and irrelevant entries were removed. Subsequently, the remaining texts underwent a normalisation procedure to address orthographic variations and common informal writing conventions typical of computer-mediated communication. The core of the annotation scheme involved the manual classification of each textual unit into discrete sentiment categories: sadness, happiness, anger, disgust, fear, surprise and sarcastic. This categorisation was performed by native Kurdish speakers, who were trained to interpret linguistic cues, contextual nuances, and cultural subtleties. The focus of the annotation extended beyond simple lexical polarity (e.g., the presence of positive words) to encompass a more holistic assessment of the author's intent and overall opinion, thereby adding a layer of pragmatic understanding to the dataset. The resulting annotated corpus serves as a critical resource for advancing Kurdish language technology. It provides a reliable ground-truth dataset for training and evaluating machine learning and deep learning models tailored for sentiment analysis. This work establishes a foundational benchmark for future research in Kurdish NLP and contributes to the broader effort of developing inclusive language technologies for under-resourced languages.

Files

Steps to reproduce

Data sourcing was initiated by identifying prominent Kurdish social media pages known for high user engagement. Within these pages, selection criteria were applied to individual posts to ensure a rich source of subjective user opinions. Primary selection was based on posts that explicitly solicited public opinion, as these are inherently designed to motivate users to express personal viewpoints and feelings in the comments section. A secondary criterion was the total volume of comments on a post; a higher comment count was used as an indicator of a more active discussion and was presumed to increase the likelihood of capturing a broad spectrum of expressed sentiments. The data extraction process was facilitated by EXPORTCOMMENTS, a specialised web-based platform that provides automated tools for harvesting user comments from various social media APIs. During the initial data collection phase, observational analysis indicated that Facebook consistently yielded a higher density of comments containing explicit emotional expression compared to other social media platforms examined. This disparity in sentiment-rich content led to the decision to focus the collection efforts exclusively on public posts from Kurdish Facebook pages to maximise the efficiency and density of relevant data acquisition.

Institutions

  • Salahaddin University - Erbil College of Science
    Kurdistan, Erbil

Categories

Social Media, Sentiment Analysis

Licence