Madurese-Language Social Media Comment Dataset for Public Service Analysis

Published: 14 August 2026| Version 1 | DOI: 10.17632/yndvjm5bsv.1
Contributor:

Description

This dataset contains 1,510 public comments written primarily in Madurese, collected from social media platforms including TikTok, Facebook, Instagram, and X. The comments discuss various public services, government programs, infrastructure, and community issues in Madura, Indonesia. Each record includes the following information: Platform: The social media platform where the comment was posted. Kategori_Layanan: The public-service or issue category associated with the comment, such as roads and bridges, public lighting, clean water, healthcare, education, waste management, parking, flooding, public transportation, and other government-related issues. Teks Mentah: The original comment text, preserved in its collected form, including local spelling variations, informal language, abbreviations, punctuation, emojis, and code-switching. Code Mixing: A binary annotation indicating whether the comment contains code-mixing between Madurese and another language, primarily Indonesian. Sarkasme: A binary annotation indicating whether the comment contains sarcastic or ironic expression. The dataset is intended to support research in low-resource language processing, social media analysis, public-service evaluation, code-mixing detection, sarcasm identification, and the study of public opinion in Madurese-speaking communities. It may also be useful for developing and evaluating natural language processing models involving regional languages, informal user-generated content, and multilingual communication.

Files

Steps to reproduce

1. Identify publicly available social media posts and comment sections discussing public services, government programs, infrastructure, and community issues in Madura, Indonesia. 2. Manually browse relevant content on TikTok, Facebook, Instagram, and X using the platforms’ standard interfaces. No web crawlers, bots, automated scripts, or scraping tools were used. 3. Select comments that are publicly accessible and relevant to the research topics. Record the platform, original comment text, and the corresponding public-service or issue category. 4. Preserve the original form of each comment, including informal spelling, Madurese and Indonesian expressions, abbreviations, punctuation, emojis, and other user-generated language characteristics. 5. Manually annotate each comment for Code Mixing: - 1 if the comment contains a mixture of Madurese and another language, primarily Indonesian. - 0 if no code-mixing is identified. 6. Manually annotate each comment for Sarkasme: - 1 if the comment expresses sarcasm or irony in its context. - 0 if sarcasm is not identified. 7. Review the recorded entries and annotations for consistency, then organize the data into a semicolon-delimited CSV file containing the fields Platform, Kategori_Layanan, Teks Mentah, Code Mixing, and Sarkasme. Because social media content may change or become unavailable, exact replication of the original collection process may not be possible. The dataset is provided to support validation, reuse, and further research. All data collection was conducted manually from publicly accessible content, with attention to applicable ethical and platform-specific requirements.

Institutions

Categories

Social Sciences, Public Administration, Natural Language Processing, Language

Licence