Madurese-Language Social Media Comment Dataset for Public Service Analysis
Description
This dataset contains 1,510 public comments written primarily in Madurese, collected from social media platforms including TikTok, Facebook, Instagram, and X. The comments discuss public services, government programs, infrastructure, and community issues in Madura, Indonesia. Each record includes the following fields: platform: The social media platform where the comment was posted. kategori_layanan: The public-service or issue category associated with the comment, such as roads and bridges, public lighting, clean water, healthcare, education, waste management, parking, flooding, public transportation, and other government-related issues. teks_mentah: The original comment text, preserved in its collected form, including local spelling variations, informal language, abbreviations, punctuation, emojis, and code-switching. code_mixing: A binary annotation indicating whether the comment contains code-mixing between Madurese and another language, primarily Indonesian. sarkasme: A binary annotation indicating whether the comment contains sarcastic or ironic expression. cleaning: A cleaned-text version of the original comment used for preprocessing consistency. baku: A normalized text version in standardized spelling/form. baku_steming: A normalized-and-stemmed text version for downstream NLP modeling. label: The final binary class label used for supervised learning experiments. The dataset is intended to support research in low-resource language processing, social media analysis, public-service evaluation, code-mixing detection, sarcasm identification, and public opinion studies in Madurese-speaking communities. It is also suitable for developing and evaluating NLP models on regional languages, informal user-generated content, and multilingual communication scenarios.
Files
Steps to reproduce
1. Identify public posts/comments about public services, government programs, infrastructure, and community issues in Madura. 2. Manually browse TikTok, Facebook, Instagram, and X using normal platform features only (no scraping, bots, or automation). 3. Select relevant public comments and record platform, service category, and original text. 4. Keep the original comment as teks_mentah (including informal spelling, mixed language, abbreviations, punctuation, and emojis). 5. Annotate code_mixing: 1 if Madurese is mixed with another language (mainly Indonesian), 0 otherwise. 6. Annotate sarkasme: 1 if sarcasm/irony is present in context, 0 otherwise. 7. Produce derived text fields: cleaning (cleaned text), baku (normalized text), and baku_steming (normalized + stemmed text). 8. Assign final binary label (label), perform sample-based inter-rater validation, review consistency, and export a semicolon-delimited CSV with fields: platform, kategori_layanan, teks_mentah, code_mixing, sarkasme, cleaning, baku, baku_steming, label. Note: Social media content may be edited, removed, or unavailable over time, so exact recollection may not be fully replicable. Data were collected manually from public content with attention to ethics and platform policies.
Institutions
- Universitas MaduraEast Java, Pamekasan
Categories
Funders
- Direktorat Riset, Teknologi, dan Pengabdian kepada Masyarakat (DRTPM)Grant ID: Penelitian Dosen Pemula (PDP) - 286/C3/DT.05.00/PL-BARU/2026