AUST-Sarc (Emoji-Aware Bangla Sarcasm & Offensiveness)

Published: 23 September 2025| Version 2 | DOI: 10.17632/7ryvn5gw88.2
Contributors:
,
,
,
, Jaasia Anjum

Description

AUST-SARC is a Bangla text dataset for sarcasm and offensiveness detection with explicit emoji context. Each entry is a short user-generated sentence containing ≥1 emoji, labeled with Sarcasm (1/0) and Offensive (1/0). Size: 2,649 sentences Sarcasm distribution: 1 = 1,496, 0 = 1,153 Offensive distribution: 1 = [fill], 0 = [fill] Language: Bangla (with natural Bangla–English code-mixing) Preprocessing: emojis preserved; light normalization of URLs/mentions/hashtags; whitespace cleanup. Intended use: sarcasm detection, offensive/abusive language detection, emoji-aware pragmatics, multitask learning. Ethics: personally identifying info removed; texts may contain sensitive or offensive content due to the task.

Files

Steps to reproduce

Data source and contributors • Four human contributors (native Bangla speakers) created and curated the texts. • Sentences came from two channels: (1) original authoring (our own humor, everyday phrasing, code-mix) and (2) publicly visible Facebook comments (no private groups or DMs). • Only sentences that contained at least one emoji were eligible. Inclusion / exclusion rules • Include: short Bangla or Bangla–English code-mixed sentences with ≥1 emoji. • Exclude: entries without emojis; near-duplicates; content containing PII (names, phone numbers, emails, addresses), URLs with tracking, or images/media. • For social-media items, we removed user handles and links, and paraphrased if needed to avoid re-identifying individuals while keeping meaning.

Institutions

  • Ahsanullah University of Science and Technology

Categories

Natural Language Processing

Licence