IndoEng-MT Sentiment Dataset

Published: 7 August 2026| Version 2 | DOI: 10.17632/p6prw6chsz.2
Contributors:
,
,

Description

The IndoEng-MT Sentiment Dataset is a parallel Indonesian–English sentiment analysis dataset designed for cross-lingual NLP research, machine translation evaluation, and multilingual large language model (LLM) benchmarking. Each sample contains an original Indonesian text and its English machine-translated counterpart, enabling researchers to study sentiment preservation and consistency across languages. The dataset includes the following columns: - text: Original Indonesian text. - translated_text: English machine-translated version of the original text. - sentiment: Sentiment label associated with the text. - has_slang: Indicates whether the Indonesian text contains slang, informal expressions, or non-standard vocabulary. - has_codeswitch: Indicates whether the text contains code-switching between Indonesian and another language (e.g., English). - has_ambivalent: Indicates whether the text exhibits mixed, conflicting, or ambiguous sentiment cues that may complicate sentiment classification.

Files

Steps to reproduce

To ensure reproducibility, the complete prompt templates used to evaluate multilingual Large Language Models (LLMs) with this dataset are provided in the folder. The experiments employed three prompting strategies: Baseline Zero-Shot, Guided Zero-Shot, and Few-Shot.

Institutions

Categories

Indonesian, Sentiment Analysis, Multilingual LLM

Licence