IndoEng-MT Sentiment Dataset
Description
The IndoEng-MT Sentiment Dataset is a parallel Indonesian–English sentiment analysis dataset designed for cross-lingual NLP research, machine translation evaluation, and multilingual large language model (LLM) benchmarking. Each sample contains an original Indonesian text and its English machine-translated counterpart, enabling researchers to study sentiment preservation and consistency across languages. The dataset includes the following columns: - text: Original Indonesian text. - translated_text: English machine-translated version of the original text. - sentiment: Sentiment label associated with the text. - has_slang: Indicates whether the Indonesian text contains slang, informal expressions, or non-standard vocabulary. - has_codeswitch: Indicates whether the text contains code-switching between Indonesian and another language (e.g., English). - has_ambivalent: Indicates whether the text exhibits mixed, conflicting, or ambiguous sentiment cues that may complicate sentiment classification.
Files
Steps to reproduce
To ensure reproducibility, the complete prompt templates used to evaluate multilingual Large Language Models (LLMs) with this dataset are provided in the folder. The experiments employed three prompting strategies: Baseline Zero-Shot, Guided Zero-Shot, and Few-Shot.
Institutions
- Binus UniversityJakarta, Jakarta