Assamese-English Code Mixed dataset with Emotion Labelling using Plutchik’s model

Published: 30 June 2026| Version 1 | DOI: 10.17632/b8cnwznwxv.1
Contributors:
Afsana Laskar, Shikhar Kumar Sarma

Description

The rising trend of multilingual discourse on digital channels has led to the creation of a substantial corpus of text mixed with Assamese and English languages. Despite the growing sophistication in multilingual natural language processing (NLP), freely available datasets on emotions annotated for code-mixed text have not been extensively explored yet. Assamese is a low resource langauge and finding pure code mixed data is a trivial task.The dataset on Assamese-English code mixed data serves as a balanced code-mixed corpus of Assamese and English for multi-class emotion recognition in a low-resource multilingual environment. The dataset features linguistic patterns at the token level, code-mixing statistics, structural analysis using part-of-speech tagging, and emotion labeling, which aid in the study of computational linguistics. The data set can be repurposed for benchmarking machine learning, deep learning, and transformers-based multilingual natural language processing models. This dataset is useful for studies on code-mixing analysis, emotion detection, processing of low-resource languages, and cross-lingual representation learning. Code mixing is of two types inter sentential and intra sentential. The dataset contains 9456 intra-sentential switching instances, indicating that most bilingual mixing occurs within sentence boundaries, rest of the data is Inter-sentential switching among the 10,000 sentences. Parts of speech tagging of the sentence is also done.

Files

Institutions

Categories

Natural Language Processing

Licence