Assamese-English Code Mixed dataset with Emotion Labelling using Plutchik’s model
Description
The rising trend of multilingual discourse on digital channels has led to the creation of a substantial corpus of text mixed with Assamese and English languages. Despite the growing sophistication in multilingual natural language processing (NLP), freely available datasets on emotions annotated for code-mixed text have not been extensively explored yet. Assamese is a low resource langauge and finding pure code mixed data is a trivial task.The dataset on Assamese-English code mixed data serves as a balanced code-mixed corpus of Assamese and English for multi-class emotion recognition in a low-resource multilingual environment. The dataset features linguistic patterns at the token level, code-mixing statistics, structural analysis using part-of-speech tagging, and emotion labeling, which aid in the study of computational linguistics. The data set can be repurposed for benchmarking machine learning, deep learning, and transformers-based multilingual natural language processing models. This dataset is useful for studies on code-mixing analysis, emotion detection, processing of low-resource languages, and cross-lingual representation learning. Code mixing is of two types inter sentential and intra sentential. The dataset contains 9456 intra-sentential switching instances, indicating that most bilingual mixing occurs within sentence boundaries, rest of the data is Inter-sentential switching among the 10,000 sentences. Parts of speech tagging of the sentence is also done.
Files
Institutions
- Gauhati UniversityAssam, Guwahati