Assamese-English Code Mix PoS Tagged Sentiment label data

Published: 30 June 2026| Version 1 | DOI: 10.17632/42ddr9tv3h.1
Contributors:
Afsana Laskar,

Description

This dataset provides a comprehensive resource for studying Assamese-English code-mixed language, making it particularly valuable for tasks such as part-of-speech (POS) tagging, emotion detection, Named-Entity Recognition, sentiment analysis, etc. Code-mixed text, which blends words and phrases from multiple languages within a single sentence, presents unique linguistic challenges and opportunities. This dataset specifically focuses on Assamese-English code-mixed text, offering a diverse set of sentences collected from newspapers, journals, and storybooks. The choice of sources ensures a wide range of linguistic contexts and realistic language usage, making the dataset suitable for exploring the complexities of code-mixed languages and their applications. Finding a proper dataset for low resourced languages is tough. The dataset is composed of Assamese-English code-mixed sentences, where linguistic elements from both languages coexist within the same sentence. Such intermixing captures the natural bilingual expressions found in real-world communication. The sentences were carefully gathered from sources like Assamese and English newspapers, journals, and storybooks. These sentences were then manually converted to Assamese-English code-mixed texts/sentences, ensuring that the data reflects authentic and meaningful usage of both languages. This makes the dataset particularly relevant for various NLP tasks like POS tagging and sentiment analysis, which requires accurate identification of linguistic components across languages. Although the dataset was initially designed for POS tagging, its structure and annotations make it highly suitable for a wide range of NLP applications such as sentiment analysis. The diversity of sources and the natural intermingling of languages offer a rich ground for studying how sentiments are expressed in bilingual contexts. Each sentence is annotated with linguistic details, allowing researchers to explore how sentiments are expressed in bilingual contexts, particularly in a code-mixed format.

Files

Institutions

Categories

Natural Language Processing

Licence