Dataset for Sentiment Analysis in Code-Mixed Language

Published: 10 October 2025| Version 1 | DOI: 10.17632/xxprggbztd.1
Contributors:
, Dr Muhammad Nouman Noor

Description

This data is for sentiment analysis based on low resource code-mixed languages like Roman Urdu, Roman Hindi, and Roman English. The dataset is initially collected from ecommerce platform and is based on user reviews on the platform. After collection, data is cleaned and can be used for sentiment analysis research. This dataset classifies the reviews on three classes Positive, Negative and Neutral.

Files

Institutions

  • National University of Computer and Emerging Sciences

Categories

Natural Language Processing, Sentiment Analysis

Licence