Hinglish_dataset

Published: 30 March 2026| Version 1 | DOI: 10.17632/wtp7htm3ss.1
Contributors:
Adeeb Karbhari Adeeb, Jawad Roshan, Riyan Malakji, Gitanjali Shinde, Grishma Bobhate, Sonal Fantangare

Description

The dataset used in this study consists of 2766 Hinglish sentences, collected from publicly available online sources such as social media platforms. Hinglish is a code-mixed language that combines Hindi and English words written in Roman script. Each sentence in the dataset is manually labeled into one of three sentiment classes: positive, negative, or neutral. The dataset contains informal language patterns, including slang, abbreviations, and mixed vocabulary, making it suitable for evaluating sentiment analysis models on real-world multilingual text.

Files

Categories

Natural Language Processing, Machine Learning, Deep Learning

Licence