Hinglish_dataset
Published: 30 March 2026| Version 1 | DOI: 10.17632/wtp7htm3ss.1
Contributors:
Adeeb Karbhari Adeeb, Jawad Roshan, Riyan Malakji, Gitanjali Shinde, Grishma Bobhate, Sonal FantangareDescription
The dataset used in this study consists of 2766 Hinglish sentences, collected from publicly available online sources such as social media platforms. Hinglish is a code-mixed language that combines Hindi and English words written in Roman script. Each sentence in the dataset is manually labeled into one of three sentiment classes: positive, negative, or neutral. The dataset contains informal language patterns, including slang, abbreviations, and mixed vocabulary, making it suitable for evaluating sentiment analysis models on real-world multilingual text.
Files
Categories
Natural Language Processing, Machine Learning, Deep Learning