MaSERC: A Dual-Scenario Marathi Emotional Speech Corpus for Speech Emotion Recognition Research

Published: 22 June 2026| Version 1 | DOI: 10.17632/b9jrrpdxgb.1
Contributors:
,

Description

MaSERC (Marathi Speech Emotion Recognition Corpus) is the first emotional speech dataset for the Marathi language, contributed by students and faculty members of MIT Art, Design and Technology University, Pune, India. The corpus covers six basic emotion categories — Happy (आनंद), Sad (दु:ख), Angry (राग), Fear (भीती), Surprise (आश्चर्य), and Neutral (तटस्थ) — and contains approximately 804 audio recordings structured around two distinct speaking scenarios. In the first scenario, speakers deliver the single carrier sentence तू खरंच असं केलंस (Tu Kharach Asa Kelas, "You really did that?") once per emotion in a casual everyday conversation context; this sentence was chosen because its meaning shifts entirely with prosodic delivery, isolating acoustic-prosodic variation from lexical content. In the second scenario, speakers deliver a six-sentence Marathi monologue describing a student's examination experience, with each sentence embedding one of the six emotions in sequence — opening with happiness on seeing familiar questions, moving through surprise at the paper's difficulty, fear as questions became harder, anger at the paper's design, sadness at underperformance, and finally neutral reflection on the experience as a whole. All speakers are native Marathi speakers affiliated with MIT Art, Design and Technology University, Pune, aged 18 to 50 years, spanning students across academic years and a small number of faculty members. As a residential university drawing students from across Maharashtra and other Indian states, the cohort represents diverse Marathi dialectal backgrounds including Pune Metropolitan, Vidarbha, Marathwada, and Konkan varieties, alongside speakers from other linguistic regions who acquired Marathi during their studies, giving the corpus dialectal diversity uncommon in single-region corpora. Recordings were made in a partially controlled indoor environment using a professional condenser microphone at 44.1 kHz, 16-bit PCM stereo, released at 16 kHz mono WAV, with mean utterance durations of 3.93–5.52 seconds and a total duration of approximately 68 minutes. Files follow the naming convention [ID][SpeakerName][Year][Scenario][Emotion]-[RepNumber].wav. Emotion labels were validated by three independent annotators using majority vote, achieving Cohen's κ ≥ 0.78. Predefined train/validation/test splits, speaker metadata, and scenario labels are included in the release under CC BY 4.0.

Files

Categories

Emotion Representation, Meta Dataset

Licence