UFTBESD: University of Frontier Technology, Bangladesh-Bangla Emotional Speech Dataset

Published: 23 July 2026| Version 4 | DOI: 10.17632/h3f8pjbw73.4
Contributors:
,
,
,
,
,
,
,

Description

UFTBESD (University of Frontier Technology, Bangladesh - Bangla Emotional Speech Dataset) is a Bangla-language speech emotion recognition dataset developed to capture realistic emotional speech under everyday acoustic conditions. The dataset consists of 1,400 audio recordings collected from 100 native Bangla speakers aged 15 to 38 years (mean age 22.6 years), with a balanced gender distribution (50 male and 50 female), recruited from all eight administrative divisions of Bangladesh. Each participant read two predefined Bangla sentences, with each sentence recorded once under each of seven target emotional states: angry, disgust, fear, happy, neutral, sad, and surprise. This yields 100 speakers x 2 sentences x 7 emotions = 1,400 audio clips. Every speaker contributes exactly 14 recordings, and every emotion category contains exactly 200 recordings (100 female, 100 male). All recordings were made using participants' own consumer smartphones in real indoor and outdoor environments (dormitory rooms, kitchens, courtyards, roadsides, etc.), without acoustic isolation or signal conditioning at any stage. A substantial portion of the corpus therefore contains natural background noise, device-specific microphone characteristics, and room reverberation, making it well suited to real-world speech emotion recognition research. Audio files are stored uncompressed as WAV (44.1 kHz, 16-bit PCM, mono), with durations ranging from 0.74 to 6.04 seconds (mean 2.68 s). This version corrects folder and file naming inconsistencies present in the previous release and adds a full metadata CSV file (one row per recording), covering speaker ID, pseudonym, age, gender, division, district, native dialect, education level, sentence text (Bangla and English), emotion label, recording environment, device brand/model, and audio properties. UFTBESD is intended to support the development and evaluation of Bangla speech emotion recognition systems and can be used with common machine learning and deep learning architectures such as CNN, LSTM, BiLSTM, and transformer-based models.

Files

Steps to reproduce

1. Speaker recruitment: 100 native Bangla speakers (50 male, 50 female), aged 15-38 years, were recruited from all eight administrative divisions of Bangladesh (Dhaka, Rajshahi, Khulna, Mymensingh, Rangpur, Barishal, Chittagong, Sylhet) through university-affiliated networks. Written informed consent was obtained from all participants (parental consent for the one minor participant, aged 15). 2. Script preparation: Two Bangla sentences with affectively neutral/underspecified propositional content were selected so that either sentence could be rendered under any of the seven target emotions without semantic conflict: S1: তুমি কিছুই বলছো না কেন? (Why aren't you saying anything?) S2: এটা আমি কখনো ভাবিনি। (I never thought about this.) 3. Recording session: Each participant was briefed on the seven target emotions (angry, disgust, fear, happy, neutral, sad, surprise) and encouraged to recall a relevant personal memory to support authentic portrayal. Each of the two sentences was then recorded once under each emotion, using the participant's own smartphone, in a naturally occurring indoor or outdoor location, with no directional microphones, studio equipment, or acoustic treatment. This produced 100 x 2 x 7 = 1,400 raw recordings. 4. Quality screening: Every recording was manually reviewed to confirm (i) the full target sentence was spoken without omission, (ii) the file was not truncated, corrupted, or affected by major dropouts, and (iii) overall audio quality was sufficient for analysis. Recordings failing any criterion were discarded and re-recorded. 5. Metadata annotation: Each accepted recording was logged with speaker ID, anonymised pseudonym, age, gender, division, district, native dialect, education level, sentence ID and text, emotion label, recording environment, device brand/model, and technical audio properties (sample rate, bit depth, channels, duration), compiled into a single metadata CSV. 6. Folder/file organization: Files were organized as UFTBESD/FEMALE|MALE/PERSON_N(PSEUDONYM)/SENTENCE_1|SENTENCE_2/, with filenames encoding participant index, gender, pseudonym, sentence index, and emotion label (e.g., P_12_F_12_MAYABI_S_02_ANGRY.wav). Naming inconsistencies from the prior version were identified and corrected before this release.

Institutions

Categories

Computer Science, Engineering, Machine Learning, Hidden Markov Model, Audio Signal Processing, Human-Computer Interaction, Bangladesh, Convolutional Neural Network, Long Short-Term Memory Network

Licence