AudioHate-DB

Published: 22 June 2026| Version 1 | DOI: 10.17632/hhd57tyy8y.1
Contributor:

Description

This dataset contains audio recordings of English speech spoken by Bangladeshi speakers, which can be used for research on hate speech detection, processing of accented-English speech and audio-based natural language understanding. The dataset aims to fill an important gap in the current resources, as it records English hate and non-hate speech as expressed in the phonetic and prosodic features of the Bangladeshi accent. The data set consists of a total of 3018 audio samples, with 1515 samples of hate speech and 1503 samples of non-hate speech. The almost balanced distribution allows to train and test binary classification models without employing elaborate techniques to balance the classes. The data contained within it are divided into two kinds: 1. Hate Speech: 1,515 audio samples with English utterances that are classified as hate speech. 2. Non-Hate Speech: 1,503 audiosamples that include neutral or non-offensive English speech. Both versions of the data are available. Original recordings of audio data in their natural state without changes and processed data: audio recordings that have been cleaned and redacted for any potential noise to provide standardized audio that can be directly used in any pipelines. Audio files are available in WAV format and are sorted into distinctly labeled folders based on the given class (hate / non_hate) and version (raw / processed).

Files

Categories

Natural Language Processing, Speech Recognition

Licence