Phonics Exercise Audio Dataset

Published: 11 June 2025| Version 1 | DOI: 10.17632/rmkznfhfsh.1
Contributors:
Fatma Yehia, Rana Abubakr, Ali Mohamed, Saif Allah Mostafa, Merna Fouad , Hend Ali, Peter Elia, Gehad Ismail Sayed

Description

The Phonics Exercise Audio Dataset is a custom-built audio corpus developed to support speech-based learning and pronunciation correction. It currently consists of 198 audio samples that represent correct and incorrect pronunciations of Arabic letters from "أ" to "ت". These samples are recorded in a standardized format to support consistent feature extraction and model training. The dataset enables the development of phoneme classification systems aimed at improving auditory discrimination skills in dyslexic learners. Ongoing data collection efforts seek to expand this resource to include more letters and speakers, increasing the system’s effectiveness in varied learning contexts.

Files

Institutions

  • Canadian College International Institute

Categories

Learning, Pronunciation, Speech Correction

Licence