KurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech
Description
KurFemTTS is a collaborative initiative between the University of Kurdistan Hewlêr (UKH) and Kurdish Academia (KA) aimed at advancing speech and language technologies for Central Kurdish (Sorani), a low-resource language. The corpus consists of 10 hours of high-quality speech recordings collected from a single native female speaker with a Sorani accent. The recordings were conducted in a professional studio environment at Helen Radio, Erbil, Kurdistan Region, Iraq, ensuring high-quality and consistent speech data suitable for advanced speech processing research. The dataset is distributed with comprehensive metadata in CSV format, containing speech-text alignment information and recording details required for training and evaluating speech AI models. The audio files are provided in WAV format with the following technical specifications: File Format: WAV Encoding Type: WAV (uncompressed audio) Channel: Mono Sample Rate: 22,024 Hz (standard configuration for TTS systems) Speaker Profile: Single native female speaker Language/Dialect: Central Kurdish (Sorani) Accent: Sorani Kurdish accent Total Duration: 10 hours Metadata Format: CSV file containing audio-text pairs and associated recording information The primary objectives of KurFemTTS are to provide a valuable linguistic resource for the development and evaluation of modern speech AI systems, including: Text-to-Speech (TTS) Synthesis: Supporting the development of natural and high-quality Kurdish speech generation systems. Automatic Speech Recognition (ASR): Enabling the training and evaluation of Kurdish speech recognition models. Speaker Verification: Providing data for developing speaker identification and authentication systems. Speaker Translation: Supporting speech-to-speech translation and multilingual communication models. Voice Conversion: Facilitating research on transforming and adapting speaker characteristics while preserving linguistic content. Speech and Language Models: Providing a foundation for training and evaluating advanced AI models for Kurdish language processing. Low-Resource Language Advancement: Addressing the scarcity of high-quality Kurdish speech datasets and promoting research in underrepresented languages. Inclusive and Multilingual AI Development: Contributing to culturally representative and accessible artificial intelligence technologies for Kurdish and other low-resource languages. By providing a large-scale, high-quality Kurdish female speech resource, KurFemTTS aims to bridge the data scarcity gap in Kurdish speech technologies and accelerate the development of next-generation AI-driven language and speech applications.
Files
Institutions
- University of Kurdistan HewlerErbil, Erbil