KurFemTTS is a collaborative initiative between the University of Kurdistan Hewlêr (UKH) and Kurdish Academia (KA) aimed at advancing speech and language technologies for Central Kurdish (Sorani), a low-resource language. The corpus consists of 10 hours of high-quality speech recordings collected from a single native female speaker with a Sorani accent. The recordings were conducted in a professional studio environment at Helen Radio, Erbil, Kurdistan Region, Iraq, ensuring high-quality and consistent speech data suitable for advanced speech processing research.
The dataset is distributed with comprehensive metadata in CSV format, containing speech-text alignment information and recording details required for training and evaluating speech AI models. The audio files are provided in WAV format with the following technical specifications:
File Format: WAV
Encoding Type: WAV (uncompressed audio)
Channel: Mono
Sample Rate: 22,024 Hz (standard configuration for TTS systems)
Speaker Profile: Single native female speaker
Language/Dialect: Central Kurdish (Sorani)
Accent: Sorani Kurdish accent
Total Duration: 10 hours
Metadata Format: CSV file containing audio-text pairs and associated recording information
The primary objectives of KurFemTTS are to provide a valuable linguistic resource for the development and evaluation of modern speech AI systems, including:
Text-to-Speech (TTS) Synthesis: Supporting the development of natural and high-quality Kurdish speech generation systems.
Automatic Speech Recognition (ASR): Enabling the training and evaluation of Kurdish speech recognition models.
Speaker Verification: Providing data for developing speaker identification and authentication systems.
Speaker Translation: Supporting speech-to-speech translation and multilingual communication models.
Voice Conversion: Facilitating research on transforming and adapting speaker characteristics while preserving linguistic content.
Speech and Language Models: Providing a foundation for training and evaluating advanced AI models for Kurdish language processing.
Low-Resource Language Advancement: Addressing the scarcity of high-quality Kurdish speech datasets and promoting research in underrepresented languages.
Inclusive and Multilingual AI Development: Contributing to culturally representative and accessible artificial intelligence technologies for Kurdish and other low-resource languages.
By providing a large-scale, high-quality Kurdish female speech resource, KurFemTTS aims to bridge the data scarcity gap in Kurdish speech technologies and accelerate the development of next-generation AI-driven language and speech applications.