Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

KurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech

Domain:

natural language processing

Record type:

dataset
Creator:
AbdRasSyaFat
Editor:
Uni
Publisher:
Men
Host:avatar
KurFemTTS is a collaborative initiative between the University of Kurdistan Hewlêr (UKH) and Kurdish Academia (KA) aimed at advancing speech and language technologies for Central Kurdish (Sorani), a low-resource language. The corpus consists of 10 hours of high-quality speech recordings collected from a single native female speaker with a Sorani accent. The recordings were conducted in a professional studio environment at Helen Radio, Erbil, Kurdistan Region, Iraq, ensuring high-quality and consistent speech data suitable for advanced speech processing research. The dataset is distributed with comprehensive metadata in CSV format, containing speech-text alignment information and recording details required for training and evaluating speech AI models. The audio files are provided in WAV format with the following technical specifications: File Format: WAV Encoding Type: WAV (uncompressed audio) Channel: Mono Sample Rate: 22,024 Hz (standard configuration for TTS systems) Speaker Profile: Single native female speaker Language/Dialect: Central Kurdish (Sorani) Accent: Sorani Kurdish accent Total Duration: 10 hours Metadata Format: CSV file containing audio-text pairs and associated recording information The primary objectives of KurFemTTS are to provide a valuable linguistic resource for the development and evaluation of modern speech AI systems, including: Text-to-Speech (TTS) Synthesis: Supporting the development of natural and high-quality Kurdish speech generation systems. Automatic Speech Recognition (ASR): Enabling the training and evaluation of Kurdish speech recognition models. Speaker Verification: Providing data for developing speaker identification and authentication systems. Speaker Translation: Supporting speech-to-speech translation and multilingual communication models. Voice Conversion: Facilitating research on transforming and adapting speaker characteristics while preserving linguistic content. Speech and Language Models: Providing a foundation for training and evaluating advanced AI models for Kurdish language processing. Low-Resource Language Advancement: Addressing the scarcity of high-quality Kurdish speech datasets and promoting research in underrepresented languages. Inclusive and Multilingual AI Development: Contributing to culturally representative and accessible artificial intelligence technologies for Kurdish and other low-resource languages. By providing a large-scale, high-quality Kurdish female speech resource, KurFemTTS aims to bridge the data scarcity gap in Kurdish speech technologies and accelerate the development of next-generation AI-driven language and speech applications.

Visit

doi.org

Tasks

text to speechspeech processing

Tags

Robust Speech RecognitionSpeech AdaptationCorpus LinguisticsKurdSpeech Synthesis

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution Non Commercial No Derivatives 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-nd/4.0/legalcode

Similar

A Comprehensive Kurdish Speech Corpus for Speaker Identification and VerificationKurdish (Sorani) Speech to Text: Presenting an Experimental DatasetWAXAL: A Large-Scale Multilingual African Language Speech CorpusAn Amharic speech corpus for large vocabulary continuous speech recognitionTunArTTS: Tunisian Arabic Text-To-Speech CorpusDesign of a Yoruba Language Speech Corpus for the Purposes of Text-to-Speech (TTS) Synthesis

A Comprehensive Kurdish Speech Corpus for Speaker Identification and Verification

Abstract / General Description: This dataset comprises a proprietary acoustic corpus specifically de

Kurdish (Sorani) Speech to Text: Presenting an Experimental Dataset

We present an experimental dataset, Basic Dataset for Sorani Kurdish Automatic Speech Recognition (B

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech datase

An Amharic speech corpus for large vocabulary continuous speech recognition

TunArTTS: Tunisian Arabic Text-To-Speech Corpus

Design of a Yoruba Language Speech Corpus for the Purposes of Text-to-Speech (TTS) Synthesis