Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kurdish (Sorani) Speech to Text: Presenting an Experimental Dataset

Domain:

natural language processing

Record type:

paperdataset
Creator:
QadHas
Host:avatar
We present an experimental dataset, Basic Dataset for Sorani Kurdish Automatic Speech Recognition (BD-4SK-ASR), which we used in the first attempt in developing an automatic speech recognition for Sorani Kurdish. The objective of the project was to develop a system that automatically could recognize simple sentences based on the vocabulary which is used in grades one to three of the primary schools in the Kurdistan Region of Iraq. We used CMUSphinx as our experimental environment. We developed a dataset to train the system. The dataset is publicly available for non-commercial use under the CC BY-NC-SA 4.0 license. 4 pages, 1 figure

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similar

Automatic Text Summarization (ATS) for Research Documents in Sorani KurdishKHLD: A Large-Scale Benchmark of the Kurdish Handwritten Lines Dataset for Low-Resource Central Kurdish (Sorani)KurdishMCQ: A multiple-choice question dataset for the Kurdish language (Sorani dialect)KurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-SpeechNamed Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative AnalysisDomain-Specific Machine Translation to Translate Medicine Brochures in English to Sorani Kurdish

Automatic Text Summarization (ATS) for Research Documents in Sorani Kurdish

Extracting concise information from scientific documents aids learners, researchers, and practitione

KHLD: A Large-Scale Benchmark of the Kurdish Handwritten Lines Dataset for Low-Resource Central Kurdish (Sorani)

The Kurdish Handwritten Lines Dataset (KHLD) is a large-scale image dataset aiming to facilitate han

KurdishMCQ: A multiple-choice question dataset for the Kurdish language (Sorani dialect)

This dataset is the first large-scale, publicly available collection of multiple-choice questions in

KurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech

KurFemTTS is a collaborative initiative between the University of Kurdistan Hewlêr (UKH) and Kurdish

Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative Analysis

This work contributes towards balancing the inclusivity and global applicability of natural language

Domain-Specific Machine Translation to Translate Medicine Brochures in English to Sorani Kurdish

Access to Kurdish medicine brochures is limited, depriving Kurdish-speaking communities of critical