Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language

Domain:

natural language processinghealthcare

Record type:

paperdataset
Creator:
WiaEkpSalAts
Host:avatar
The lack of impaired speech data hinders advancements in the development of inclusive speech technologies, particularly in low-resource languages such as Akan. To address this gap, this study presents a curated corpus of speech samples from native Akan speakers with speech impairment. The dataset comprises of 50.01 hours of audio recordings cutting across four classes of impaired speech namely stammering, cerebral palsy, cleft palate, and stroke induced speech disorder. Recordings were done in controlled supervised environments were participants described pre-selected images in their own words. The resulting dataset is a collection of audio recordings, transcriptions, and associated metadata on speaker demographics, class of impairment, recording environment and device. The dataset is intended to support research in low-resource automatic disordered speech recognition systems and assistive speech technology.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

Akan

Tags

SoundArtificial Intelligence

Similar

UGAkan-ImpairedSpeechData: A Dataset of Impaired Speech in the Akan LanguageSomali Automatic Speech Recognition DatasetSagalee: an Open Source Automatic Speech Recognition Dataset for Oromo LanguageVoxMg: An Automatic Speech Recognition Dataset for MalagasyAfriVox: An African benchmark dataset for Automatic Speech Translation and Speech RecognitionKARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

UGAkan-ImpairedSpeechData: A Dataset of Impaired Speech in the Akan Language

The UGAkan-ImpairedSpeechData is a speech dataset from indigenous speakers of Akan with different fo

Somali Automatic Speech Recognition Dataset

This dataset contains audio recordings and corresponding transcriptions in Somali, designed for auto

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoke

VoxMg: An Automatic Speech Recognition Dataset for Malagasy

African languages are not well-represented in Natural Language Processing (NLP). The main reason is a lack of resources for training models. Low-resource languages, such as Malagasy, cannot benefit from modern NLP methods if no datasets are available. This paper pr

AfriVox: An African benchmark dataset for Automatic Speech Translation and Speech Recognition

This project creates a benchmark dataset for evaluating Automatic Speech Translation and Speech reco

KARAKALPAK SPEECH CORPUS: THE FIRST BENCHMARK DATASET FOR AUTOMATIC SPEECH RECOGNITION

While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recog