Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

UGAkan-ImpairedSpeechData: A Dataset of Impaired Speech in the Akan Language

Domain:

natural language processinghealthcare

Record type:

dataset
Creator:
WiaAbdSalEkp
Editor:
Uni
Publisher:
Men
Host:avatar
The UGAkan-ImpairedSpeechData is a speech dataset from indigenous speakers of Akan with different forms of speech impairments. It contains audio descriptions of culturally relevant images. The dataset comprises 14,312 audio files and corresponding transcriptions equivalent to 50.01 hours. Recordings were done in different environments including Outdoor (7,706 audio files), Other (3,075), Indoor (2,254), Studio (982), and Car (295). The dataset is also categorized by aetiology and gender. Male speakers contributed 6,754 files equivalent to 19.02 hours, with the highest representation from individuals with Cerebral Palsy (2,881 files, 8.84 hours), followed by Stammering, Cleft, and Stroke. Female speakers contributed 7,558 files equivalent to 30.99 hours, with most recordings coming from individuals with Cerebral Palsy (4,835 files, 15.66 hours) and Stammering (2,574 files, 13.88 hours). Stroke data was recorded only from male speakers, while Cleft speech samples were collected from both genders, with a higher volume from males. In terms of duration, the audio files vary in length. The average audio length is 12.46 seconds, with a standard deviation of 7.71 seconds, indicating moderate variability. The majority of audio files range from 6.59s to 16.00s, suggesting a right-skewed distribution. The maximum duration is 60.08s, which exceeds the upper bound of the interquartile range and is likely an outlier.

Visit

doi.orgdata.mendeley.com

Tasks

speech processing

Languages

Akan

Tags

LinguisticsFOS: Languages and literatureNatural Language ProcessingMachine TranslationSpeech RecognitionSpeech DisorderText-to-SpeechTransformer LLMLow-Resource LLM

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution Non Commercial No Derivatives 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-nd/4.0/legalcode

Similar

Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan LanguageIS LANGUAGE TRULY A RIGHT? Deaf and speech impaired people in South Africa

Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language

The lack of impaired speech data hinders advancements in the development of inclusive speech technol

IS LANGUAGE TRULY A RIGHT? Deaf and speech impaired people in South Africa

IS LANGUAGE TRULY A RIGHT? Deaf and speech impaired people in South Africa

Poster presented at the Deep Learning Indaba 2023 by tsakani shilowe