Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR

Domaine:

natural language processinghealthcare

Type de record:

datasetmodel
Créateur:
TobTejAdiChr
Hôte:avatar
Africa has a very poor doctor-to-patient ratio. At very busy clinics, doctors could see 30+ patients per day—a heavy patient burden compared with developed countries—but productivity tools such as clinical automatic speech recognition (ASR) are lacking for these overworked clinicians. However, clinical ASR is mature, even ubiquitous, in developed nations, and clinician-reported performance of commercial clinical ASR systems is generally satisfactory. Furthermore, the recent performance of general domain ASR is approaching human accuracy. However, several gaps exist. Several publications have highlighted racial bias with speech-to-text algorithms and performance on minority accents lags significantly. To our knowledge, there is no publicly available research or benchmark on accented African clinical ASR, and speech data is non-existent for the majority of African accents. We release AfriSpeech, 200hrs of Pan-African English speech, 67,577 clips from 2,463 unique speakers across 120 indigenous accents from 13 countries for clinical and general domain ASR, a benchmark test set, with publicly available pre-trained models with SOTA performance on the AfriSpeech benchmark.

Visit

figshare.com

Tasks

automatic speech recognitionspeech processing

Tags

Trustworthy Information Processing46 Information and Computing Sciences4602 Artificial Intelligence47 Language, Communication and Culture4704 LinguisticsTrustworthy Information Processing RA2Clinical ResearchBehavioral and Social Science0801 Artificial Intelligence and Image Processing1702 Cognitive Sciences+1

Licenses

CC BY 4.0

Similaires

AfriSpeech-MultiBench: A Verticalized Multidomain Multicountry Benchmark Suite for African Accented English ASRDomain Adversarial Training for Accented Speech RecognitionIntron AfriSpeech-200 Automatic Speech Recognition ChallengeAfriSpeech-200Performant ASR Models for Medical Entities in Accented SpeechImproving Accented Speech Recognition with Multi-Domain Training

AfriSpeech-MultiBench: A Verticalized Multidomain Multicountry Benchmark Suite for African Accented English ASR

Recent advances in speech-enabled AI, including Google's NotebookLM and OpenAI's speech-to-speech AP

Domain Adversarial Training for Accented Speech Recognition

In this paper, we propose a domain adversarial training (DAT) algorithm to alleviate the accented sp

Intron AfriSpeech-200 Automatic Speech Recognition Challenge

Can you create an automatic speech recognition (ASR) model for African accents, for use by doctors? African hospitals have some of the lowest doctor-patient ratios in the world. At very busy clinics, doctors could see over 30 patients a day without any of the prod

AfriSpeech-200

AFRISPEECH-200 is a 200hr Pan-African speech corpus for clinical and general domain English accented

Performant ASR Models for Medical Entities in Accented Speech

Recent strides in automatic speech recognition (ASR) have accelerated their application in the medic

Improving Accented Speech Recognition with Multi-Domain Training

5 pages, 2 figures. Accepted to ICASSP 2023 International audience Thanks to the rise