Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Synthetic Text Corpus for African Language ASR

Domain:

natural language processing

Record type:

dataset
Creator:
CLEAR Global
Host:
This dataset contains 13,488 synthetic sentences across 10 African languages (Bambara, Chichewa, Hausa, Kanuri, Luo, Nande, Somali, Twi, Wolof, Yoruba) generated using large language models (GPT-4o, GPT-4.5, Claude 3.5 Sonnet, Claude 3.7 Sonnet). Each sentence has been evaluated by human linguists on readability and naturalness (1-7 scale), translation adequacy and accuracy (1-7 scale), grammatical correctness, word validity, and presence of notable errors. Corrected versions are provided where applicable. The dataset was created by Dimagi to support ASR, NLP research, and evaluation for low-resource African languages. See: DeRenzi et al. (2025), "Synthetic Voice Data for Automatic Speech Recognition in African Languages", arXiv:2507.17578.

Visit

mozilladatacollective.com

Connected records

paper

Tasks

automatic speech recognitionspeech processing

Languages

BamanankanChichewaDholuoHausaKanembuKanuri, MangaKanuri, YerwaLameNandeSomali+2

Tags

mdcmozilla data collectiveNLPTSV

Licenses

Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)

Similar

Using ASR-Generated Text for Spoken Language ModelingAfrican Multilingual Text CorpusUrhobo language training text for NLP, ASR and TTS tasksUsable Amharic text corpus for natural language processing applicationssomali-asr-synthetic-youtubeAfroMAFT Corpus: Language Adaptation Corpus for African languages

Using ASR-Generated Text for Spoken Language Modeling

International audience This papers aims at improving spoken language modeling (LM) us

African Multilingual Text Corpus

NLP model training, translation Notes / challenges: Unequal language representation

Urhobo language training text for NLP, ASR and TTS tasks

Urhobo language training text for NLP, ASR and TTS tasks

Usable Amharic text corpus for natural language processing applications

somali-asr-synthetic-youtube

AfroMAFT Corpus: Language Adaptation Corpus for African languages

Language Adaptation Corpus for 17 African languages, English, French, and Arabic.