Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
MutMugNyaChe
Hôte:avatar
We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across four language families: Swahili, Kikuyu, Kamba, Kimeru, Luo, Maasai, Kipsigis, Somali (East Africa); Wolof (West Africa); and Fulani (West/Central Africa). The dataset contains over 601,000 approved sentence-level text annotations and over 385,000 audio recordings, collected through a dedicated community data collection platform involving over 100 contributors. To validate the dataset's utility, we train and evaluate ASR, MT, and TTS models, establishing baselines across all languages. Our best ASR system achieves 3.24% WER on Swahili (Common Voice), reducing prior academic SOTA from 8.3% to 3.24% (5.1 percentage point absolute, 61% relative reduction), and 4.3% WER on Somali. The dataset will be published on HuggingFace. We describe the collection platform, quality assurance workflows, and baseline experiments, and discuss implications for African language technology infrastructure.

Visit

arxiv.org

Tasks

automatic speech recognitionmachine translationspeech processingtext to speech

Languages

FulaFulfulde, AdamawaFulfulde, BorguFulfulde, Central-Eastern NigerFulfulde, MaasinaFulfulde, NigerianFulfulde, Western NigerGikuyuKambaKimîîru+7

Tags

Computation and LanguageMachine Learning

Similaires

Large Multimodal Models for Low-Resource Languages: A SurveyLarge Language Models Adaptation for Low-resource Languages: The Case for African LanguagesOpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource LanguagesMultimodal Sarcasm Dataset Generation for a Low-Resource Language: SwahiliWA-Vox: Building a Large Scale, Inclusive Speech Dataset for West African LanguagesAfrican Voices: Multilingual Speech Dataset for Low-Resource African Languages

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

Large Language Models Adaptation for Low-resource Languages: The Case for African Languages

David Ifeoluwa Adelani (Supervisor) Despite remarkable advances in Large Language Models (LLMs), Afr

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially

Multimodal Sarcasm Dataset Generation for a Low-Resource Language: Swahili

WA-Vox: Building a Large Scale, Inclusive Speech Dataset for West African Languages

Empowering West African Voices: Introducing WA-VOX West Africa, home to hundreds of millions across

African Voices: Multilingual Speech Dataset for Low-Resource African Languages

A large-scale multilingual speech dataset developed by Data Science Nigeria. Contains more than 3,000 hours of transcribed audio across four Nigerian languages: Hausa, Igbo, Nigerian Pidgin, and Yorùbá. The dataset supports Automatic Speech Recognition (ASR) and sp