Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
DiaNelAgbNak
Hôte:avatar
The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at Waxal NLP Datasets under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages. Initial dataset release with added TTS, some more to come

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processingtext to speech

Tags

Audio and Speech ProcessingArtificial IntelligenceComputation and Languagewaxal

Similaires

BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpusWaxal-Multilingual/speech-dataKurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-SpeechLughaGen Multilingual African Language CorpusToward the Creation of a Large-Scale Moroccan Sign Language CorpusRomanization-based Large-scale Adaptation of Multilingual Language Models

BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus

BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speec

Waxal-Multilingual/speech-data

This repository contains multi-modal speech data for African languages that can be used to train ASR

KurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech

KurFemTTS is a collaborative initiative between the University of Kurdistan Hewlêr (UKH) and Kurdish

LughaGen Multilingual African Language Corpus

LughaGen is a curated multilingual corpus for four Kenyan and East African languages: Swahili (sw),

Toward the Creation of a Large-Scale Moroccan Sign Language Corpus

Abstract: This article reports on an ongoing research project that aims to build a foundation for re

Romanization-based Large-scale Adaptation of Multilingual Language Models

Large multilingual pretrained language models (mPLMs) have become the de facto state of the art for