Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus

Domaine:

natural language processing

Type de record:

paper
BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speech models. The ten languages represented are: Akuapem Twi, Asante Twi, Chichewa, Ewe, Hausa, Kikuyu, Lingala, Luganda, Luo, and Yoruba. This corpus is a derivative work of Bible recordings made and released by the Open.Bible project from Biblica. We have aligned, cleaned, and filtered the original recordings, and additionally hand-checked a subset of the alignments for each language. We present results for text-to-speech models with Coqui TTS. The data is released under a commercial-friendly CC-BY-SA license.

Visit

arxiv.org

Connected records

dataset

Tasks

speech translationtext to speechspeech processingautomatic speech recognition

Languages

AkanBwamu, CwiChichewaDinka, SoutheasternÉwéGandaGikuyuHausaLingalaYoruba

Tags

bibletts

Licenses

CC-BY-SA

Similaires

WAXAL: A Large-Scale Multilingual African Language Speech CorpusCommon Voice: A Massively-Multilingual Speech CorpusNCHLT Auxiliary Speech Corpus - MultilingualA First South African Corpus of Multilingual Code-switched Soap Opera SpeechNoise-Robust Multilingual Speech Recognition and the Tatar Speech CorpusZambezi Voice: A Multilingual Speech Corpus for Zambian Languages

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech datase

Common Voice: A Massively-Multilingual Speech Corpus

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identi

NCHLT Auxiliary Speech Corpus - Multilingual

This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data S

A First South African Corpus of Multilingual Code-switched Soap Opera Speech

We introduce a speech corpus containing multilingual code-switching compiled from South African soap operas. The corpus contains English, isiZulu, isiXhosa, Setswana and Sesotho speech, paired into four language-balanced subcorpora containing English-isiZulu, Engli

Noise-Robust Multilingual Speech Recognition and the Tatar Speech Corpus

After focusing on individual languages for a long time, multilingual automatic speech recognition ha

Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours) and labelled data (over 80 hours) consistin