Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus

Domaine:

natural language processing

Type de record:

paper

We introduce the ÌròyìnSpeech corpus -- a new dataset influenced by a desire to increase the amount of high quality, freely available, contemporary Yorùbá speech. We release a multi-purpose dataset that can be used for both TTS and ASR tasks. We curated text sentences from the news and creative writing domains under an open license i.e., CC-BY-4.0 and had multiple speakers record each sentence. We provide 5000 of our utterances to the Common Voice platform to crowdsource transcriptions online. The dataset has 38.5 hours of data in total, recorded by 80 volunteers.

Visit

arxiv.org

Tasks

automatic speech recognitiontext to speechspeech processing

Languages

Yoruba

Tags

Similaires

MENYO-20k: A Multi-domain English - Yorùbá Corpus for Machine TranslationYoruba Multi-Speaker Speech CorpusOMAN-SPEECH: A Multi-Layer Annotated Speech Corpus for Omani Arabic DialectsA general-purpose IsiZulu speech synthesizer"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"BIG-C: a Multimodal Multi-Purpose Dataset for Bemba

MENYO-20k: A Multi-domain English - Yorùbá Corpus for Machine Translation

MENYO-20k is a multi-domain parallel dataset with texts obtained from news articles, ted ta

Yoruba Multi-Speaker Speech Corpus

Yoruba tts notebook and data

OMAN-SPEECH: A Multi-Layer Annotated Speech Corpus for Omani Arabic Dialects

Automatic Speech Recognition (ASR) has achieved strong performance in high-resource languages; howev

A general-purpose IsiZulu speech synthesizer

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"

"This dataset is a 20.03-hour Amharic speech corpus recorded from 100 native speakers and containing

BIG-C: a Multimodal Multi-Purpose Dataset for Bemba

We present BIG-C (Bemba Image Grounded Conversations), a large multimodal dataset for Bemba. While Bemba is the most populous language of Zambia, it exhibits a dearth of resources which render the development of language technologies or language processing resea