Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects

Domaine:

natural language processing

Type de record:

paper

Yorùbá an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in disparities for dialects and varieties for which there are little to no resources or tools. We take steps towards bridging this gap by introducing a new high-quality parallel text and speech corpus YORÙLECT across three domains and four regional Yorùbá dialects. To develop this corpus, we engaged native speakers, travelling to communities where these dialects are spoken, to collect text and speech data. Using our newly created corpus, we conducted extensive experiments on (text) machine translation, automatic speech recognition, and speech-to-text translation. Our results reveal substantial performance disparities between standard Yorùbá and the other dialects across all tasks. However, we also show that with dialect-adaptive finetuning, we are able to narrow this gap. We believe our dataset and experimental analysis will contribute greatly to developing NLP tools for Yorùbá and its dialects, and potentially for other African languages, by improving our understanding of existing challenges and offering a high-quality dataset for further development. We release YORÙLECT dataset and models publicly under an open license.

Visit

arxiv.orgProject on Github

Connected records

dataset

Tasks

machine translationautomatic speech recognitionspeech processingspeech translation

Languages

Yoruba

Tags

yorulect

Similaires

Iraqi Arabic NLP Toolkit (IANLP): Building Language Resources for Low-Resource DialectsSaptak: A Large-scale Multi-Regional Benchmark Dataset for Poly-Dialectal Neural Machine Translation between Standard Bangla and Regional Dialects, and among Regional DialectsDocumentation and description of unheard voices:<i>k'aʔanniʃe</i>and<i>ɗenke</i>of the GanjuleBayelemabaga: Creating Resources for Bambara NLPThe RTR Harmonic Domain in Two Dialects of YorùbáNiger Volta LTI: Yorùbá language training text for NLP, ASR and TTS tasks

Iraqi Arabic NLP Toolkit (IANLP): Building Language Resources for Low-Resource Dialects

A survey and resource-oriented study examining the current state of Natural Language Processing (NLP

Saptak: A Large-scale Multi-Regional Benchmark Dataset for Poly-Dialectal Neural Machine Translation between Standard Bangla and Regional Dialects, and among Regional Dialects

This dataset is a comprehensive parallel corpus developed for poly-dialectal neural machine translat

Documentation and description of unheard voices:<i>k'aʔanniʃe</i>and<i>ɗenke</i>of the Ganjule

Bayelemabaga: Creating Resources for Bambara NLP

The RTR Harmonic Domain in Two Dialects of Yorùbá

In this thesis, a process of vowel harmony is explored in two dialects of Yorùbá where the tongue-

Niger Volta LTI: Yorùbá language training text for NLP, ASR and TTS tasks

Yorùbá language training text for NLP, ASR and TTS tasks