Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Improving Tone Recognition Performance using Wav2vec 2.0-Based Learned Representation in Yoruba, a Low-Resourced Language

Domaine:

natural language processing

Type de record:

paper
Créateur:
SaiNorPaulin, Melatagia YontaJea
Éditeur:
Ass
Hôte:
Many sub-Saharan African languages are categorized as tone languages, and for the most part, they are classified as low-resource languages due to the limited resources and tools available to process these languages. Identifying the tone associated with a syllable is therefore a key challenge for speech recognition in these languages. We propose models that automate the recognition of tones in continuous speech that can easily be incorporated into a speech recognition pipeline for these languages. We have investigated different neural architectures as well as several feature extraction algorithms in speech (FBs (Filter Banks), LEAF (Learnable Frontend), CS (Cestrogram), MFCC (Mel-Frequency Cepstral Coefficients)). In the context of low-resource languages, we also evaluated W2V (Wav2vec 2.0) models for this task. In this work, we use a public speech recognition dataset on Yoruba. As for the results, using the combination of features obtained from CS and FBs, we obtain a minimum TER (Tone Error Rate) of 19.54%, whereas for the evaluations of the models using W2V, we have a TER of 17.72%, demonstrating that the use of W2V provides better performance than the models used in the literature for tone identification on low-resource languages.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

Yoruba

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background