Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

Domaine:

natural language processing

Type de record:

paper
Créateur:
MokMarivate, VukosiMunNet
Éditeur:
arXiv
Hôte:avatar
Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyond matched evaluation. Results revealed clear interactions between architecture and language, with W2V-BERT outperforming Whisper on Nguni languages by 3 to 4 WER points whilst Whisper performed better on Sotho-Tswana languages. W2V-BERT with tone conditioning reached 28.41% average WER across datasets and 23.79% on Xitsonga transfer. No single model suited all 6 languages, so deployment should pair model selection per language with validation across corpora.

Visit

doi.orgarxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

BirwaNgwoSetswanaTsonga

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Improved Meta Learning for Low Resource Speech RecognitionSMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech RecognitionRobust speech recognition for low-resource languagesMUST: A Multilingual Student-Teacher Learning approach for low-resource speech recognitionText-To-Speech Data Augmentation for Low Resource Speech RecognitionDEFI-COLaF/Speech-Recognition-for-Low-Resource-Languages

Improved Meta Learning for Low Resource Speech Recognition

We propose a new meta learning based framework for low resource speech recognition that improves the

SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition

Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource langu

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T

MUST: A Multilingual Student-Teacher Learning approach for low-resource speech recognition

Student-teacher learning or knowledge distillation (KD) has been previously used to address data sca

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r

DEFI-COLaF/Speech-Recognition-for-Low-Resource-Languages

# Speech-Recognition-for-Low-Resource-Languages This repository contains code to fine-tune Whisper