Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Speech Language Models for Under-Represented Languages: Insights from Wolof

Domain:

natural language processing

Record type:

papermodelsoftware
Creator:
Sy,DouCerIll
Editor:
LabNat
Publisher:
CCSD
Host:avatar
We present our journey in training a speech language model for Wolof, an underrepresented language spoken in West Africa, and share key insights. We first emphasize the importance of collecting large-scale, spontaneous, high-quality unsupervised speech data, and show that continued pretraining HuBERT on this dataset outperforms both the base model and African-centric models on ASR. We then integrate this speech encoder into a Wolof LLM to train the first Speech LLM for this language, extending its capabilities to tasks such as speech translation. Furthermore, we explore training the Speech LLM to perform multi-step Chain-of-Thought before transcribing or translating. Our results show that the Speech LLM not only improves speech recognition but also performs well in speech translation. The models and the code will be openly shared.

Visit

hal.science

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Languages

Wolof

Tags

Speech Representation, HuBERT, Speech Recognition, Speech Translation, Speech Language Models[INFO]Computer Science [cs]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

Automatic Spell Checker and Correction for Under-represented Spoken Languages: Case Study on WolofImproving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and TextSMOL: Professionally translated parallel data for 115 under-represented languagesMultilingual Rag Agents For Localized Knowledge: Adaptive Indexing For Under-Represented LanguagesAutomatic speech recognition for less-represented languages Transcription automatique de langues peu dotéesA Morphologically-Aware Dictionary-based Data Augmentation Technique for Machine Translation of Under-Represented Languages

Automatic Spell Checker and Correction for Under-represented Spoken Languages: Case Study on Wolof

This paper presents a spell checker and correction tool specifically designed for Wolof, an under-re

Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text

Whisper and other large-scale automatic speech recognition models have made significant progress in

SMOL: Professionally translated parallel data for 115 under-represented languages

We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine tr

Multilingual Rag Agents For Localized Knowledge: Adaptive Indexing For Under-Represented Languages

Abstract The democratization of information through Retrieval-Augmented Generation

Automatic speech recognition for less-represented languages Transcription automatique de langues peu dotées

With the development of technologies operating in a multilingual context, portability

A Morphologically-Aware Dictionary-based Data Augmentation Technique for Machine Translation of Under-Represented Languages

The availability of parallel texts is crucial to the performance of machine translation models. Howe