Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Dialectal Adaptation of Foundation Models for Low-Resource Speech Synthesis: A Case Study on Adamawa Fulfulde

Domaine:

natural language processing

Type de record:

modelpaper
Créateur:
RapOliMis
Éditeur:
Elsevier BV
Hôte:
The rapid evolution of neural speech synthesis has achieved near-human naturalness for high-resource languages, yet a profound digital divide persists for the linguistically diverse yet digitally underserved regions of Africa. Fulfulde, a major West and Central African macro language, exemplifies this disparity. While recent foundation models like Massively Multilingual Speech (MMS) provide zero-shot synthesis capabilities for Fulfulde, they often fail to capture the nuanced phonetic signatures of specific dialects, such as the Adamawa variety spoken in Cameroon. This study presents a robust framework for the dialectal adaptation of foundation models under extreme data scarcity. Rather than claiming a "first-of-its-kind" system, we demonstrate that targeted fine-tuning of a VITS-based architecture significantly outperforms the generalized zero-shot outputs of state-of-the-art multilingual models. Utilizing a meticulously curated 100-minute corpus and a specialized orthographic normalization protocol, our system achieves a Mean Opinion Score (MOS) of 4.25, as evaluated by a panel of 24 native speakers. The results indicate that dialect-specific fine-tuning is essential for accurately modeling glottalic airstream mechanisms (implosives) and vowel quantity contrasts, which are frequently marginalized in generalized multilingual frameworks. This work provides a rigorous methodology for localized linguistic inclusion in the era of foundation models.

Visit

doi.org

Tasks

text to speechspeech processing

Languages

Fulfulde, AdamawaFulfulde, BorguFulfulde, Central-Eastern NigerFulfulde, MaasinaFulfulde, NigerianFulfulde, Western Niger

Licenses

https://www.uspto.gov/ip-policy/copyright-policy/copyright-basics

Similaires

Multilingual Byte2Speech Models for Scalable Low-resource Speech SynthesisLow-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-StudyInvestigations on Speech Recognition Systems for Low-Resource Dialectal Arabic-English Code-Switching SpeechFoundation Models for Low-Resource Language Education (Vision Paper)Common Voice Scripted Speech 26.0 - Adamawa FulfuldeLarge Language Models Adaptation for Low-resource Languages: The Case for African Languages

Multilingual Byte2Speech Models for Scalable Low-resource Speech Synthesis

To scale neural speech synthesis to various real-world languages, we present a multilingual end-to-e

Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study

Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain

Investigations on Speech Recognition Systems for Low-Resource Dialectal Arabic-English Code-Switching Speech

Code-switching (CS), defined as the mixing of languages in conversations, has become a worldwide phe

Foundation Models for Low-Resource Language Education (Vision Paper)

Recent studies show that large language models (LLMs) are powerful tools for working with natural la

Common Voice Scripted Speech 26.0 - Adamawa Fulfulde

A collection of read speech recordings in Adamawa Fulfulde (fub).

Large Language Models Adaptation for Low-resource Languages: The Case for African Languages

David Ifeoluwa Adelani (Supervisor) Despite remarkable advances in Large Language Models (LLMs), Afr