Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Acoustic and Tonal Modeling of the tpuri Language through a Multi-Modular Hybrid Approach

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
JulPatPauKol
Éditeur:
Eur
Hôte:
Automatic speech recognition (ASR) for tonal low-resource languages remains challenging due to the scarcity of labelled data and the need to model complex prosodic systems. This paper presents a hybrid multimodular ASR architecture for tpuri, a Mboum-Day Niger-Congo language spoken in Cameroon and Chad that exhibits contrastive lexical tone, vowel length and nasalisation. The system combines a self-supervised Wav2Vec 2.0 acoustic encoder with a tonal processing module based on YIN pitch estimation and STFTderived spectral features, and an adaptive fusion mechanism that integrates acoustic and tonal representations before decoding. We pretrain the acoustic encoder on 45 hours of read and spontaneous speech and finetune it on 19h35 of scripted speech. On the scripted test set, our best configuration reaches a word error rate (WER) of 10.4%, a phone error rate (PER) of 8.7% and a tone error rate (TER) of 6.1%. Ablation experiments show that removing the tonal module (+1.5 WER, +2.3 TER) or self-supervised pretraining (+3.4 WER) substantially degrades performance, while adaptive fusion and tone-aware data augmentation yield smaller but consistent gains. A fine-grained error analysis across tonal, grammatical, syllabic and morphological dimensions indicates that the architecture is particularly effective at modelling lexical tone and clause-level syntax, but still struggles with complex syllable structures and rich morphology. Overall, the results demonstrate that competitive ASR is attainable for under-resourced tonal languages such as tpuri by tightly coupling self-supervised acoustic modelling with explicit tonal representations, and provide a reusable blueprint for extending ASR to other Niger-Congo languages.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

MbumNzakambayTupuri

Licenses

https://creativecommons.org/licenses/by-nc-sa/4.0

Similaires

DNN acoustic modeling with modular multi-lingual feature extraction networksInvestigation of Various Hybrid Acoustic Modeling Units via a Multitask Learning and Deep Neural Network Technique for LVCSR of the Low-Resource Language, AmharicOn the ranking of variable length discords through a hybrid outlier detection approachTowards Learning to Speak and Hear Through Multi-Agent Communication over a Continuous Acoustic ChannelUsing different acoustic, lexical and language modeling units for ASR of an under-resourced language - AmharicModeling the cross-linguistic variations of tonal systems

DNN acoustic modeling with modular multi-lingual feature extraction networks

In this work, we propose several deep neural network architectures that are able to leverage data fr

Investigation of Various Hybrid Acoustic Modeling Units via a Multitask Learning and Deep Neural Network Technique for LVCSR of the Low-Resource Language, Amharic

On the ranking of variable length discords through a hybrid outlier detection approach

International audience In this paper we are interested in identifying insightful chan

Towards Learning to Speak and Hear Through Multi-Agent Communication over a Continuous Acoustic Channel

Towards Learning to Speak and Hear Through Multi-Agent Communication over a Continuous Acoustic Channel

Poster presented at the Deep Learning Indaba 2022 by Kevin Eloff

Using different acoustic, lexical and language modeling units for ASR of an under-resourced language - Amharic

State-of-the-art large vocabulary continuous speech recognition systems use mostly phone based acoustic models (AMs) and word based lexical and language models. However, phone based AMs are not efficient in modeling long-term temporal dependencies and the use of wo

Modeling the cross-linguistic variations of tonal systems

This study aims to simulate the cross-linguistic variations of tonal systems with low dimensional mo