Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages

Domain:

natural language processing

Record type:

paper
Creator:
AlaHerAbdKla
Publisher:
arXiv
Host:avatar
Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space models such as Mamba. Although prior work has evaluated these architectures in multiple languages, their effectiveness in African languages remains underexplored. In this work, we evaluate Mamba for ASR on seven South African languages. In monolingual experiments, each model is trained on 50 hours of speech per language, and we compare Mamba to a Conformer baseline of similar parameter scale. Mamba achieves similar recognition accuracy to Conformer while using fewer computational resources and training faster. We further evaluate generalization in this setting and find that both models struggle to generalize to speech that is much longer than what they were trained on. We then study multilingual ASR using Mamba models, where the baseline is pooling all languages together. On top of this, we tested three extensions: training with language-family information by adding both language and language-family embeddings as biases to the downsampled acoustic representations, and multitask learning with a CTC ASR objective and a language identification (LID) head. We find that multilingual training consistently improves performance over monolingual training. However, adding explicit language information does not improve in-domain performance but does improve cross-corpus robustness. We conducted ablation studies in low-resource multilingual settings using 5-hour and 10-hour per-language training data, where we observed gains from using language embeddings and further demonstrated that removing or altering them hurt model performance. Lastly, we analysed these embeddings and find that they do not capture linguistic similarity in a typological sense, but instead act as task-specific control vectors. under review

Visit

doi.orgarxiv.org

Tasks

automatic speech recognitionlanguage identificationspeech processing

Languages

VunjoZimba

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-sa/4.0/legalcode

Similar

Linguistically Informed Evaluation of Multilingual ASR for African LanguagesMonolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Surveyfurqanx/ASR-for-african-languagesDistilling Monolingual Models from Large Multilingual TransformersMultilingual training set selection for ASR in under-resourced Malian languagesMultilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Linguistically Informed Evaluation of Multilingual ASR for African Languages

Word Error Rate (WER) mischaracterizes ASR models' performance for African languages by combining ph

Monolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a signi

furqanx/ASR-for-african-languages

# Automatic Speech Recognition for African Languages - Platform: Zindi - Competition link: https://

Distilling Monolingual Models from Large Multilingual Transformers

Although language modeling has been trending upwards steadily, models available for low-resourced la

Multilingual training set selection for ASR in under-resourced Malian languages

We present first speech recognition systems for the two severely under-resourced Malian languages Ba

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Poster presented at the Deep Learning Indaba 2023 by Samuel Rutunda