Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multilingual training set selection for ASR in under-resourced Malian languages

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
van der Westhuizen, EwaldPadNiesler, Thomas
Éditeur:
arXiv
Hôte:avatar
We present first speech recognition systems for the two severely under-resourced Malian languages Bambara and Maasina Fulfulde. These systems will be used by the United Nations as part of a monitoring system to inform and support humanitarian programmes in rural Africa. We have compiled datasets in Bambara and Maasina Fulfulde, but since these are very small, we take advantage of six similarly under-resourced datasets in other languages for multilingual training. We focus specifically on the best composition of the multilingual pool of speech data for multilingual training. We find that, although maximising the training pool by including all six additional languages provides improved speech recognition in both target languages, substantially better performance can be achieved by a more judicious choice. Our experiments show that the addition of just one language provides best performance. For Bambara, this additional language is Maasina Fulfulde, and its introduction leads to a relative word error rate reduction of 6.7%, as opposed to a 2.4% relative reduction achieved when pooling all six additional languages. For the case of Maasina Fulfulde, best performance was achieved when adding only Luganda, leading to a relative word error rate improvement of 9.4% as opposed to a 3.9% relative improvement when pooling all six languages. We conclude that careful selection of the out-of-language data is worthwhile for multilingual training even in highly under-resourced settings, and that the general assumption that more data is better does not always hold. 12 pages, 4 figures, Accepted for presentation at SPECOM 2021

Visit

doi.orgarxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

BamanankanFulfulde, AdamawaFulfulde, BorguFulfulde, Central-Eastern NigerFulfulde, MaasinaFulfulde, NigerianFulfulde, Western NigerGandaLame

Tags

Audio and Speech Processing (eess.AS)FOS: Electrical engineering, electronic engineering, information engineeringFOS: Electrical engineering, electronic engineering, information engineering

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similaires

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced LanguagesSemi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced LanguagesThe limitations of data perturbation for ASR of learner data in under-resourced languagesImproved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic LandmarksMultilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched SpeechDatasheets for Under-resourced Languages: An Example

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages

In this work, we explore the benefits of using multilingual bottleneck features (mBNF) in acoustic m

Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced Languages

This paper reports on the semi-supervised development of acoustic and language models for under-resourced, code-switched speech in five South African languages. Two approaches are considered. The first constructs four separate bilingual automatic speech recognisers

The limitations of data perturbation for ASR of learner data in under-resourced languages

Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks

Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V

Multilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched Speech

Datasheets for Under-resourced Languages: An Example

The datasheet provides an example of how to use the Datasheet standard for describing and sharing un