Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced Languages

Domain:

natural language processing

Record type:

paper
This paper reports on the semi-supervised development of acoustic and language models for under-resourced, code-switched speech in five South African languages. Two approaches are considered. The first constructs four separate bilingual automatic speech recognisers (ASRs) corresponding to four different language pairs between which speakers switch frequently. The second uses a single, unified, five-lingual ASR system that represents all the languages (English, isiZulu, isiXhosa, Setswana and Sesotho). We evaluate the effectiveness of these two approaches when used to add additional data to our extremely sparse training sets. Results indicate that batch-wise semi-supervised training yields better results than a non-batch-wise approach. Furthermore, while the separate bilingual systems achieved better recognition performance than the unified system, they benefited more from pseudolabels generated by the five-lingual system than from those generated by the bilingual systems.

Visit

aclanthology.orgarxiv.org

Tasks

speech processinglanguage modelingautomatic speech recognitioncode switching

Languages

SetswanaSotho, SouthernXhosaZulu

Tags

acl

Similar

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced LanguagesMultilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched SpeechSemi-supervised acoustic model training for five-lingual code-switched ASRFine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched SpeechSemi-supervised acoustic and language model training for English-isiZulu code-switched speech recognitionMultilingual training set selection for ASR in under-resourced Malian languages

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages

In this work, we explore the benefits of using multilingual bottleneck features (mBNF) in acoustic m

Multilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched Speech

Semi-supervised acoustic model training for five-lingual code-switched ASR

This paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS)

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguis

Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition

We present an analysis of semi-supervised acoustic and language model training for English-isiZulu c

Multilingual training set selection for ASR in under-resourced Malian languages

We present first speech recognition systems for the two severely under-resourced Malian languages Ba