Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition

Domain:

natural language processing

Record type:

paper
Creator:
BisDe Wet, F.vanNie
Host:avatar
We present an analysis of semi-supervised acoustic and language model training for English-isiZulu code-switched ASR using soap opera speech. Approximately 11 hours of untranscribed multilingual speech was transcribed automatically using four bilingual code-switching transcription systems operating in English-isiZulu, English-isiXhosa, English-Setswana and English-Sesotho. These transcriptions were incorporated into the acoustic and language model training sets. Results showed that the TDNN-F acoustic models benefit from the additional semi-supervised data and that even better performance could be achieved by including additional CNN layers. Using these CNN-TDNN-F acoustic models, a first iteration of semi-supervised training achieved an absolute mixed-language WER reduction of 3.4%, and a further 2.2% after a second iteration. Although the languages in the untranscribed data were unknown, the best results were obtained when all automatically transcribed data was used for training and not just the utterances classified as English-isiZulu. Despite reducing perplexity, the semi-supervised language model was not able to improve the ASR performance. 4th Code-Switch workshop, France

Visit

arxiv.org

Tasks

automatic speech recognitioncode switchingspeech processing

Languages

SetswanaSotho, SouthernXhosaZulu

Tags

Audio and Speech ProcessingMachine LearningSound

Similar

Semi-supervised acoustic model training for five-lingual code-switched ASREnglish-IsiZulu Code-Switched Speech Recognition DatasetImproved low-resource Somali speech recognition by semi-supervised acoustic and language model trainingMultilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched SpeechAutomatic Speech Recognition of English-isiZulu Code-switched Speech from South African Soap OperasComparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language Models

Semi-supervised acoustic model training for five-lingual code-switched ASR

This paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS)

English-IsiZulu Code-Switched Speech Recognition Dataset

Dataset for semi-supervised acoustic and language model training for English-isiZulu code-switched s

Improved low-resource Somali speech recognition by semi-supervised acoustic and language model training

We present improvements in automatic speech recognition (ASR) for Somali, a currently extremely under-resourced language. This forms part of a continuing United Nations (UN) effort to employ ASR-based keyword spotting systems to support humanitarian relief programm

Multilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched Speech

Automatic Speech Recognition of English-isiZulu Code-switched Speech from South African Soap Operas

Comparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language Models

International audience This paper investigates the potential of improving a hybrid au