Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Semi-supervised acoustic model training for five-lingual code-switched ASR

Domain:

natural language processing

Record type:

papermodel
Creator:
BisYılDe van der Westhuizen, Ewald
Publisher:
arXiv
Host:avatar
This paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS) speech in multiple South African languages. We consider two approaches. The first constructs separate bilingual acoustic models corresponding to language pairs (English-isiZulu, English-isiXhosa, English-Setswana and English-Sesotho). The second constructs a single unified five-lingual acoustic model representing all the languages (English, isiZulu, isiXhosa, Setswana and Sesotho). For these two approaches we consider the effectiveness of semi-supervised training to increase the size of the very sparse acoustic training sets. Using approximately 11 hours of untranscribed speech, we show that both approaches benefit from semi-supervised training. The bilingual TDNN-F acoustic models also benefit from the addition of CNN layers (CNN-TDNN-F), while the five-lingual system does not show any significant improvement. Furthermore, because English is common to all language pairs in our data, it dominates when training a unified language model, leading to improved English ASR performance at the expense of the other languages. Nevertheless, the five-lingual model offers flexibility because it can process more than two languages simultaneously, and is therefore an attractive option as an automatic transcription system in a semi-supervised training pipeline. Accepted for publication at Interspeech 2019

Visit

doi.orgarxiv.org

Tasks

automatic speech recognitioncode switchingspeech processing

Languages

SetswanaSotho, SouthernXhosaZulu

Tags

Computation and Language (cs.CL)Sound (cs.SD)Audio and Speech Processing (eess.AS)FOS: Computer and information sciencesFOS: Computer and information sciencesFOS: Electrical engineering, electronic engineering, information engineeringFOS: Electrical engineering, electronic engineering, information engineering

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/