Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Domain:

natural language processing

Record type:

papermodel
Creator:
FroMorvanNiesler, Thomas
Host:avatar
Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguistic expertise. This is partly due to the large number of language combinations that may appear within and across utterances, which might require several annotators with different linguistic expertise to consider an utterance sequentially. This is time-consuming and costly. It would be useful if the spoken languages in an utterance and the boundaries thereof were known before annotation commences, to allow segments to be assigned to the relevant language experts in parallel. To address this, we investigate the development of a continuous multilingual language diarizer using fine-tuned speech representations extracted from a large pre-trained self-supervised architecture (WavLM). We experiment with a code-switched corpus consisting of five South African languages (isiZulu, isiXhosa, Setswana, Sesotho and English) and show substantial diarization error rate improvements for language families, language groups, and individual languages over baseline systems. Presented at SACAIR 2022

Visit

arxiv.org

Tasks

code switchinglanguage identificationspeech processing

Languages

SetswanaSotho, SouthernXhosaZulu

Tags

Audio and Speech ProcessingArtificial IntelligenceSound

Similar

An analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and EnglishSemi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced LanguagesMultilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitchingBenchmarking Self-Supervised Speech Models on Multilingual Nigerian SpeechSemi-supervised acoustic and language model training for English-isiZulu code-switched speech recognitionCorpus of multilingual code-switched soap opera speech

An analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and English

An analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and English

Poster presented at the Deep Learning Indaba 2022 by Tolúlọpẹ́ Ògúnrẹ̀mí

Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced Languages

This paper reports on the semi-supervised development of acoustic and language models for under-resourced, code-switched speech in five South African languages. Two approaches are considered. The first constructs four separate bilingual automatic speech recognisers

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching

While many speakers of low-resource languages regularly code-switch between their languages and othe

Benchmarking Self-Supervised Speech Models on Multilingual Nigerian Speech

Self-supervised speech models such as Whisper and wav2vec 2.0 have significantly advanced automatic

Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition

We present an analysis of semi-supervised acoustic and language model training for English-isiZulu c

Corpus of multilingual code-switched soap opera speech

The corpus comprises 26.9 hours of annotated multilingual speech that contains examples of code-swit