Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Corpus of multilingual code-switched soap opera speech

Domaine:

natural language processing

Type de record:

dataset
Créateur:
van der Westhuizen, EwaldNiesler, Thomas
Éditeur:
Stellenbosch University
Hôte:avatar
The corpus comprises 26.9 hours of annotated multilingual speech that contains examples of code-switching in isiZulu, isiXhosa, Setswana, Sesotho and English. The speech was obtained from South African soap operas. Code-switching between English and one of the Bantu languages is by far most prevalent in the data. Although not very common, switches between the Bantu languages themselves also occur. An initial attempt to align the audio extracted from soap opera episodes with the corresponding scripts revealed that actors very often perform ad lib. The speech and the examples of code-switching it contains can therefore be considered to be spontaneous.

Visit

hdl.handle.net

Tasks

code switchingspeech processing

Languages

SetswanaSotho, SouthernXhosaZulu

Tags

code-switching, spontaneous speech, South African languages, isiZulu, isiXhosa, Setswana, Sesotho

Licenses

Research only.