Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Afrikaans-English code-switched sentences with language identification and part-of-speech tags

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
MicKayVuk
Hôte:avatar

As highlighted in recent surveys, one of the biggest barriers to progress in code-switching (CS) research is the limited availability—both in quantity and quality—of annotated code-switched text (Doğruöz et al., 2023; Mondal et al., 2022). Winata et al. (2022) further show that this issue is particularly acute in the South African context. The survey reports that only a small body of CS research exists for South African languages and that data resources remain limited or absent. Notably, for Afrikaans–English specifically, there are no publicly available code-switched datasets.

Traditionally, researchers have relied on naturally occurring CS data from social media, speech recordings and manual or automated transcriptions. While these sources offer authenticity and sociolinguistic richness, they each introduce challenges such as ethical and privacy concerns, noise and inconsistency, costly annotation processes, and domain imbalance. Parallel corpora and substitutive generation methods also offer avenues for artificial CS creation, but for Afrikaans–English these corpora are either highly domain-specific or overly general and do not necessarily reflect naturally occurring switching patterns.

Motivated by these limitations, the thesis explores an alternative and increasingly relevant solution: generating synthetic CS data using multilingual large language models (LLMs). A controlled prompting framework is developed using GPT-4o and Gemini 2.0 Flash to produce aligned Afrikaans, English, correct CS and incorrect CS sentences (constituting a sentence set) across diverse topics.

In addition to addressing the data scarcity gap, evaluating the quality of synthetically generated sentences remains a challenge. the thesis further aims to develop quality evaluation strategies for synthetic data to be used in language learning applications. For this purpose, word-level language identification (LID) and part-of-speech (POS) tagging are required. Tagging was done using a combination of existing taggers, LLMs and human annotation.

The final data set used for training and testing consists of 5750 sets unvalidated sets of sentences and 624 validated sets of sentences, both with LID and POS tags. The validated set is made up of 464 grammar-based sentence sets (sentences that were specifically generated to add diversity to the types of incorrect CS sentences), 60 human-validated sentence sets and 100 gold-standard sets of sentences used for model testing.

Visit

figshare.com

Tasks

code switchinglanguage identificationpart of speech tagging

Languages

Afrikaans

Tags

Knowledge representation and reasoningNatural language processingAfrican languagesOther language, communication and culture not elsewhere classifiedAfrikaans-English code-switchingGPTGeminiLarge language models (LLMs)Natural language processingQuality evaluation+3

Licenses

CC BY-SA 4.0

Similaires

Large language models (LLMs)-generated Afrikaans-English code-switched dataIntegration of Phonotactic Features for Language Identification on Code-Switched SpeechChichewa-English Code-Switched Speech DatasetEnglish-IsiZulu Code-Switched Speech Recognition DatasetGhana English-Twi Code-switched Speech CorpusEffects of Language Modelling for Sepedi-English Code-Switched Speech in Automatic Speech Recognition System

Large language models (LLMs)-generated Afrikaans-English code-switched data

As highlighted in recent surveys, one of the biggest barriers to progress in code-switc

Integration of Phonotactic Features for Language Identification on Code-Switched Speech

Integration of Phonotactic Features for Language Identification on Code-Switched

Chichewa-English Code-Switched Speech Dataset

A speech dataset containing 247 audio recordings of Chichewa-English code-switched phrases. Code-swi

English-IsiZulu Code-Switched Speech Recognition Dataset

Dataset for semi-supervised acoustic and language model training for English-isiZulu code-switched s

Ghana English-Twi Code-switched Speech Corpus

Gold-standard English-Twi code-switched speech with transcripts and speaker meta

Effects of Language Modelling for Sepedi-English Code-Switched Speech in Automatic Speech Recognition System