Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Prompting Towards Alleviating Code-Switched Data Scarcity in Under-Resourced Languages with GPT as a Pivot

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
TerOlaMarivate, Vukosi
Hôte:avatar
Many multilingual communities, including numerous in Africa, frequently engage in code-switching during conversations. This behaviour stresses the need for natural language processing technologies adept at processing code-switched text. However, data scarcity, particularly in African languages, poses a significant challenge, as many are low-resourced and under-represented. In this study, we prompted GPT 3.5 to generate Afrikaans--English and Yoruba--English code-switched sentences, enhancing diversity using topic-keyword pairs, linguistic guidelines, and few-shot examples. Our findings indicate that the quality of generated sentences for languages using non-Latin scripts, like Yoruba, is considerably lower when compared with the high Afrikaans-English success rate. There is therefore a notable opportunity to refine prompting guidelines to yield sentences suitable for the fine-tuning of language models. We propose a framework for augmenting the diversity of synthetically generated code-switched data using GPT and propose leveraging this technology to mitigate data scarcity in low-resourced languages, underscoring the essential role of native speakers in this process. To be published in the Proceedings of SIGUL 2024: 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages

Visit

arxiv.org

Tasks

code switching

Languages

AfrikaansYoruba

Tags

Computation and Language

Similaires

Optimised Code-Switched Language Model Data Augmentation in Four Under-Resourced South African LanguagesCode-Switched Language Modelling Using a Code Predictive Lstm in Under-Resourced South African LanguagesMultilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced LanguagesSemi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced LanguagesImproving N-Best Rescoring in Under-Resourced Code-Switched Speech Recognition Using Pretraining and Data AugmentationMultilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched Speech

Optimised Code-Switched Language Model Data Augmentation in Four Under-Resourced South African Languages

Code-Switched Language Modelling Using a Code Predictive Lstm in Under-Resourced South African Languages

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages

In this work, we explore the benefits of using multilingual bottleneck features (mBNF) in acoustic m

Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced Languages

This paper reports on the semi-supervised development of acoustic and language models for under-resourced, code-switched speech in five South African languages. Two approaches are considered. The first constructs four separate bilingual automatic speech recognisers

Improving N-Best Rescoring in Under-Resourced Code-Switched Speech Recognition Using Pretraining and Data Augmentation

We present improvements in n-best rescoring of code-switched speech achieved by n-gram augmentation

Multilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched Speech