Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Insights into Low-Resource Language Modelling: Improving Model Performances for South African Languages

Domain:

natural language processing

Record type:

papermodel
Creator:
Visser, RuanGrobler, TriekoDunaiski, Marcel
Publisher:
Journal of Universal Computer Science
Host:avatar
To address the gap in natural language processing for Southern African languages, our paper presents an in-depth analysis of language model development under resource-constrained conditions. We investigate the interplay between model size, pretraining objectives, and multilingual dataset composition in the context of low-resource languages such as Zulu and Xhosa. In our approach, we initially pretrain language models from scratch on specific low-resource languages using a variety of model configurations, and incrementally add related languages to explore the effect of additional languages on the performance of these models. We demonstrate that smaller data volumes can be effectively leveraged, and that the choice of pretraining objective and multilingual dataset composition significantly influences model performance. Our monolingual and multilingual models, exhibit competitive, and in some cases superior, performance compared to established multilingual models such as XLM-R-base and AfroXLM-R-base.

Visit

doi.org

Tasks

language modeling

Languages

Xhosa

Tags

Language ModellingLow-Resource LanguagesTransformersMultilingualPretraining

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Low-Resource Language Modelling of South African LanguagesImproving transformer model translation for low resource South African languages using BERTEnhancing African low-resource languages: Swahili data for language modellingInkubaLM: A small language model for low-resource African languagesUnsupervised Language Model Adaptation for Low-Resource LanguagesContinual-learning for Modelling Low-Resource Languages from Large Language Models

Low-Resource Language Modelling of South African Languages

Language models are the foundation of current neural network-based models for natural language under

Improving transformer model translation for low resource South African languages using BERT

Enhancing African low-resource languages: Swahili data for language modelling

Language modelling using neural networks requires adequate data to guarantee quality word representation which is important for natural language processing (NLP) tasks. However, African languages, Swahili in particular, have been disadvantaged and most of them are

InkubaLM: A small language model for low-resource African languages

High-resource language models often fall short in the African context, where there is a critical nee

Unsupervised Language Model Adaptation for Low-Resource Languages

This paper introduces a two-way neural machine translation system from Bengali to English and vice v

Continual-learning for Modelling Low-Resource Languages from Large Language Models

Modelling a language model for a multi-lingual scenario includes several potential challenges, among