Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Low-Resource Language Modelling of South African Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
MesHayShaBuys, Jan
Éditeur:
arXiv
Hôte:avatar
Language models are the foundation of current neural network-based models for natural language understanding and generation. However, research on the intrinsic performance of language models on African languages has been extremely limited, which is made more challenging by the lack of large or standardised training and evaluation sets that exist for English and other high-resource languages. In this paper, we evaluate the performance of open-vocabulary language models on low-resource South African languages, using byte-pair encoding to handle the rich morphology of these languages. We evaluate different variants of n-gram models, feedforward neural networks, recurrent neural networks (RNNs), and Transformers on small-scale datasets. Overall, well-regularized RNNs give the best performance across two isiZulu and one Sepedi datasets. Multilingual training further improves performance on these datasets. We hope that this research will open new avenues for research into multilingual and low-resource language modelling for African languages. AfricaNLP workshop at EACL 2021

Visit

doi.orgarxiv.org

Tasks

language modeling

Languages

Sotho, NorthernZulu

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Insights into Low-Resource Language Modelling: Improving Model Performances for South African LanguagesEnhancing African low-resource languages: Swahili data for language modellingContinual-learning for Modelling Low-Resource Languages from Large Language ModelsInkubaLM: A small language model for low-resource African languagesRonaldKato/Luganda-Language-Analysis-Framework-for-Low-Resource-African-LanguagesOvercoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review

Insights into Low-Resource Language Modelling: Improving Model Performances for South African Languages

To address the gap in natural language processing for Southern African languages, our paper presents

Enhancing African low-resource languages: Swahili data for language modelling

Language modelling using neural networks requires adequate data to guarantee quality word representation which is important for natural language processing (NLP) tasks. However, African languages, Swahili in particular, have been disadvantaged and most of them are

Continual-learning for Modelling Low-Resource Languages from Large Language Models

Modelling a language model for a multi-lingual scenario includes several potential challenges, among

InkubaLM: A small language model for low-resource African languages

High-resource language models often fall short in the African context, where there is a critical nee

RonaldKato/Luganda-Language-Analysis-Framework-for-Low-Resource-African-Languages

This repository implements a complete NLP pipeline for analyzing Luganda, a Bantu language spoken by

Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review

Generative language modelling has surged in popularity with the emergence of services such as ChatGP