Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review

Domain:

natural language processing

Record type:

paper
Creator:
McGNik
Host:avatar
Generative language modelling has surged in popularity with the emergence of services such as ChatGPT and Google Gemini. While these models have demonstrated transformative potential in productivity and communication, they overwhelmingly cater to high-resource languages like English. This has amplified concerns over linguistic inequality in natural language processing (NLP). This paper presents the first systematic review focused specifically on strategies to address data scarcity in generative language modelling for low-resource languages (LRL). Drawing from 54 studies, we identify, categorise and evaluate technical approaches, including monolingual data augmentation, back-translation, multilingual training, and prompt engineering, across generative tasks. We also analyse trends in architecture choices, language family representation, and evaluation methods. Our findings highlight a strong reliance on transformer-based models, a concentration on a small subset of LRLs, and a lack of consistent evaluation across studies. We conclude with recommendations for extending these methods to a wider range of LRLs and outline open challenges in building equitable generative language systems. Ultimately, this review aims to support researchers and developers in building inclusive AI tools for underrepresented languages, a necessary step toward empowering LRL speakers and the preservation of linguistic diversity in a world increasingly shaped by large-scale language technologies. This work is currently under review. Please do not cite without permission

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and LanguageArtificial Intelligence

Similar

Enhancing African low-resource languages: Swahili data for language modellingComparing Methods for Overcoming Data Scarcity in Less-Resourced Natural Language ProcessingLLM Safety Alignment in Low-Resource Languages: A Systematic Literature ReviewOvercoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource LanguagesHybrid neural model for Hausa text auto completion: addressing data scarcity for a low-resource languageBabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context

Enhancing African low-resource languages: Swahili data for language modelling

Language modelling using neural networks requires adequate data to guarantee quality word representation which is important for natural language processing (NLP) tasks. However, African languages, Swahili in particular, have been disadvantaged and most of them are

Comparing Methods for Overcoming Data Scarcity in Less-Resourced Natural Language Processing

Natural language processing (NLP) tasks like named entity recognition (NER) and automatic text summa

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safet

Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages

Multilingual ASR models such as Whisper perform well on high-resource languages but exhibit substant

Hybrid neural model for Hausa text auto completion: addressing data scarcity for a low-resource language

The Hausa language, a major linguistic vehicle for over 70 million people in West Africa, remains cr

BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context

The BabyLM challenge called on participants to develop sample-efficient language models. Submissions