Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
KeiDiarra, SebastienHomDia
Host:avatar
Effective text generation and chat interfaces for low-resource languages (LRLs) remain a challenge for state-of-the-art large language models (LLMs) to support. This is mainly due to the difficulty of curating high-quality instruction datasets for LRLs, a limitation prevalent in the languages spoken across the African continent and other regions. Current approaches, such as automated translation and synthetic data generation, frequently yield outputs that lack fluency or even orthographic consistency. In this paper, we introduce InstructLR, a novel framework designed to generate high-quality instruction datasets for LRLs. Our approach integrates LLM-driven text generation with a dual-layer quality filtering mechanism: an automated filtering layer based on retrieval-augmented-generation (RAG)-based n-shot prompting, and a human-in-the-loop validation layer. Drawing inspiration from benchmarks such as MMLU in task definition, InstructLR has facilitated the creation of three multi-domain instruction benchmarks: ZarmaInstruct-50k, BambaraInstruct-50k, and FulfuldeInstruct-50k.

Visit

arxiv.org

Tags

Machine Learning

Similar

Automatic speech recognition for under-resourced languages: A surveyA primer on getting neologisms from foreign languages to under-resourced languagesDatasheets for Under-resourced Languages: An ExampleEnd-to-End Text-To-Speech synthesis for under resourced South African languagesShort Text Language Identification for Under Resourced LanguagesTranslation-Based Dictionary Alignment for Under-Resourced Bantu Languages

Automatic speech recognition for under-resourced languages: A survey

(Impact-F 1.28 estim. in 2012) International audience no abstract

A primer on getting neologisms from foreign languages to under-resourced languages

Mainly due to lack of support, most under-resourced languages have a reduced lexicon in most realms

Datasheets for Under-resourced Languages: An Example

The datasheet provides an example of how to use the Datasheet standard for describing and sharing un

End-to-End Text-To-Speech synthesis for under resourced South African languages

Short Text Language Identification for Under Resourced Languages

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text languag

Translation-Based Dictionary Alignment for Under-Resourced Bantu Languages

Despite a large number of active speakers, most Bantu languages can be considered as under- or less-