Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions

Domain:

natural language processing

Record type:

paperdataset
Creator:
KökThaImaÜst
Host:avatar
Instruction tuning enhances large language models (LLMs) by aligning them with human preferences across diverse tasks. Traditional approaches to create instruction tuning datasets face serious challenges for low-resource languages due to their dependence on data annotation. This work introduces a novel method, Multilingual Reverse Instructions (MURI), which generates high-quality instruction tuning datasets for low-resource languages without requiring human annotators or pre-existing multilingual models. Utilizing reverse instructions and a translation pipeline, MURI produces instruction-output pairs from existing human-written texts in low-resource languages. This method ensures cultural relevance and diversity by sourcing texts from different native domains and applying filters to eliminate inappropriate content. Our dataset, MURI-IT, includes more than 2 million instruction-output pairs across 200 languages. Evaluation by native speakers and fine-tuning experiments with mT5 models demonstrate the approach's effectiveness for both NLU and open-ended generation. We publicly release datasets and models at github.com.

Visit

arxiv.org

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similar

InstructAlign: High-and-Low Resource Language Alignment via Continual Crosslingual Instruction TuningMulti-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMsTuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource LanguagesInstructAlign: High-and-Low Resource Language Alignment via Continual CrosslingualInstruction TuningX-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual InstructionsFine Tuning Methods for Low-resource Languages

InstructAlign: High-and-Low Resource Language Alignment via Continual Crosslingual Instruction Tuning

Large language models (LLMs) that are tuned with instructions have demonstrated remarkable capabilit

Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs

Audio large language models (LLMs) enable unified speech understanding and generation, but adapting

Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages

This article introduces contrastive alignment instructions (AlignInstruct) to address two challenges

InstructAlign: High-and-Low Resource Language Alignment via Continual CrosslingualInstruction Tuning

Large language models (LLMs) that are tuned with instructions have demonstrated remarkable capabilit

X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions

Large language models respond well in high-resource languages like English but struggle in low-resou

Fine Tuning Methods for Low-resource Languages

The rise of Large Language Models has not been inclusive of all cultures. The models are mostly trai