Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AFRILANGTUTOR: Advancing Language Tutoring and Culture Education in Low-Resource Languages with Large Language Models

Domaine:

natural language processingeducation

Type de record:

paperdatasetmodelsoftware
Créateur:
BelNahAziMon
Hôte:avatar
How can language learning systems be developed for languages that lack sufficient training resources? This challenge is increasingly faced by developers across the African continent who aim to build AI systems capable of understanding and responding in local languages. To address this gap, we introduce AFRILANGDICT, a collection of 194.7K African language-English dictionary entries designed as seed resources for generating language-learning materials, enabling us to automatically construct large-scale, diverse, and verifiable student-tutor question-answer interactions suitable for training AI-assisted language tutors. Using AFRILANGDICT, we build AFRILANGEDU, a dataset of 78.9K multi-turn training examples for Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). Using AFRILANGEDU, we train language tutoring models collectively referred to as AFRILANGTUTOR. We fine-tune two multilingual LLMs: Llama-3-8B-IT and Gemma-3-12B-IT on AFRILANGEDU across 10 African languages and evaluate their performance. Our results show that models trained on AFRILANGEDU consistently outperform their base counterparts, and combining SFT and DPO yields substantial improvements, with gains ranging from 1.8% to 15.5% under LLM-as-a-judge evaluations across four criteria. To facilitate further research on low-resource languages, all resources are available at huggingface.co.

Visit

arxiv.org

Tags

Computation and Language

Similaires

Interactive Machine Translation with Large Language Models for Low-resource LanguagesAdaptive and Efficient Large Language Models for Low-Resource African LanguagesZero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource LanguagesNusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language ModelsContinual-learning for Modelling Low-Resource Languages from Large Language ModelsEmploying large language models in Swahili, a low-resource language

Interactive Machine Translation with Large Language Models for Low-resource Languages

Large language models (LLM) have been applied to machine translation with notable success. However,

Adaptive and Efficient Large Language Models for Low-Resource African Languages

PAIDeF SuperAI 2025 Conference

Adaptive and Efficient Large Language Mod

Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Large language models (LLMs) have shown impressive zero-shot capabilities in various document rerank

NusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language Models

Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-res

Continual-learning for Modelling Low-Resource Languages from Large Language Models

Modelling a language model for a multi-lingual scenario includes several potential challenges, among

Employing large language models in Swahili, a low-resource language