Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languages

Domain:

natural language processing

Record type:

papermodel
Creator:
RajS, WalRag
Host:avatar
Language model compression through knowledge distillation has emerged as a promising approach for deploying large language models in resource-constrained environments. However, existing methods often struggle to maintain performance when distilling multilingual models, especially for low-resource languages. In this paper, we present a novel hybrid distillation approach that combines traditional knowledge distillation with a simplified attention matching mechanism, specifically designed for multilingual contexts. Our method introduces an extremely compact student model architecture, significantly smaller than conventional multilingual models. We evaluate our approach on five African languages: Kinyarwanda, Swahili, Hausa, Igbo, and Yoruba. The distilled student model; AfroXLMR-Comet successfully captures both the output distribution and internal attention patterns of a larger teacher model (AfroXLMR-Large) while reducing the model size by over 85%. Experimental results demonstrate that our hybrid approach achieves competitive performance compared to the teacher model, maintaining an accuracy within 85% of the original model's performance while requiring substantially fewer computational resources. Our work provides a practical framework for deploying efficient multilingual models in resource-constrained environments, particularly benefiting applications involving African languages.

Visit

arxiv.org

Languages

HausaKinyarwandaSwahiliYoruba

Tags

Computation and LanguageArtificial IntelligenceInformation RetrievalMachine Learning

Similar

Extracting General-use Transformers for Low-resource Languages via Knowledge DistillationCross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource LanguagesAdapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via AdaptersMultilingual Knowledge Graphs and Low-Resource Languages: A ReviewOptimal Transport Distillation for Low-Resource Language Alignment in Multilingual RetrievalMultilingual Distillation Robustness in Low-Resource Cross-Lingual NER

Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation

In this paper, we propose the use of simple knowledge distillation to produce smaller and more effic

Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages

While impressive performance has been achieved on the task of Answer Sentence Selection (AS2) for En

Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters

This paper explores the integration of graph knowledge from linguistic ontologies into multilingual

Multilingual Knowledge Graphs and Low-Resource Languages: A Review

There is a lack of multilingual data to support applications in a large number of languages, especia

Optimal Transport Distillation for Low-Resource Language Alignment in Multilingual Retrieval

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Multilingual Distillation Robustness in Low-Resource Cross-Lingual NER

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident