Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
GurVykvanOst
Hôte:avatar
Low-resource languages (LRLs) face significant challenges in natural language processing (NLP) due to limited data. While current state-of-the-art large language models (LLMs) still struggle with LRLs, smaller multilingual models (mLMs) such as mBERT and XLM-R offer greater promise due to a better fit of their capacity to low training data sizes. This study systematically investigates parameter-efficient adapter-based methods for adapting mLMs to LRLs, evaluating three architectures: Sequential Bottleneck, Invertible Bottleneck, and Low-Rank Adaptation. Using unstructured text from GlotCC and structured knowledge from ConceptNet, we show that small adaptation datasets (e.g., up to 1 GB of free-text or a few MB of knowledge graph data) yield gains in intrinsic (masked language modeling) and extrinsic tasks (topic classification, sentiment analysis, and named entity recognition). We find that Sequential Bottleneck adapters excel in language modeling, while Invertible Bottleneck adapters slightly outperform other methods on downstream tasks due to better embedding alignment and larger parameter counts. Adapter-based methods match or outperform full fine-tuning while using far fewer parameters, and smaller mLMs prove more effective for LRLs than massive LLMs like LLaMA-3, GPT-4, and DeepSeek-R1-based distilled models. While adaptation improves performance, pre-training data size remains the dominant factor, especially for languages with extensive pre-training coverage. Pre-print

Visit

arxiv.org

Tags

Computation and Language

Similaires

Efficient multilingual and domain adaptation of language models under resource constraintsAdaptive and Efficient Large Language Models for Low-Resource African LanguagesSmall Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced LanguagesSmall Data? No Problem: Exploring the Viability of Multilingual Pretrained Language Models for Low-resourced LanguagesImpact of Intermediate-Task Training on Low-Resource Languages in Multilingual Language ModelsA Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

Efficient multilingual and domain adaptation of language models under resource constraints

Neural networks trained for language modeling, which is the task of finding the missing words in a g

Adaptive and Efficient Large Language Models for Low-Resource African Languages

PAIDeF SuperAI 2025 Conference

Adaptive and Efficient Large Language Mod

Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages

Pretrained multilingual language models have been shown to work well on many languages for a variety of downstream NLP tasks. However, these models are known to require a lot of training data. This consequently leaves out a huge percentage of the world{'}s language

Small Data? No Problem: Exploring the Viability of Multilingual Pretrained Language Models for Low-resourced Languages

Pretrained multilingual language models have been shown to work well on many languages for a variety

Impact of Intermediate-Task Training on Low-Resource Languages in Multilingual Language Models

Accuracy of English-language Question Answering (QA) systems has improved significantly in recent ye

A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

Multilingual short-text classification supports operational systems such as content moderation, cust