Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages

Domaine:

natural language processing

Type de record:

paperdatasetmodelsoftware
Créateur:
Li,ZhaLi,Wan
Hôte:avatar
Large language models (LLMs) continue to struggle with low-resource languages, primarily due to limited training data, translation noise, and unstable cross-lingual alignment. To address these challenges, we propose LiRA (Linguistic Robust Anchoring for LLMs)-a plug-and-play framework that requires only lightweight fine-tuning on top of existing pretrained backbones. LiRA jointly optimizes representation stability and cross-lingual semantic consistency by combining two key components: Arca (Anchored Representation Composition Architecture), which aligns low-resource inputs to a shared English semantic space through anchor-based alignment and collaborative encoding; and LaSR (Language-coupled Semantic Reasoner), a lightweight, language-aware head that enforces consistency regularization for unified cross-lingual understanding, retrieval, and reasoning. We theoretically show that under controlled anchoring error and translation-induced bias, LiRA guarantees bounded representation deviation and stable downstream performance under local Lipschitz continuity. To facilitate research, we release a new multilingual product retrieval dataset covering five Southeast Asian and two South Asian languages. Extensive experiments across diverse low-resource benchmarks demonstrate consistent improvements in retrieval, ranking, question answering, and reasoning tasks. Code will be publicly available on GitHub, and the dataset will be hosted on Hugging Face. Accepted by ICML 2026

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similaires

Multilingual jailbreaking of LLMs using low-resource languagesMultilingual and Multimodal LLMs in the Wild: Building for Low-Resource LanguagesRethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languageschinmayjainnnn/LLMs-for-Translation-of-Low-Resource-LanguagesToward robust representation for low-resource automatic speech recognitionRobust speech recognition for low-resource languages

Multilingual jailbreaking of LLMs using low-resource languages

Large Language Models (LLMs) remain vulnerable to jailbreak attempts that circumvent safety guardrai

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

Multimodal LLMs are evolving from vision-language to tri-modality that see, hear, and read, yet pipe

Rethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languages

Realignment is a promising strategy to improve cross-lingual transfer in multilingual language model

chinmayjainnnn/LLMs-for-Translation-of-Low-Resource-Languages

Machine translation from assamese to english and vice versa using state of the art LLM's # Hindi-En

Toward robust representation for low-resource automatic speech recognition

Vers une représentation robuste pour la reconnaissance automatique de la parole des langues peu doté

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T