Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages

Domain:

natural language processing

Record type:

paperdatasetmodelsoftware
Creator:
Li,ZhaLi,Wan
Host:avatar
Large language models (LLMs) continue to struggle with low-resource languages, primarily due to limited training data, translation noise, and unstable cross-lingual alignment. To address these challenges, we propose LiRA (Linguistic Robust Anchoring for LLMs)-a plug-and-play framework that requires only lightweight fine-tuning on top of existing pretrained backbones. LiRA jointly optimizes representation stability and cross-lingual semantic consistency by combining two key components: Arca (Anchored Representation Composition Architecture), which aligns low-resource inputs to a shared English semantic space through anchor-based alignment and collaborative encoding; and LaSR (Language-coupled Semantic Reasoner), a lightweight, language-aware head that enforces consistency regularization for unified cross-lingual understanding, retrieval, and reasoning. We theoretically show that under controlled anchoring error and translation-induced bias, LiRA guarantees bounded representation deviation and stable downstream performance under local Lipschitz continuity. To facilitate research, we release a new multilingual product retrieval dataset covering five Southeast Asian and two South Asian languages. Extensive experiments across diverse low-resource benchmarks demonstrate consistent improvements in retrieval, ranking, question answering, and reasoning tasks. Code will be publicly available on GitHub, and the dataset will be hosted on Hugging Face. Accepted by ICML 2026

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similar

Multilingual jailbreaking of LLMs using low-resource languagesMultilingual and Multimodal LLMs in the Wild: Building for Low-Resource LanguagesRethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languageschinmayjainnnn/LLMs-for-Translation-of-Low-Resource-LanguagesToward robust representation for low-resource automatic speech recognitionRobust speech recognition for low-resource languages

Multilingual jailbreaking of LLMs using low-resource languages

Large Language Models (LLMs) remain vulnerable to jailbreak attempts that circumvent safety guardrai

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

Multimodal LLMs are evolving from vision-language to tri-modality that see, hear, and read, yet pipe

Rethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languages

Realignment is a promising strategy to improve cross-lingual transfer in multilingual language model

chinmayjainnnn/LLMs-for-Translation-of-Low-Resource-Languages

Machine translation from assamese to english and vice versa using state of the art LLM's # Hindi-En

Toward robust representation for low-resource automatic speech recognition

Vers une représentation robuste pour la reconnaissance automatique de la parole des langues peu doté

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T