Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG

Domain:

natural language processing

Record type:

paper
Creator:
Ki,CarMcNKha
Host:avatar
Multilingual Retrieval-Augmented Generation (mRAG) systems enable language models to answer knowledge-intensive queries with citation-supported responses across languages. While such systems have been proposed, an open questions is whether the mixture of different document languages impacts generation and citation in unintended ways. To investigate, we introduce a controlled methodology using model internals to measure language preference while holding other factors such as document relevance constant. Across eight languages and six open-weight models, we find that models preferentially cite English sources when queries are in English, with this bias amplified for lower-resource languages and for documents positioned mid-context. Crucially, we find that models sometimes trade-off document relevance for language preference, indicating that citation choices are not always driven by informativeness alone. Our findings shed light on how language models leverage multilingual context and influence citation behavior. 33 pages, 20 figures

Visit

arxiv.org

Tags

Computation and Language

Similar

Language-Preference-Based Re-ranking for Multilingual Swahili Information RetrievalCORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAGKingsleyTechie/MARAG-Adaptive-Multilingual-RAG-for-Low-Resource-LanguagesMultilingual Rag Agents For Localized Knowledge: Adaptive Indexing For Under-Represented Languageswill-genius/Multilingual-RAG-based-swahili-dictionaryAre Multilingual Language Models an Off-ramp for Under-resourced Languages? Will we arrive at Digital Language Equality in Europe in 2030?

Language-Preference-Based Re-ranking for Multilingual Swahili Information Retrieval

CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG

Multilingual retrieval-augmented generation (mRAG) is often implemented within a fixed retrieval spa

KingsleyTechie/MARAG-Adaptive-Multilingual-RAG-for-Low-Resource-Languages

Official implementation of "Multilingual Adaptive RAG for Localized Knowledge" (ACM TOIS submission)

Multilingual Rag Agents For Localized Knowledge: Adaptive Indexing For Under-Represented Languages

Abstract The democratization of information through Retrieval-Augmented Generation

will-genius/Multilingual-RAG-based-swahili-dictionary

# Multilingual-RAG-based-swahili-dictionary

Are Multilingual Language Models an Off-ramp for Under-resourced Languages? Will we arrive at Digital Language Equality in Europe in 2030?

Large language models (LLMs) demonstrate unprecedented capabilities and define the state of the art