Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models

Domain:

natural language processing

Record type:

paper
Creator:
ShaMurXia
Host:avatar
Although the multilingual capability of LLMs offers new opportunities to overcome the language barrier, do these capabilities translate into real-life scenarios where linguistic divide and knowledge conflicts between multilingual sources are known occurrences? In this paper, we studied LLM's linguistic preference in a cross-language RAG-based information search setting. We found that LLMs displayed systemic bias towards information in the same language as the query language in both document retrieval and answer generation. Furthermore, in scenarios where no information is in the language of the query, LLMs prefer documents in high-resource languages during generation, potentially reinforcing the dominant views. Such bias exists for both factual and opinion-based queries. Our results highlight the linguistic divide within multilingual LLMs in information search systems. The seemingly beneficial multilingual capability of LLMs may backfire on information parity by reinforcing language-specific information cocoons or filter bubbles further marginalizing low-resource views. NAACL 2025

Visit

arxiv.org

Tags

Computation and LanguageArtificial IntelligenceInformation Retrieval

Similar

Fairness in Multilingual Large Language Models: Addressing the Language Disparity Gap in AI SystemsQuantifying Language Disparities in Multilingual Large Language ModelsBridging language gaps in multilingual large language modelsIsolating Culture Neurons in Multilingual Large Language ModelsMultilingual Emotion Neurons in Large Audio-Language ModelsM-IFEval: On Multilingual Instruction-Following Capability of Large Language Models

Fairness in Multilingual Large Language Models: Addressing the Language Disparity Gap in AI Systems

Current Large Language Models (LLMs) exhibit significant performance disparities across languages, w

Quantifying Language Disparities in Multilingual Large Language Models

Results reported in large-scale multilingual evaluations are often fragmented and confounded by fact

Bridging language gaps in multilingual large language models

Large language models (LLMs) have revolutionized natural language processing, yet significant perfor

Isolating Culture Neurons in Multilingual Large Language Models

Language and culture are deeply intertwined, yet it has been unclear how and where multilingual larg

Multilingual Emotion Neurons in Large Audio-Language Models

Emotion is central to human communication, and its expression varies across languages. Large audio-l

M-IFEval: On Multilingual Instruction-Following Capability of Large Language Models

Instruction-following capability has become a major ability to be evaluated for Large Language Model