Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
ShaBenSinSil
Hôte:avatar
Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy. Although well-studied in English, its manifestation in other languages remains largely unexamined, leaving billions of non-English speakers potentially vulnerable to model-validated misinformation. We present the first large-scale, multi-model evaluation of cross-lingual sycophancy, benchmarking \textbf{six instruction-tuned models} across \textbf{1.1 million instances} spanning \textbf{38 languages} and \textbf{33 topic categories}. We identify a consistent resource-tier effect: sycophancy rates spike sharply in low-resource and zero-shot language settings. Critically, this degradation is topic-agnostic, as models fail uniformly across both benign and safety-critical prompts, offering no additional protection where it is most needed. We further identify tokenizer fertility as a structural driver of this alignment collapse. Collectively, our results demonstrate that prevailing alignment methodologies generalize poorly beyond high-resource languages, underscoring the urgent need for equitable multilingual safety techniques. 19 pages, 9 figures, 7 tables

Visit

arxiv.org

Tasks

text classification

Tags

Computation and LanguageArtificial Intelligence

Similaires

radheshj/Multilingual-Sycophancy-BenchmarkExploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages?Towards Scalable Multilingual AI:Benchmarking CompressedTransformer Models Across LanguagesBridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South LanguagesHow Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons PerspectiveMultilingual Biosecurity Safety Evaluation of Open-Weight Language Models: Evidence from African Languages

radheshj/Multilingual-Sycophancy-Benchmark

A multilingual sycophancy benchmark exposing how AI safety guardrails fail in low-resource languages

Exploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages?

Prior research has revealed that certain abstract concepts are linearly represented as directions in

Towards Scalable Multilingual AI:Benchmarking CompressedTransformer Models Across Languages

Large Language Models (LLMs) have demonstrated remarkable proficiency in natural language processing

Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages

Large language models (LLMs) are being deployed across the Global South, where everyday use involves

How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective

Multilingual Alignment is an effective and representative paradigm to enhance LLMs' multilingual cap

Multilingual Biosecurity Safety Evaluation of Open-Weight Language Models: Evidence from African Languages