Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context

Domain:

natural language processing

Record type:

paper
Creator:
VerBha
Host:avatar
Alignment tuning has enabled large language models to excel in reasoning, instruction-following, and minimizing harmful generations. However, despite their widespread deployment, these models exhibit a monolingual bias, raising concerns about the effectiveness of alignment across languages. Current alignment methods predominantly focus on English, leaving it unclear how alignment mechanism generalize to multilingual settings. To address this, we conduct a systematic analysis of distributional shifts in the embedding space of LLMs before and after alignment, uncovering its impact on model behavior across diverse languages. We leverage the alignment-induced separation in safety space as a quantitative tool to measure how alignment enforces safety constraints. Our study evaluates seven LLMs using balanced toxicity datasets and parallel text-detoxification benchmarks, revealing substantial disparities in the latent representation space between high-resource and low-resource languages. These findings underscore the need for language-specific fine-tuning to ensure fair, reliable and robust multilingual alignment. Our insights provide a foundation for developing truly safe multilingual LLMs, emphasizing the urgency of addressing alignment gaps in underrepresented languages. 14 pages, 11 Figures, 2 Tables, currently under review at ACL 2025

Visit

arxiv.org

Tags

Computation and Language

Similar

EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual ContextContext-Aware Large Language Models for Multilingual UnderstandingLLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual FeedbackControlling Language Confusion in Multilingual LLMsCan LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacksCan LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval over haystacks

EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context

Large Language Models (LLMs) have achieved impressive progress across a wide range of tasks, yet the

Context-Aware Large Language Models for Multilingual Understanding

Multilingual large language models (LLMs) have demonstrated strong performance in cross-lingual task

LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback

To democratize large language models (LLMs) to most natural languages, it is imperative to make thes

Controlling Language Confusion in Multilingual LLMs

Large language models often suffer from language confusion, a phenomenon in which responses are part

Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks

Existing multilingual long-context benchmarks, often based on the popular needle-in-a-haystack test,

Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval over haystacks

Existing multilingual long-context benchmarks, often based on the popular needle-in-a-haystack test,