Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
ZahAsg
Hôte:avatar
We introduce MENAValues, a novel benchmark designed to evaluate the cultural alignment and multilingual biases of large language models (LLMs) with respect to the beliefs and values of the Middle East and North Africa (MENA) region, an underrepresented area in current AI evaluation efforts. Drawing from large-scale, authoritative human surveys, we curate a structured dataset that captures the sociocultural landscape of MENA with population-level response distributions from 16 countries. To probe LLM behavior, we evaluate diverse models across multiple conditions formed by crossing three perspective framings (neutral, personalized, and third-person/cultural observer) with two language modes (English and localized native languages: Arabic, Persian, Turkish). Our analysis reveals three critical phenomena: "Cross-Lingual Value Shifts" where identical questions yield drastically different responses based on language, "Reasoning-Induced Degradation" where prompting models to explain their reasoning worsens cultural alignment, and "Logit Leakage" where models refuse sensitive questions while internal probabilities reveal strong hidden preferences. We further demonstrate that models collapse into simplistic linguistic categories when operating in native languages, treating diverse nations as monolithic entities. MENAValues offers a scalable framework for diagnosing cultural misalignment, providing both empirical insights and methodological tools for developing more culturally inclusive AI.

Visit

arxiv.org

Tags

Computation and Language

Similaires

CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket AnalyticsRethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMsPak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu BenchmarkFraming Political Bias in Multilingual LLMs Across Pakistani LanguagesEvaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and EnglishAssessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs

CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics

Cricket is the second most popular sport worldwide, with billions of fans seeking advanced statistic

Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs

Cross-lingual alignment (CLA) aims to align multilingual representations, enabling Large Language Mo

Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark

Large language models (LLMs) demonstrate strong Helpfulness, Harmlessness, and Honesty (3H) alignmen

Framing Political Bias in Multilingual LLMs Across Pakistani Languages

Large Language Models (LLMs) increasingly shape public discourse, yet most evaluations of political

Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English

Although LLMs have been extremely effective in a large number of complex tasks, their understanding

Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs

Recent trends in LLMs development clearly show growing interest in the use and application of sovere