Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AbdAl-ElSTay
Éditeur:
arXiv
Hôte:avatar
The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acute for Arabic sociolinguistic knowledge: credible grading requires not only linguistic fluency but deep cultural familiarity that cannot be approximated by surface-level metrics. We address this with a cross-evaluation framework instantiated on two underrepresented Arabic dialect communities: Egyptian and Iraqi Arabic. We contribute 103 validated prompt-rubric pairs (70 Egyptian, 33 Iraqi; 53 Cultural, 50 Linguistic), authored and graded by native-speaker SMEs using penalty-weighted rubrics distinguishing positive content requirements from answer-specific negative error criteria. Three frontier LLMs serve as target models (graded by human SMEs across 302 unique prompt-response pairs), while five frontier LLMs serve as automated judges enforcing a provider-level self-evaluation guard. A dual-metric scheme combining Mean Absolute Deviation (MAD) with Signed Mean Error separates directional grading bias from symmetric noise. Across 1,307 judge evaluations: GPT-5.4 is the most reliable judge (MADj = 10.21 pp, Signed Error = -1.12%); four of five judges show systematic leniency (+2.01% to +6.56%); Cultural tasks are harder to grade than Linguistic tasks for all judges (MAD gap 1.83-4.78 pp); and models substantially outperform on Egyptian prompts compared to Iraqi prompts. However, given leniency differences between Iraqi and Egyptian SMEs, we cannot solely attribute this gap to model knowledge. We therefore emphasize findings that do not assume identical leniency across human graders. Across all samples, implicit cultural reasoning -- requiring models to simulate native-speaker judgment rather than rely on lexical verification -- emerges as the primary failure mode for automated grading across all judge models.

Visit

doi.orgarxiv.org

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

CULEMO: Cultural Lenses on Emotion -- Benchmarking LLMs for Cross-Cultural Emotion UnderstandingUnderstanding Slang with LLMs: Modelling Cross-Cultural Nuances through ParaphrasingPerformance Evaluation and Validation of an AI-Driven Hyperspectral Remote Sensing Framework for Locust Surveillance Using Ground-Truth Dataamharic-ocr-ground-truthBenchmarking Large Language Models on Egyptian Arabic: Dialectal Gaps, Evaluation Challenges, and Practical InsightsCross-Lingual Auto Evaluation for Assessing Multilingual LLMs

CULEMO: Cultural Lenses on Emotion -- Benchmarking LLMs for Cross-Cultural Emotion Understanding

NLP research has increasingly focused on subjective tasks such as emotion analysis. However, existin

Understanding Slang with LLMs: Modelling Cross-Cultural Nuances through Paraphrasing

In the realm of social media discourse, the integration of slang enriches communication, reflecting

Performance Evaluation and Validation of an AI-Driven Hyperspectral Remote Sensing Framework for Locust Surveillance Using Ground-Truth Data

Desert locust outbreaks continue to threaten agricultural productivity and food security across East

amharic-ocr-ground-truth

Benchmarking Large Language Models on Egyptian Arabic: Dialectal Gaps, Evaluation Challenges, and Practical Insights

Abstract—Large Language Models (LLMs) have demonstrated remarkable performance across a wide range

Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs

Evaluating machine-generated text remains a significant challenge in NLP, especially for non-English