Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Al-AbdRasLeg
Éditeur:
arXiv
Hôte:avatar
Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and dialect encodes social meaning that MSA-centric evaluation cannot capture. We present a rubric-based benchmark for the Saudi dialect, comprising 31 expert-authored prompts spanning idiomatic, pragmatic, lexical, and culturally-embedded phenomena, each paired with an expert-established ground truth. Our methodology separates evaluation into a model-agnostic phase, in which atomic, MECE positive criteria are derived solely from the ground truth, and a model-specific phase, in which four state-of-the-art systems -- Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 -- are scored against those criteria and penalised for errors they actively introduce. Across 124 model-prompt evaluations we catalogue 466 error instances under a nine-category taxonomy. The four systems cluster within a narrow macro-average band (42.7%-53.1%), with no model exceeding 55% and every model recording at least one negative-scoring prompt, confirming that Saudi dialectal competence remains broadly unsolved. Notably, Ambiguous Framing is the dominant failure mode (37.3% of errors) while outright Hallucination accounts for only 11.2%, indicating that models fail less by stating falsehoods than by distorting register and flattening pragmatic nuance. We further observe a consistency-versus-ceiling trade-off and model-distinctive error signatures. We release the full prompt set, ground truths, and scored rubrics to support reproducible dialectal evaluation.

Visit

doi.org

Tags

Computation and Language (cs.CL)Artificial Intelligence (cs.AI)Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in UrduAraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language ModelsDziriFake: A Dataset and Comparative Benchmark of Large Language Models for Fake News Detection In Algerian DialectBeyond Aesthetics: Cultural Competence in Text-to-Image ModelsM3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language ModelsFrom Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories: grammar, morphology, spelling, reading comprehension, and syntax, through 150 ex

DziriFake: A Dataset and Comparative Benchmark of Large Language Models for Fake News Detection In Algerian Dialect

Beyond Aesthetics: Cultural Competence in Text-to-Image Models

Text-to-Image (T2I) models are being increasingly adopted in diverse global communities where they c

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Despite the existence of various benchmarks for evaluating natural language processing models, we ar

From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

Recent progress in NLP research has demonstrated remarkable capabilities of large language models (L