Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models via Multi-Component Metrics

Domain:

natural language processing

Record type:

dataset
Creator:
FatHum
Publisher:
Spr
Host:
Abstract Large Language Models (LLMs) have achieved strong multilingual capabilities, yet their behavior across Arabic dialects remains insufficiently studied, particularly with respect to cultural and stereotypical bias. Existing bias benchmarks are largely English-centric and typically treat Arabic as a unified language, overlooking substantial dialectal variation. In this paper, we present AraRegBias, a multidimensional framework for evaluating dialectal and stereotypical bias in Arabic LLMs across six dialect groups: Modern Standard Arabic (MSA), Egyptian, Gulf, Levantine, Iraqi, and Maghrebi Arabic. The framework integrates lexical, semantic, and toxicity-based signals into a unified normalized metric, the AraRegBias Index (ARI), enabling consistent and interpretable cross-dialect comparison. We construct a controlled benchmark dataset comprising neutral, cultural, and stereotype-oriented prompts to systematically probe model behavior under different socio-cultural contexts. Experiments conducted on an instruction-tuned transformer model reveal statistically significant dialect dependent variations in generated outputs, with stereotype prompts amplifying disparities across dialect groups. Results show that underrepresented dialects, particularly Iraqi Arabic, exhibit stronger negative bias tendencies compared to higher-resource dialects. Comprehensive statistical analyses further confirm robust inter-dialect differences, while correlation analysis indicates that lexical and semantic components dominate the ARI, with toxicity contributing minimally. These findings suggest that bias in Arabic LLMs is predominantly implicit rather than overtly toxic, underscoring the need for dialect-aware fairness evaluation frameworks in Arabic NLP systems.

Visit

doi.org

Languages

Arabic, Libyan SpokenArabic, Moroccan Spoken

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language ModelsCross-dialectal Arabic translation: comparative analysis on large language modelsEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"{S}tereo{S}et: Measuring stereotypical bias in pretrained language modelsBenchmarking Large Language Models on Egyptian Arabic: Dialectal Gaps, Evaluation Challenges, and Practical InsightsSociolinguistic Bias and Language Inequality in Large Language Models

AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models

Existing AI bias evaluation benchmarks largely reflect Western perspectives, leaving African context

Cross-dialectal Arabic translation: comparative analysis on large language models

Introduction Exploring Arabic dialects in Natural Language Processing (NLP) is essential to underst

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

{S}tereo{S}et: Measuring stereotypical bias in pretrained language models

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or African Americans are athletic. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real-world d

Benchmarking Large Language Models on Egyptian Arabic: Dialectal Gaps, Evaluation Challenges, and Practical Insights

Abstract—Large Language Models (LLMs) have demonstrated remarkable performance across a wide range

Sociolinguistic Bias and Language Inequality in Large Language Models

Large language models (LLMs) are increasingly deployed across multilingual applications, yet persist