Abstract
Large Language Models (LLMs) have achieved strong multilingual capabilities, yet their behavior across Arabic dialects remains insufficiently studied, particularly with respect to cultural and stereotypical bias. Existing bias benchmarks are largely English-centric and typically treat Arabic as a unified language, overlooking substantial dialectal variation. In this paper, we present AraRegBias, a multidimensional framework for evaluating dialectal and stereotypical bias in Arabic LLMs across six dialect groups: Modern Standard Arabic (MSA), Egyptian, Gulf, Levantine, Iraqi, and Maghrebi Arabic. The framework integrates lexical, semantic, and toxicity-based signals into a unified normalized metric, the AraRegBias Index (ARI), enabling consistent and interpretable cross-dialect comparison. We construct a controlled benchmark dataset comprising neutral, cultural, and stereotype-oriented prompts to systematically probe model behavior under different socio-cultural contexts. Experiments conducted on an instruction-tuned transformer model reveal statistically significant dialect dependent variations in generated outputs, with stereotype prompts amplifying disparities across dialect groups. Results show that underrepresented dialects, particularly Iraqi Arabic, exhibit stronger negative bias tendencies compared to higher-resource dialects. Comprehensive statistical analyses further confirm robust inter-dialect differences, while correlation analysis indicates that lexical and semantic components dominate the ARI, with toxicity contributing minimally. These findings suggest that bias in Arabic LLMs is predominantly implicit rather than overtly toxic, underscoring the need for dialect-aware fairness evaluation frameworks in Arabic NLP systems.