Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Benchmarking Large Language Models on Egyptian Arabic: Dialectal Gaps, Evaluation Challenges, and Practical Insights

Domain:

natural language processing

Record type:

datasetpaper
Creator:
YasMoh
Publisher:
Zenodo
Host:avatar
Abstract—Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language understanding tasks. However, their effectiveness on Arabic dialects—particularly Egyptian Arabic—remains insufficiently studied. Egyptian Arabic (EA) differs substantially from Modern Standard Arabic (MSA) in phonology, morphology, syntax, and lexicon, and it is one of the most widely spoken Arabic varieties in the world. In this paper, we present a systematic evaluation of five prominent LLMs—GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet, LLaMA-3 70B, and AraBERT—on a curated benchmark of 1,200 Egyptian Arabic tasks spanning sentiment analysis, reading comprehension, paraphrase detection, and conversational response generation. Our results reveal a consistent performance gap between MSA and EA across all models, with accuracy drops ranging from 8% to 21% depending on the task type. We further analyze common error patterns, including code-switching blind spots and morphological ambiguity, and propose evaluation guidelines tailored to low-resource Arabic dialects. Our dataset and evaluation scripts are made publicly available to encourage further research in this space. Index Terms—Egyptian Arabic, Large Language Models, NLP benchmarking, low-resource dialects, Arabic NLP, dialectal Ara- bic

Visit

doi.orgzenodo.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Cross-dialectal Arabic translation: comparative analysis on large language modelsTounsiBench: Benchmarking Large Language Models for Tunisian ArabicBALSAM: A Platform for Benchmarking Arabic Large Language ModelsEvaluation of Arabic Large Language Models on Moroccan DialectSouha-BH/TounsiBench-Benchmarking-Large-Language-Models-for-Tunisian-ArabicAraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models via Multi-Component Metrics

Cross-dialectal Arabic translation: comparative analysis on large language models

Introduction Exploring Arabic dialects in Natural Language Processing (NLP) is essential to underst

TounsiBench: Benchmarking Large Language Models for Tunisian Arabic

BALSAM: A Platform for Benchmarking Arabic Large Language Models

The impressive advancement of Large Language Models (LLMs) in English has not been matched across al

Evaluation of Arabic Large Language Models on Moroccan Dialect

Large Language Models (LLMs) have shown outstanding performance in many Natural Language Processing

Souha-BH/TounsiBench-Benchmarking-Large-Language-Models-for-Tunisian-Arabic

# TounsiBench-Benchmarking-Large-Language-Models-for-Tunisian-Arabic Thank you for using the Tounsi

AraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models via Multi-Component Metrics

Abstract Large Language Models (LLMs) have achieved strong multilingual capabiliti