Abstract—Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language understanding tasks. However, their effectiveness on Arabic dialects—particularly Egyptian Arabic—remains insufficiently studied. Egyptian Arabic (EA) differs substantially from Modern Standard Arabic (MSA) in phonology, morphology, syntax, and lexicon, and it is one of the most widely spoken Arabic varieties in the world. In this paper, we present a systematic evaluation of five prominent LLMs—GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet, LLaMA-3 70B, and AraBERT—on a curated benchmark of 1,200 Egyptian Arabic tasks spanning sentiment analysis, reading comprehension, paraphrase detection, and conversational response generation. Our results reveal a consistent performance gap between MSA and EA across all models, with accuracy drops ranging from 8% to 21% depending on the task type. We further analyze common error patterns, including code-switching blind spots and morphological ambiguity, and propose evaluation guidelines tailored to low-resource Arabic dialects. Our dataset and evaluation scripts are made publicly available to encourage further research in this space. Index Terms—Egyptian Arabic, Large Language Models, NLP benchmarking, low-resource dialects, Arabic NLP, dialectal Ara- bic