Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Survey of Large Language Models for Arabic Language and its Dialects

Domain:

natural language processing

Record type:

paper
Creator:
MasAl-Al-
Host:avatar
This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, decoder-only, and encoder-decoder models, along with the datasets used for pre-training, spanning Classical Arabic, Modern Standard Arabic, and Dialectal Arabic. The study also explores monolingual, bilingual, and multilingual LLMs, analyzing their architectures and performance across downstream tasks, such as sentiment analysis, named entity recognition, and question answering. Furthermore, it assesses the openness of Arabic LLMs based on factors, such as source code availability, training data, model weights, and documentation. The survey highlights the need for more diverse dialectal datasets and attributes the importance of openness for research reproducibility and transparency. It concludes by identifying key challenges and opportunities for future research and stressing the need for more inclusive and representative models. Submitted to ACM Transactions on Asian and Low-Resource Language Information Processing

Visit

arxiv.org

Tasks

information extractionlanguage modelingnamed entity recognitionsentiment analysistext classification

Tags

Computation and LanguageArtificial Intelligence

Similar

Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its DialectsBALSAM: A Platform for Benchmarking Arabic Large Language ModelsArabic Large Language Models for Medical Text GenerationTounsiBench: Benchmarking Large Language Models for Tunisian ArabicMulti-Label Emotion Recognition in Low-Resource Dialects: A Case Study on Algerian Arabic with Large Language ModelsLarge Language Models for Arabic Sentiment Analysis and Dialect Detection: A Systematic Review

Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its Dialects

We present state-of-the-art results on morphosyntactic tagging across different varieties of Arabic

BALSAM: A Platform for Benchmarking Arabic Large Language Models

The impressive advancement of Large Language Models (LLMs) in English has not been matched across al

Arabic Large Language Models for Medical Text Generation

Efficient hospital management systems (HMS) are critical worldwide to address challenges such as ove

TounsiBench: Benchmarking Large Language Models for Tunisian Arabic

Multi-Label Emotion Recognition in Low-Resource Dialects: A Case Study on Algerian Arabic with Large Language Models

Large Language Models for Arabic Sentiment Analysis and Dialect Detection: A Systematic Review

## Overview and Motivation This research project is a systematic review that consolidates and criti