Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Large Language Models for Arabic Sentiment Analysis and Dialect Detection: A Systematic Review

Domaine:

natural language processing

Type de record:

paper
Créateur:
HesManReh
Éditeur:
Cen
Éditeur:
OSF
Hôte:avatar
## Overview and Motivation This research project is a systematic review that consolidates and critically evaluates recent work on large language models (LLMs) and transformer-based methods applied to two closely related tasks in Arabic natural language processing: sentiment analysis (SA), the automatic identification of the emotional polarity of text, and dialect identification (DI), the recognition of which regional variety of Arabic a text is written in. The work is motivated by a persistent gap in the field. Arabic is the official language of more than twenty countries and the native tongue of over 400 million speakers, yet it has been historically under-served in NLP. It presents a distinctive cluster of difficulties: it is morphologically rich (a root-and-pattern system generates many surface forms, producing data sparsity and high out-of-vocabulary rates), diacritics are routinely dropped in everyday writing, and — most importantly — it is diglossic. Formal Modern Standard Arabic (MSA) coexists with a continuum of spoken dialects that dominate social media, reviews, and conversational text but lack consistent spelling. Dialects such as Moroccan Darija and Egyptian Arabic can differ as much as separate Romance languages, and sentiment cues like negation and intensifiers are themselves dialect-specific, so a model trained on one variety can be actively misled by another. Compounding this, annotated dialectal resources are scarce and unevenly distributed, with Gulf and Maghrebi varieties under-represented precisely where commercial demand (tourism, hospitality, services) is high. ## Purpose The central purpose is to determine whether efficient, smaller, dialect-aware Arabic models can rival large, general-purpose, proprietary LLMs — treating computational and data efficiency not as a secondary convenience but as a primary axis of comparison. Existing surveys of Arabic SA largely predate the LLM era, tend to treat sentiment and dialect separately, and rarely contrast Arabic-specific versus multilingual models, open versus proprietary systems, or full fine-tuning versus parameter-efficient methods through an efficiency lens. This review sets out to fill that gap. ## Methodology The review follows the PRISMA 2020 guidelines for systematic reviews. The authors identified 219 records across ACL Anthology, IEEE Xplore, ScienceDirect, and Google Scholar (complemented by Springer, ACM, Scopus, and Web of Science), using Boolean queries combining language-scope terms, task terms, and method terms. After removing duplicates and applying two stages of screening against predefined eligibility criteria, 21 primary studies (spanning 2020–2026) were retained. Each was coded systematically for its tasks, datasets, model families, adaptation strategies, and reported evaluation metrics (accuracy, precision, recall, F1, and macro-F1). The project also builds a structured taxonomy of approaches, organizing the literature into encoder fine-tuning, generative prompting, hybrid/ensemble systems, and efficient/multi-task methods. ## Key Findings The review's principal conclusions are that dialect-aware Arabic encoders — models such as AraBERT, MARBERTv2, CAMeLBERT-DA, QARiB, and AraELECTRA — remain the most accurate and reproducible backbone for dialectal sentiment analysis when fine-tuned or ensembled, reaching macro-F1 scores up to roughly 0.81 on multi-dialect hotel reviews. Generative LLMs (GPT-4o, Gemini 1.5, LLaMA-3, Qwen2.5, and the Arabic-centric Jais) are convenient and flexible but frequently underperform fine-tuned encoders under zero- and few-shot prompting. Efficient adaptation narrows this gap substantially: LoRA and retrieval-style prompting improve generative-model results, and contrastive few-shot methods such as SetFit approach full fine-tuning accuracy with as few as 32 labeled examples per class. The overarching outcome is an evidence-based argument that competitive Arabic dialectal sentiment analysis does not require the largest models. The project delivers five concrete contributions: a PRISMA-based synthesis of the 21 studies; a taxonomy of methods with model-type comparisons; a critical comparison of full fine-tuning versus parameter-efficient adaptation and of large versus small models; a discussion of persistent challenges (data scarcity, annotation inconsistency, dialect overlap, code-switching, domain shift, bias, and compute limits); and a forward-looking roadmap toward Arabic-native foundation models, efficient small language models, continual and cross-dialect learning, multi-task formulations, and retrieval augmentation. The review is timely, coinciding with a concentrated burst of relevant work at venues such as the 2025 RANLP shared task on sentiment analysis for Arabic dialects in the hospitality domain (AHaSIS), which provides a common dialectal benchmark for directly comparing these competing approaches.

Visit

doi.org

Tasks

language identificationsentiment analysistext classification

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Tags

Physical Sciences and MathematicsComputer SciencesArtificial Intelligence and Robotics

Licenses

Creative Commons Zero v1.0 Universalhttps://creativecommons.org/publicdomain/zero/1.0/legalcode