Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Addressing the Complexity of Dialectal Arabic: An Enhanced Encoder-Decoder Ensemble Approach for Optimized Sentiment Analysis

Domain:

natural language processing

Record type:

datasetpaper
Creator:
DjaRimSamMoh
Publisher:
Ass
Host:
Sentiment analysis has become an essential tool in understanding global narratives across social, economic, political, and commercial sectors. As social media platforms increasingly produce vast amounts of uncontrolled textual data, the need to analyze content in regional languages has grown significantly. This article focuses on the unique challenges posed by sentiment analysis in Arabic and dialectal Arabic, with a special emphasis on the Algerian Dialects often referred to as Darija, which is characterized by its linguistic diversity and multilingual nature. One of the primary issues addressed is the lack of annotated datasets for this dialect and the complexities of accurately interpreting this diverse language. To tackle these challenges, we present DZDialect, a new dataset comprising 117,569 annotated comments, and we explore various advanced methodologies for sentiment analysis. Furthermore, the study compares the performance of machine learning algorithms (SVM, NBM, and KNN), deep learning algorithms (LSTM and CNN) using wor2vec as word embedding tool, and transformer-based classifiers (AraBERT Base, AraBERT Mini, AraBERT Medium, DistilBERT Multilingual, and AraGPT-2) in classifying social media posts into positive or negative sentiment categories. In addition, we introduce an innovative ensemble architecture that combines multiple pre-trained models for enhanced performance. The system leverages DistilBERT and AraBERT Base as encoders, while AraGPT-2 serves as the decoder. These components work in concert through a sophisticated stacking and voting mechanism. The results reveal promising accuracy rates, with the AraBERT Base model achieving 87.9%, the LSTM model 85%, and the SVM classifier 82%. The stacking model attained an accuracy of 91.1%, while the majority voting model reached 90%. This research contributes valuable insights into sentiment analysis for dialectal Arabic, with practical implications for real-world applications across various sectors.

Visit

doi.org

Tasks

sentiment analysistext classification

Languages

Arabic, Algerian Spoken