Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

From Active Learning to Semantic Data Augmentation: Exploring the Limits of Named Entity Recognition in Low-Resource Arabic Dialects

Domaine:

natural language processing

Type de record:

paper
Créateur:
ManChi
Éditeur:
Ass
Hôte:
Named Entity Recognition (NER) for Arabic dialects faces persistent challenges due to the scarcity of annotated resources, severe class imbalance, and high linguistic variability. These factors hinder both model accuracy and cross-dialect generalization, creating a pressing need for more efficient annotation strategies and robust learning approaches. Motivated by these challenges, this work investigates the integration of active learning strategies with a semantic oversampling method to enhance performance on Algerian and Moroccan dialectal corpora. Three sampling strategies, Random, Uncertainty, and Diversity, are evaluated at incremental annotation levels (20%, 40%, 60%, and 80%), both with and without oversampling. The proposed semantic oversampling approach generates contextually coherent synthetic examples to alleviate underrepresented entity classes. Experiments conducted with three pre-trained language models, AraBERT, MARBERT, and Multi-dialect-BERT-Base-Arabic, demonstrate that semantic oversampling provides substantial early-stage improvements, particularly in recall, with consistent benefits observed across models. However, overall F1-scores remain modest (≤ 55%), and cross-dialect transfer performance is still limited. These findings indicate that while combining active learning with semantic oversampling improves annotation efficiency and model robustness, further progress in dialectal NER will require richer, more diverse datasets and dialect-aware modeling techniques.

Visit

doi.org

Tasks

information extractionnamed entity recognition

Languages

Arabic, Algerian Spoken

Licenses

https://creativecommons.org/licenses/by/4.0/legalcode

Similaires

ANEA: Distant Supervision for Low-Resource Named Entity RecognitionA Hybrid Method for Low-Resource Named Entity RecognitionSALT-Style Self-Augmentation for Cross-Lingual Named Entity Recognition F1 Scores on WikiANN in Low-Resource SettingsPerformance comparison of cross-lingual transfer learning methods for Named Entity Recognition in low-resource African languagesAnalysing the effects of transfer learning on low-resourced named entity recognition performanceText-To-Speech Data Augmentation for Low Resource Speech Recognition

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

Distant supervision allows obtaining labeled training corpora for low-resource settings where only l

A Hybrid Method for Low-Resource Named Entity Recognition

Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse a

SALT-Style Self-Augmentation for Cross-Lingual Named Entity Recognition F1 Scores on WikiANN in Low-Resource Settings

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Performance comparison of cross-lingual transfer learning methods for Named Entity Recognition in low-resource African languages

Cross-lingual transfer learning enables NLP for low-resource languages by leveraging labeled data fr

Analysing the effects of transfer learning on low-resourced named entity recognition performance

Transfer learning has led to large gains in performance for nearly all NLP tasks while making downstream models easier and faster to train. This has also been extended to low-resourced languages, with some success. We investigate the properties of transfer learning

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r