Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Construction of a comprehensive dataset for named entity recognition and entity linking in Algerian Dialectal Arabic

Domain:

natural language processing

Record type:

dataset
Creator:
WisAdeIlhHad
Publisher:
SAG
Host:
Algerian Arabic (Darija) dominates digital communication in North Africa yet remains severely under-resourced in Natural Language Processing (NLP), hindering the development of robust applications for social media analysis and e-commerce. This paper addresses this scarcity by presenting a systematic framework for constructing and benchmarking Named Entity Recognition (NER) and Entity Linking (EL) resources tailored to the dialect’s linguistic complexity. We introduce a large-scale, multi-script dataset constructed through a novel hybrid methodology that integrates manual annotation of authentic texts, automated knowledge graph extraction from Wikidata, and rule-based synthetic generation. This approach ensures diverse coverage across ten semantic categories while explicitly addressing the challenges of code-switching and orthographic variation (Arabizi and Arabic script). A transformer-based model (XLM-RoBERTa) fine-tuned on this resource achieves state-of-the-art performance, demonstrating significant robustness compared to existing baselines. Beyond the dataset, we provide a practical deployment interface and comprehensive evaluation metrics, establishing a crucial foundation for advancing NLP capabilities in North African dialects and facilitating downstream tasks such as content moderation and cultural heritage preservation.

Visit

doi.org

Tasks

code switchinginformation extractionnamed entity recognition

Languages

Arabic, Algerian Spoken

Licenses

https://journals.sagepub.com/page/policies/text-and-data-mining-license

Similar

ELNER-DZ: A Dataset for Named Entity Recognition and Entity Linking in Algerian Arabic DialectDzNER: A large Algerian Named Entity Recognition datasetNaijaNER : Comprehensive Named Entity Recognition for 5 Nigerian LanguagesNaijaNER: Comprehensive Named Entity Recognition for 5 Nigerian LanguagesNERAMazigh: A Named Entity Recognition Dataset for the Amazigh Language MasakhaNER: Named Entity Recognition Dataset for 20 African languages.

ELNER-DZ: A Dataset for Named Entity Recognition and Entity Linking in Algerian Arabic Dialect

ELNER-DZ is the first large-scale dataset for Named Entity Recognition (NER) and Entity Linking (EL)

DzNER: A large Algerian Named Entity Recognition dataset

NaijaNER : Comprehensive Named Entity Recognition for 5 Nigerian Languages

Most of the common applications of Named Entity Recognition (NER) is on English and other highly available languages. In this work, we present our findings on Named Entity Recognition for 5 Nigerian Languages (Nigerian English, Nigerian Pidgin English, Igbo, Yoruba

NaijaNER: Comprehensive Named Entity Recognition for 5 Nigerian Languages

This work explores individual and combined Named Entity Recognition in 5 Nigerian languages - Nigerian English, Pidgin, Yoruba, Hausa and Igbo.

NERAMazigh: A Named Entity Recognition Dataset for the Amazigh Language

NERAMazigh is a manually annotated Named Entity Recognition (NER) dataset for the Amazigh language,

MasakhaNER: Named Entity Recognition Dataset for 20 African languages.