Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

NERAMazigh: A Named Entity Recognition Dataset for the Amazigh Language

Domaine:

natural language processing

Type de record:

dataset
Créateur:
BanAmr
Éditeur:
Zenodo
Hôte:avatar
NERAMazigh is a manually annotated Named Entity Recognition (NER) dataset for the Amazigh language, a severely under-resourced member of the Afro-Asiatic language family. The corpus contains approximately 90,000 tokens collected from diverse textual sources, including educational materials, literary works, institutional publications, and news articles. The dataset is annotated using the BIO tagging scheme with a fine-grained taxonomy of 16 entity categories. The annotation process follows a rigorous two-stage workflow involving three annotators and achieves a high level of inter-annotator agreement (Fleiss’ κ = 0.92), ensuring the reliability and consistency of the annotations. To facilitate comparability with existing NER benchmarks, the dataset also includes an additional version formatted according to the CoNLL-2003 schema. In this version, the original fine-grained entity categories are mapped to the standard CoNLL entity types (PER, ORG, LOC), while the remaining categories are grouped under the MISC label. By providing both a fine-grained annotation scheme and a CoNLL-compatible version, NERAMazigh supports a wide range of experimental settings and enables benchmarking across different NER architectures and evaluation protocols. This resource aims to support future research in Amazigh NLP and contribute to the broader development of language technologies for low-resource languages.

Visit

doi.orgzenodo.org

Tasks

information extractionnamed entity recognition

Languages

AmazighBerber

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative AnalysisDzNER: A large Algerian Named Entity Recognition dataset MasakhaNER: Named Entity Recognition Dataset for 20 African languages.Amazigh Linguistic Dataset: Part-of-Speech Tagging, Named Entity Recognition, and Parallel Corpus (Tifinagh-English)ELNER-DZ: A Dataset for Named Entity Recognition and Entity Linking in Algerian Arabic DialectNamed entity recognition for African languages a focus on the Igbo language

Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative Analysis

This work contributes towards balancing the inclusivity and global applicability of natural language

DzNER: A large Algerian Named Entity Recognition dataset

MasakhaNER: Named Entity Recognition Dataset for 20 African languages.

Amazigh Linguistic Dataset: Part-of-Speech Tagging, Named Entity Recognition, and Parallel Corpus (Tifinagh-English)

This dataset is a comprehensive linguistic resource for the Amazigh (Berber) language, focusing on t

ELNER-DZ: A Dataset for Named Entity Recognition and Entity Linking in Algerian Arabic Dialect

ELNER-DZ is the first large-scale dataset for Named Entity Recognition (NER) and Entity Linking (EL)

Named entity recognition for African languages a focus on the Igbo language

Named Entity Recognition (NER) is a crucial task for many downstream NLP applications, including te