NERAMazigh is a manually annotated Named Entity Recognition (NER) dataset for the Amazigh language, a severely under-resourced member of the Afro-Asiatic language family. The corpus contains approximately 90,000 tokens collected from diverse textual sources, including educational materials, literary works, institutional publications, and news articles. The dataset is annotated using the BIO tagging scheme with a fine-grained taxonomy of 16 entity categories.
The annotation process follows a rigorous two-stage workflow involving three annotators and achieves a high level of inter-annotator agreement (Fleiss’ κ = 0.92), ensuring the reliability and consistency of the annotations.
To facilitate comparability with existing NER benchmarks, the dataset also includes an additional version formatted according to the CoNLL-2003 schema. In this version, the original fine-grained entity categories are mapped to the standard CoNLL entity types (PER, ORG, LOC), while the remaining categories are grouped under the MISC label.
By providing both a fine-grained annotation scheme and a CoNLL-compatible version, NERAMazigh supports a wide range of experimental settings and enables benchmarking across different NER architectures and evaluation protocols. This resource aims to support future research in Amazigh NLP and contribute to the broader development of language technologies for low-resource languages.