Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

WikiANN

Domain:

natural language processing

Record type:

dataset
Creator:
uni
Host:
WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, and test splits of Rahimi et al. (2019), which supports 176 of the 282 languages from the original WikiANN corpus.

Visit

huggingface.co

Tasks

information extractionnamed entity recognition

Languages

AfrikaansAmharicArabic, Egyptian SpokenIgboKinyarwandaLingalaMalagasySomaliSwahiliYoruba

Licenses

unknown

Similar

WikiannComparison of Projection-Based Cross-Lingual NER Methods on WikiAnn Benchmark for Low-Resource LanguagesScaling Source Language Diversity in Multi-Source Cross-Lingual NER for Low-Resource WikiAnn Performance

Wikiann

WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, an

Comparison of Projection-Based Cross-Lingual NER Methods on WikiAnn Benchmark for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Scaling Source Language Diversity in Multi-Source Cross-Lingual NER for Low-Resource WikiAnn Performance

To better tackle the named entity recognition (NER) problem on languages with little/no labeled data