Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

WikiANN

Domaine:

natural language processing

Type de record:

dataset
Créateur:
uni
Hôte:
WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, and test splits of Rahimi et al. (2019), which supports 176 of the 282 languages from the original WikiANN corpus.

Visit

huggingface.co

Tasks

information extractionnamed entity recognition

Languages

AfrikaansAmharicArabic, Egyptian SpokenIgboKinyarwandaLingalaMalagasySomaliSwahiliYoruba

Licenses

unknown

Similaires

WikiannComparison of Projection-Based Cross-Lingual NER Methods on WikiAnn Benchmark for Low-Resource LanguagesScaling Source Language Diversity in Multi-Source Cross-Lingual NER for Low-Resource WikiAnn Performance

Wikiann

WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, an

Comparison of Projection-Based Cross-Lingual NER Methods on WikiAnn Benchmark for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Scaling Source Language Diversity in Multi-Source Cross-Lingual NER for Low-Resource WikiAnn Performance

To better tackle the named entity recognition (NER) problem on languages with little/no labeled data