Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Cross-lingual Name Tagging and Linking for 282 Languages

Domain:

natural language processing

Record type:

paper
The ambitious goal of this work is to develop a cross-lingual name tagging and linking framework for 282 languages that exist in Wikipedia. Given a document in any of these languages, our framework is able to identify name mentions, assign a coarse-grained or fine-grained type to each mention, and link it to an English Knowledge Base (KB) if it is linkable. We achieve this goal by performing a series of new KB mining methods: generating “silver-standard” annotations by transferring annotations from English to other languages through cross-lingual links and KB properties, refining annotations through self-training and topic selection, deriving language-specific morphology features from anchor links, and mining word translation pairs from cross-lingual links. Both name tagging and linking results for 282 languages are promising on Wikipedia data and on-Wikipedia data.

Visit

aclanthology.org

Connected records

dataset

Tasks

named entity recognitioninformation extractiontransfer learning

Languages

AfrikaansAmharicArabic, Egyptian SpokenIgboKinyarwandaLingalaSomaliSwahiliYoruba

Licenses

Similar

Joint Multilingual Supervision for Cross-lingual Entity LinkingLwazi II Cross-lingual Proper Name CorpusZero Resource Cross-Lingual Part Of Speech TaggingNeural Cross-Lingual Coreference Resolution and its Application to Entity Linkingzig-kwin-hu/Low-Resource-Name-TaggingUnsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora Only

Joint Multilingual Supervision for Cross-lingual Entity Linking

Cross-lingual Entity Linking (XEL) aims to ground entity mentions written in any language to an Engl

Lwazi II Cross-lingual Proper Name Corpus

Prompted audio recordings of personal names in different languages, produced by 20 speakers with dif

Zero Resource Cross-Lingual Part Of Speech Tagging

Part of speech tagging in zero-resource settings can be an effective approach for low-resource langu

Neural Cross-Lingual Coreference Resolution and its Application to Entity Linking

We propose an entity-centric neural cross-lingual coreference model that builds on multi-lingual emb

zig-kwin-hu/Low-Resource-Name-Tagging

This is the repository of paper "low-resource Name Tagging Learned with Weakly Labeled Data" accepte

Unsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora Only

Due to the scarcity of part-of-speech annotated data, existing studies on low-resource languages typ