Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SevImaDufSch
Hôte:avatar
Parallel corpora are ideal for extracting a multilingual named entity (MNE) resource, i.e., a dataset of names translated into multiple languages. Prior work on extracting MNE datasets from parallel corpora required resources such as large monolingual corpora or word aligners that are unavailable or perform poorly for underresourced languages. We present CLC-BN, a new method for creating an MNE resource, and apply it to the Parallel Bible Corpus, a corpus of more than 1000 languages. CLC-BN learns a neural transliteration model from parallel-corpus statistics, without requiring any other bilingual resources, word aligners, or seed data. Experimental results show that CLC-BN clearly outperforms prior work. We release an MNE resource for 1340 languages and demonstrate its effectiveness in two downstream tasks: knowledge graph augmentation and bilingual lexicon induction. LREC 2022

Visit

arxiv.org

Tasks

information extractionnamed entity recognition

Tags

Computation and Language

Similaires

Parameter-Efficient Fine-Tuning for Equitable Named Entity Recognition in African Languages: A Lora-Based ApproachA Hybrid Method for Low-Resource Named Entity RecognitionLinguoNER: A Language-Agnostic Framework for Named Entity Recognition in Low-Resource Languages with a Focus on YambetaMasakhaNER: Named Entity Recognition for African LanguagesA NAMED ENTITY RECOGNITION SYSTEM FOR BASSA, EBIRA, AND OKUN LANGUAGESTowards a Crowdsourcing Platform for Low Resource Languages -- A Collectivist Approach

Parameter-Efficient Fine-Tuning for Equitable Named Entity Recognition in African Languages: A Lora-Based Approach

Named entity recognition (NER) remains a challenge for Africa

A Hybrid Method for Low-Resource Named Entity Recognition

Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse a

LinguoNER: A Language-Agnostic Framework for Named Entity Recognition in Low-Resource Languages with a Focus on Yambeta

This paper presents LinguoNER, a practical and extensible framework for bootstrapping Named Entity R

MasakhaNER: Named Entity Recognition for African Languages

We take a step towards addressing the under-representation of the African continent in NLP research by creating the first large publicly available high-quality dataset for named entity recognition (NER) in ten African languages, bringing together a variety of stake

A NAMED ENTITY RECOGNITION SYSTEM FOR BASSA, EBIRA, AND OKUN LANGUAGES

This study focuses on developing a Named Entity Recognition (NER) system tailored specifically for l

Towards a Crowdsourcing Platform for Low Resource Languages -- A Collectivist Approach

This work demonstrates how semi-supervised learning and human-in-the-loop crowdsourcing can help neu