Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
PluBerRan
Hôte:avatar
We present judgeWEL, a dataset for named entity recognition (NER) in Luxembourgish, automatically labelled and subsequently verified using large language models (LLM) in a novel pipeline. Building datasets for under-represented languages remains one of the major bottlenecks in natural language processing, where the scarcity of resources and linguistic particularities make large-scale annotation costly and potentially inconsistent. To address these challenges, we propose and evaluate a novel approach that leverages Wikipedia and Wikidata as structured sources of weak supervision. By exploiting internal links within Wikipedia articles, we infer entity types based on their corresponding Wikidata entries, thereby generating initial annotations with minimal human intervention. Because such links are not uniformly reliable, we mitigate noise by employing and comparing several LLMs to identify and retain only high-quality labelled sentences. The resulting corpus is approximately five times larger than the currently available Luxembourgish NER dataset and offers broader and more balanced coverage across entity categories, providing a substantial new resource for multilingual and low-resource NER research. Accepted at LREC 2026

Visit

arxiv.org

Tasks

information extractionnamed entity recognition

Tags

Computation and LanguageArtificial Intelligence

Similaires

Do "English" Named Entity Recognizers Work Well on Global Englishes?How Well Do LLMs Understand Tunisian Arabic?DzNER: A large Algerian Named Entity Recognition datasetNERAMazigh: A Named Entity Recognition Dataset for the Amazigh Language MasakhaNER: Named Entity Recognition Dataset for 20 African languages.YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset

Do "English" Named Entity Recognizers Work Well on Global Englishes?

The vast majority of the popular English named entity recognition (NER) datasets contain American or

How Well Do LLMs Understand Tunisian Arabic?

Large Language Models (LLMs) are the engines driving today's AI agents. The better these models unde

DzNER: A large Algerian Named Entity Recognition dataset

NERAMazigh: A Named Entity Recognition Dataset for the Amazigh Language

NERAMazigh is a manually annotated Named Entity Recognition (NER) dataset for the Amazigh language,

MasakhaNER: Named Entity Recognition Dataset for 20 African languages.

YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset

Named Entity Recognition (NER) is a foundational NLP task, yet research in Yorùbá has been constrain