Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Impact of Data Corruption on Named Entity Recognition for Low-resourced Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
FokBeu
Hôte:avatar
Data availability and quality are major challenges in natural language processing for low-resourced languages. In particular, there is significantly less data available than for higher-resourced languages. This data is also often of low quality, rife with errors, invalid text or incorrect annotations. Many prior works focus on dealing with these problems, either by generating synthetic data, or filtering out low-quality parts of datasets. We instead investigate these factors more deeply, by systematically measuring the effect of data quantity and quality on the performance of pre-trained language models in a low-resourced setting. Our results show that having fewer completely-labelled sentences is significantly better than having more sentences with missing labels; and that models can perform remarkably well with only 10% of the training data. Importantly, these results are consistent across ten low-resource languages, English, and four pre-trained models.

Visit

arxiv.org

Tasks

information extractionnamed entity recognition

Tags

Computation and LanguageArtificial Intelligence

Similaires

Deep Learning Transformer Architecture for Named Entity Recognition on Low Resourced Languages: State of the art resultsAnalysing the effects of transfer learning on low-resourced named entity recognition performanceAnalysing Cross-Lingual Transfer in Low-Resourced African Named Entity RecognitionMasakhaNER: Named Entity Recognition for African LanguagesNamed entity recognition for African languages a focus on the Igbo languageANEA: Distant Supervision for Low-Resource Named Entity Recognition

Deep Learning Transformer Architecture for Named Entity Recognition on Low Resourced Languages: State of the art results

This paper reports on the evaluation of Deep Learning (DL) transformer architecture models for Named

Analysing the effects of transfer learning on low-resourced named entity recognition performance

Transfer learning has led to large gains in performance for nearly all NLP tasks while making downstream models easier and faster to train. This has also been extended to low-resourced languages, with some success. We investigate the properties of transfer learning

Analysing Cross-Lingual Transfer in Low-Resourced African Named Entity Recognition

Transfer learning has led to large gains in performance for nearly all NLP tasks while making downst

MasakhaNER: Named Entity Recognition for African Languages

We take a step towards addressing the under-representation of the African continent in NLP research by creating the first large publicly available high-quality dataset for named entity recognition (NER) in ten African languages, bringing together a variety of stake

Named entity recognition for African languages a focus on the Igbo language

Named Entity Recognition (NER) is a crucial task for many downstream NLP applications, including te

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

Distant supervision allows obtaining labeled training corpora for low-resource settings where only l