Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Starved by scarcity: A three scarcities framework and decision roadmap for medical named entity recognition in low-resource languages

Domaine:

natural language processinghealthcare

Type de record:

paper
Créateur:
EliPol
Éditeur:
ACC
Hôte:
Named entity recognition (NER) is a foundational task in medical natural language processing (NLP), enabling the extraction of clinically relevant information from unstructured text and supporting downstream applications such as information retrieval and clinical decision support. While pre-trained language models such as BioBERT, ClinicalBERT, and PubMedBERT have achieved state-of-the-art performance for English medical NER, progress in low-resource languages remains fragmented and uneven, even though healthcare systems worldwide generate vast amounts of textual data whose value depends on the availability of language technologies. Existing surveys provide valuable overviews but largely treat low-resource settings as a single category, overlooking important differences in the constraints practitioners face. To address this gap, we conducted a systematic review of 46 studies on medical NER in low-resource languages published since 2020. Based on the synthesized evidence, we introduce the three scarcities framework, which conceptualizes low-resource medical NLP as the interaction among data, model, and infrastructure scarcities. We also propose a practitioner-oriented decision roadmap that maps appropriate modeling and data augmentation strategies to different resource configurations. Our review develops a structured taxonomy of approaches, including cross-lingual transfer, in-domain pre-training, annotation projection, back-translation, and large language model-based synthetic data generation, and shows that the effectiveness of techniques depends less on the method itself than on the underlying scarcity profile. We also identify persistent weaknesses in evaluation practices, including inconsistent reporting and limited use of statistical significance testing.

Visit

doi.org

Tasks

named entity recognitioninformation extraction

Licenses

https://creativecommons.org/licenses/

Similaires

LinguoNER: A Language-Agnostic Framework for Named Entity Recognition in Low-Resource Languages with a Focus on YambetaA Hybrid Method for Low-Resource Named Entity RecognitionANEA: Distant Supervision for Low-Resource Named Entity RecognitionNamed Entity Recognition in Low-resource Languages using Cross-lingual distributional word representationMeta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine LanguagesNamed Entity Recognition with Word Embeddings and Wikipedia Categories for a Low-Resource Language

LinguoNER: A Language-Agnostic Framework for Named Entity Recognition in Low-Resource Languages with a Focus on Yambeta

This paper presents LinguoNER, a practical and extensible framework for bootstrapping Named Entity R

A Hybrid Method for Low-Resource Named Entity Recognition

Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse a

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

Distant supervision allows obtaining labeled training corpora for low-resource settings where only l

Named Entity Recognition in Low-resource Languages using Cross-lingual distributional word representation

Named Entity Recognition (NER) is a fundamental task in many NLP applications that seek to identify

Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages

Named-entity recognition (NER) in low-resource languages is usually tackled by finetuning very large

Named Entity Recognition with Word Embeddings and Wikipedia Categories for a Low-Resource Language

In this article, we propose a word embedding--based named entity recognition (NER) approach. NER is