Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints

Domaine:

natural language processing

Type de record:

paper
Créateur:
Marivate, Vukosi
Hôte:avatar
Over the past decade, low-resource natural language processing (NLP) has experienced explosive growth, propelled by cross-lingual transfer, massively multilingual models, and the rapid proliferation of benchmarks. Yet this apparent progress masks a critical, insufficiently examined tension: the deep sociolinguistic expertise required to evaluate increasingly complex generative systems is severely strained, inequitably distributed, and structurally marginalised. We present a critical narrative survey of low-resource NLP evaluation (2014--present), tracing its evolution across three phases: early heuristic optimism, the illusions of top-down benchmark scaling, and the current era of generative bottlenecks. We conceptualise the \emph{Annotation Scarcity Paradox}, the structural friction arising when the technical capacity to scale models vastly outpaces the sovereign human infrastructure required to authentically evaluate them. By examining extractive data pipelines, undercompensated ``ghost work'', and language data flaring, we argue that this paradox threatens the epistemic validity of reported progress. We survey emerging responses -- including data augmentation, model-based evaluation, participatory curation, and annotation-efficient approaches via item response theory and active learning -- and assess their equity and validity trade-offs. We close with a practitioner call to action, arguing that overcoming this bottleneck requires a paradigm shift from transactional data extraction to relational, community-embedded evaluation rooted in epistemic governance, data sovereignty, and shared ownership. Under Review

Visit

arxiv.org

Tags

Computation and Language

Similaires

Reducing Annotation Dependency in Low-Resource NLP Using Bootstrap Techniques: A Shona Case StudyMetaphor Annotation in Sesotho Text Corpus: Towards the Representation of Resource-Scarce Languages in NLPDigital Entrepreneurship and the Acceleration of Sustainable Development in Nigeria: Opportunities and Constraintsoyinkanchekwas/low-resource-nlp-toolkitToluClassics/Low-Resource-NLP-Tutorialsellis-nlp/low-resource-MT

Reducing Annotation Dependency in Low-Resource NLP Using Bootstrap Techniques: A Shona Case Study

Metaphor Annotation in Sesotho Text Corpus: Towards the Representation of Resource-Scarce Languages in NLP

Digital Entrepreneurship and the Acceleration of Sustainable Development in Nigeria: Opportunities and Constraints

Digital entrepreneurship in Nigeria holds immense potential to drive sustainable development, yet it

oyinkanchekwas/low-resource-nlp-toolkit

Selective language routing and code-switch audits, evaluated on a reproducible AfriSenti benchmark.

ToluClassics/Low-Resource-NLP-Tutorials

Getting started in NLP for low resource languages # Low-Resource-NLP The goal of this repository i

ellis-nlp/low-resource-MT

Work on machine translation in low resource scenarios and for minority and under-resourced languages