Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

UzNER-100K: A Human-Reviewed Uzbek NER Benchmark with Gazetteer-Augmented Transformer Modeling and Robustness Analysis

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Bob
Éditeur:
Zenodo
Hôte:avatar
UzNER-100K is a human-reviewed benchmark dataset for Uzbek named entity recognition (NER). The released collection contains 114,269 sentences, 1,184,426 tokens, and 200,083 entity mentions. The benchmark includes a 100,000-sentence training split annotated with 18 fine-grained entity types under a BIOES tagging scheme, together with development, standard test, gold candidate, and hard-evaluation subsets. The final training configuration follows a mixed-origin design with 70,000 human-reviewed real sentences and 30,000 reviewed synthetic sentences generated through an LLM-assisted pipeline. The benchmark was designed for both standard model comparison and robustness-oriented evaluation. All splits are fully disjoint, and sentence-level auditing confirms zero train–test overlap. The release includes benchmark files, preprocessing utilities, evaluation tools, annotation-support materials, and reproducibility documentation. The resource is intended to support research on Uzbek NER, low-resource NLP, multilingual and monolingual transformer benchmarking, hybrid NER modeling, and robustness analysis.

Visit

doi.orgzenodo.org

Tasks

information extractionnamed entity recognition

Tags

Uzbek NERnamed entity recognitionbenchmark datasetlow-resource NLPBIOESCRF

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (c) 2026 Bobur Saidov and co-authors.http://rightsstatements.org/vocab/InC/1.0/

Similaires

TRANSFORMER-BASED SPAM DETECTION FOR UZBEK LANGUAGE: A COMPARATIVE STUDY WITH TRADITIONAL MACHINE LEARNING MODELSCross-lingual NER Robustness with Variable Source Language Counts in XTREME-NERRobustness of Cross-Lingual NER in Low-Resource Languages with Typological DivergenceCross-lingual NER robustness in low-resource languages with source language diversity**Cross-lingual Transfer Robustness in Low-Resource African NER with Script Variation**Direct English-to-Yoruba speech Translation model using Transformer with Augmented Attention Mechanism

TRANSFORMER-BASED SPAM DETECTION FOR UZBEK LANGUAGE: A COMPARATIVE STUDY WITH TRADITIONAL MACHINE LEARNING MODELS

Spam detection remains a critical challenge in natural language processing, particularly for low-res

Cross-lingual NER Robustness with Variable Source Language Counts in XTREME-NER

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Robustness of Cross-Lingual NER in Low-Resource Languages with Typological Divergence

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Cross-lingual NER robustness in low-resource languages with source language diversity

Multilingual Language Models (MLLMs) exhibit robust cross-lingual transfer capabilities, or the abil

**Cross-lingual Transfer Robustness in Low-Resource African NER with Script Variation**

Cross-lingual transfer learning enables NLP for low-resource languages by leveraging labeled data fr

Direct English-to-Yoruba speech Translation model using Transformer with Augmented Attention Mechanism