Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

UzThemeLex Dataset: An Uzbek Thematic Lexicon for Domain Terminology and Weakly Supervised NER

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Sai
Éditeur:
NovUrg
Éditeur:
Men
Hôte:avatar
UzThemeLex is a curated Uzbek-language thematic lexicon dataset designed for domain terminology mining and weakly supervised named entity recognition (NER). The release contains 4,945 unique terminological entries organized into 3 top-level domains (Agronomy, Economics and Business, Law and Governance) and 30 subcategories. Each entry provides the Uzbek term in Latin script, a normalized form for matching, a paraphrased Uzbek definition, domain and subcategory labels, provenance pointers to authoritative sources, and lightweight quality-control signals (heuristic confidence score, review flag, ambiguity flag). Optional fields include aliases and example sentences. The dataset is distributed in multiple formats to support both manual inspection and machine processing. It includes a flat CSV file and a multi-sheet Excel workbook, together with a data dictionary that documents all columns and label sets. For training and pipeline integration, the release also provides JSON/JSONL exports, taxonomy metadata, and ready-to-use pattern files for dictionary-based tagging and weak supervision (e.g., spaCy EntityRuler patterns). A validation script is included to help users verify schema consistency and detect formatting issues (e.g., residual Cyrillic characters and apostrophe normalization). UzThemeLex can be used as (i) a domain dictionary for keyword-based classification and information extraction in Uzbek texts and (ii) a gazetteer for generating weak labels to train or fine-tune NER models. The resource is intended to support Uzbek NLP research and applied text analytics in agriculture, economics, and legal/governance domains.

Visit

doi.orgdata.mendeley.com

Tasks

information extractionnamed entity recognition

Tags

Computational LinguisticsNatural Language ProcessingInformation ExtractionLexicographyText MiningLow-Resource LLM

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

An Uzbek Medical-Domain Dataset for Aspect-Based Sentiment AnalysisWeakly Supervised Domain Adaptation for Built-up Region Segmentation in Aerial and Satellite ImageryA Weakly Supervised Dataset of Fine-Grained Emotions in PortugueseWARM: A Weakly (+Semi) Supervised Model for Solving Math word ProblemsWeakly Supervised Whole Cardiac Segmentation via Attentional CNNUzEDSA: Uzbek Emotion and Sentiment Analysis Dataset

An Uzbek Medical-Domain Dataset for Aspect-Based Sentiment Analysis

UzMedSentiment is a manually annotated Uzbek medical-domain dataset designed for sentiment classific

Weakly Supervised Domain Adaptation for Built-up Region Segmentation in Aerial and Satellite Imagery

This paper proposes a novel domain adaptation algorithm to handle the challenges posed by the satell

A Weakly Supervised Dataset of Fine-Grained Emotions in Portuguese

Affective Computing is the study of how computers can recognize, interpret and simulate human affect

WARM: A Weakly (+Semi) Supervised Model for Solving Math word Problems

Solving math word problems (MWPs) is an important and challenging problem in natural language proces

Weakly Supervised Whole Cardiac Segmentation via Attentional CNN

Part 2: Machine Learning International audience Whole-heart segmentation aims to deli

UzEDSA: Uzbek Emotion and Sentiment Analysis Dataset

UzEDSA (Uzbek Emotion and Sentiment Analysis Dataset) is an open-access dataset designed for emotion