Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek language

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SalKurGóm
Hôte:avatar
Semantic relatedness between words is one of the core concepts in natural language processing, thus making semantic evaluation an important task. In this paper, we present a semantic model evaluation dataset: SimRelUz - a collection of similarity and relatedness scores of word pairs for the low-resource Uzbek language. The dataset consists of more than a thousand pairs of words carefully selected based on their morphological features, occurrence frequency, semantic relation, as well as annotated by eleven native Uzbek speakers from different age groups and gender. We also paid attention to the problem of dealing with rare words and out-of-vocabulary words to thoroughly evaluate the robustness of semantic models. Final version, published in the proceedings of SIGUL workshop of LREC 2022

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and Language

Similaires

African Wordnet as a tool to identify semantic relatedness and semantic similarityIntroducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and RelatednessA SEMI-SUPERVISED FRAMEWORK NAMED AUGSBERT-UZ FOR HIGH-PERFORMANCE SEMANTIC TEXTUAL SIMILARITY IN UZBEKCASS: A Comprehensive Arabic Semantic Similarity DatasetDeep Contextualized Pairwise Semantic Similarity for Arabic Language QuestionsMulti-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic Similarity

African Wordnet as a tool to identify semantic relatedness and semantic similarity

Introducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness

We present two novel datasets for the low-resource language Vietnamese to assess models of semantic

A SEMI-SUPERVISED FRAMEWORK NAMED AUGSBERT-UZ FOR HIGH-PERFORMANCE SEMANTIC TEXTUAL SIMILARITY IN UZBEK

Semantic Textual Similarity (STS) is one of the fundamental task of Natural Language Processing (NLP

CASS: A Comprehensive Arabic Semantic Similarity Dataset

The Comprehensive Arabic Semantic Similarity (CASS) dataset is a large-sc

Deep Contextualized Pairwise Semantic Similarity for Arabic Language Questions

Question semantic similarity is a challenging and active research problem that is very useful in man

Multi-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic Similarity

We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data set