Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Introducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
NguWalVu,
Hôte:avatar
We present two novel datasets for the low-resource language Vietnamese to assess models of semantic similarity: ViCon comprises pairs of synonyms and antonyms across word classes, thus offering data to distinguish between similarity and dissimilarity. ViSim-400 provides degrees of similarity across five semantic relations, as rated by human judges. The two datasets are verified through standard co-occurrence and neural network models, showing results comparable to the respective English datasets. The 16th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT 2018)

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and Language

Similaires

Introducing various Semantic Models for Amharic: Experimentation and Evaluation with multiple Tasks and DatasetsAfrican Wordnet as a tool to identify semantic relatedness and semantic similaritySimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek languageSemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 LanguagesNeural Models for Detecting Binary Semantic Textual Similarity for Algerian and MSASinhala Similarity and Semantic Plagiarism Detection Using Modern Embedding Models

Introducing various Semantic Models for Amharic: Experimentation and Evaluation with multiple Tasks and Datasets

The availability of different pre-trained semantic models enabled the quick development of machine learning components for downstream applications. Despite the availability of abundant text data for low resource languages, only a few semantic models are publicly av

African Wordnet as a tool to identify semantic relatedness and semantic similarity

SimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek language

Semantic relatedness between words is one of the core concepts in natural language processing, thus

SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages

Exploring and quantifying semantic relatedness is central to representing language and holds signifi

Neural Models for Detecting Binary Semantic Textual Similarity for Algerian and MSA

Sinhala Similarity and Semantic Plagiarism Detection Using Modern Embedding Models

Abstract While plagiarism detection has developed to some degree in English and in some high-resour