Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CASS: A Comprehensive Arabic Semantic Similarity Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
KhrJam
Editor:
HacDia
Publisher:
Zenodo
Host:avatar

The Comprehensive Arabic Semantic Similarity (CASS) dataset is a large-scale resource for Arabic Semantic Textual Similarity (STS). It comprises 3,048 manually annotated Modern Standard Arabic (MSA) sentence pairs with fine-grained similarity scores (0–5), spanning six semantic categories (geography, history, law, sports, health, and essay-style prose) and 42 subcategories capturing targeted linguistic phenomena (morphological variation, syntactic transformation, lexical substitution, negation, temporal and spatial modification, and entity variation).

CASS is four times larger than existing Arabic STS datasets and provides structured taxonomic coverage supporting systematic model evaluation. Each anchor sentence is paired with at least six labeled variants spanning the full similarity continuum. Annotation was performed by three native Arabic speakers with backgrounds in linguistics and computational linguistics, with an inter-rater reliability of Krippendorff's alpha = 0.82 (substantial agreement).

This dataset accompanies the paper "CASS: A Comprehensive Arabic Semantic Similarity Dataset with LLMs Systematic Evaluation" and establishes essential infrastructure for Arabic STS research and applications in education, legal technology, and content moderation.

Visit

doi.org

Tasks

embeddings

Languages

Ndasa

Tags

Arabic Natural Language ProcessingSemantic Textual SimilarityBenchmark DatasetsModel EvaluationCross-lingual Transfer LearningArabic Language Resources

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Deep Contextualized Pairwise Semantic Similarity for Arabic Language QuestionsNSURL-2019 Shared Task 8: Semantic Question Similarity in ArabicSimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek languageWord Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English SentencesArabicaQA: A Comprehensive Dataset for Arabic Question AnsweringThe Inception Team at NSURL-2019 Task 8: Semantic Question Similarity in Arabic

Deep Contextualized Pairwise Semantic Similarity for Arabic Language Questions

Question semantic similarity is a challenging and active research problem that is very useful in man

NSURL-2019 Shared Task 8: Semantic Question Similarity in Arabic

Question semantic similarity (Q2Q) is a challenging task that is very useful in many NLP application

SimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek language

Semantic relatedness between words is one of the core concepts in natural language processing, thus

Word Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English Sentences

International audience Semantic Textual Similarity (STS) is an important component in

ArabicaQA: A Comprehensive Dataset for Arabic Question Answering

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources

The Inception Team at NSURL-2019 Task 8: Semantic Question Similarity in Arabic

This paper describes our method for the task of Semantic Question Similarity in Arabic in the worksh