
The Comprehensive Arabic Semantic Similarity (CASS) dataset is a large-scale resource for Arabic Semantic Textual Similarity (STS). It comprises 3,048 manually annotated Modern Standard Arabic (MSA) sentence pairs with fine-grained similarity scores (0–5), spanning six semantic categories (geography, history, law, sports, health, and essay-style prose) and 42 subcategories capturing targeted linguistic phenomena (morphological variation, syntactic transformation, lexical substitution, negation, temporal and spatial modification, and entity variation).
CASS is four times larger than existing Arabic STS datasets and provides structured taxonomic coverage supporting systematic model evaluation. Each anchor sentence is paired with at least six labeled variants spanning the full similarity continuum. Annotation was performed by three native Arabic speakers with backgrounds in linguistics and computational linguistics, with an inter-rater reliability of Krippendorff's alpha = 0.82 (substantial agreement).
This dataset accompanies the paper "CASS: A Comprehensive Arabic Semantic Similarity Dataset with LLMs Systematic Evaluation" and establishes essential infrastructure for Arabic STS research and applications in education, legal technology, and content moderation.