Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Word Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English Sentences

Domain:

natural language processing

Record type:

paper
Creator:
BilFerSchChe
Editor:
GroLabComUni
Publisher:
CCSD
Host:avatar
International audience Semantic Textual Similarity (STS) is an important component in many Natural Language Processing (NLP) applications, and plays an important role in diverse areas such as information retrieval, machine translation, information extraction and plagiarism detection. In this paper we propose two word embedding-based approaches devoted to measuring the semantic similarity between Arabic-English cross-language sentences. The main idea is to exploit Machine Translation (MT) and an improved word embedding representations in order to capture the syntactic and semantic properties of words. MT is used to translate English sentences into Arabic language in order to apply a classical monolingual comparison. Afterwards, two word embedding-based methods are developed to rate the semantic similarity. Additionally, Words Alignment (WA), Inverse Document Frequency (IDF) and Part-of-Speech (POS) weighting are applied on the examined sentences to support the identification of words that are most descriptive in each sentence. The performances of our approaches are evaluated on a cross-language dataset containing more than 2400 Arabic-English pairs of sentence. Moreover, the proposed methods are confirmed through the Pearson correlation between our similarity scores and human ratings.

Visit

hal.science

Tasks

embeddingsmachine translation

Tags

Semantic Sentences SimilarityMachine TranslationWord EmbeddingsCross-LanguageArabic-English[INFO]Computer Science [cs][INFO.INFO-TT]Computer Science [cs]/Document and Text Processing[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

A comparative study of root-based and stem-based approaches for measuring the similarity between arabic words for arabic text mining applicationsGATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss TrainingDeep Contextualized Pairwise Semantic Similarity for Arabic Language QuestionsSinhala Similarity and Semantic Plagiarism Detection Using Modern Embedding ModelsCASS: A Comprehensive Arabic Semantic Similarity DatasetUnsupervised Cross-lingual Word Embedding Representation for English-isiZulu

A comparative study of root-based and stem-based approaches for measuring the similarity between arabic words for arabic text mining applications

Representation of semantic information contained in the words is needed for any Arabic Text Mining a

GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training

Semantic textual similarity (STS) is a critical task in natural language processing (NLP), enabling

Deep Contextualized Pairwise Semantic Similarity for Arabic Language Questions

Question semantic similarity is a challenging and active research problem that is very useful in man

Sinhala Similarity and Semantic Plagiarism Detection Using Modern Embedding Models

Abstract While plagiarism detection has developed to some degree in English and in some high-resour

CASS: A Comprehensive Arabic Semantic Similarity Dataset

The Comprehensive Arabic Semantic Similarity (CASS) dataset is a large-sc

Unsupervised Cross-lingual Word Embedding Representation for English-isiZulu

In this study, we investigate the effectiveness of using cross-lingual word embeddings for zero-shot