Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Deep Contextualized Pairwise Semantic Similarity for Arabic Language Questions

Domain:

natural language processing

Record type:

paper
Creator:
Al-FarMusSee
Host:avatar
Question semantic similarity is a challenging and active research problem that is very useful in many NLP applications, such as detecting duplicate questions in community question answering platforms such as Quora. Arabic is considered to be an under-resourced language, has many dialects, and rich in morphology. Combined together, these challenges make identifying semantically similar questions in Arabic even more difficult. In this paper, we introduce a novel approach to tackle this problem, and test it on two benchmarks; one for Modern Standard Arabic (MSA), and another for the 24 major Arabic dialects. We are able to show that our new system outperforms state-of-the-art approaches by achieving 93% F1-score on the MSA benchmark and 82% on the dialectical one. This is achieved by utilizing contextualized word representations (ELMo embeddings) trained on a text corpus containing MSA and dialectic sentences. This in combination with a pairwise fine-grained similarity layer, helps our question-to-question similarity model to generalize predictions on different dialects while being trained only on question-to-question MSA data. Accepted at ICTAI 2019

Visit

arxiv.org

Tags

Computation and LanguageMachine Learning

Similar

CASS: A Comprehensive Arabic Semantic Similarity DatasetWord Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English SentencesDeep Learning Model for Answering Why-Questions in ArabicNSURL-2019 Shared Task 8: Semantic Question Similarity in ArabicSimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek languageIntegrating Lesk Algorithm with Cosine Semantic Similarity to Resolve Polysemy for Setswana Language

CASS: A Comprehensive Arabic Semantic Similarity Dataset

The Comprehensive Arabic Semantic Similarity (CASS) dataset is a large-sc

Word Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English Sentences

International audience Semantic Textual Similarity (STS) is an important component in

Deep Learning Model for Answering Why-Questions in Arabic

The subfield of natural language processing (NLP) known as question answering (QA) involves providin

NSURL-2019 Shared Task 8: Semantic Question Similarity in Arabic

Question semantic similarity (Q2Q) is a challenging task that is very useful in many NLP application

SimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek language

Semantic relatedness between words is one of the core concepts in natural language processing, thus

Integrating Lesk Algorithm with Cosine Semantic Similarity to Resolve Polysemy for Setswana Language