Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A comparative study of root-based and stem-based approaches for measuring the similarity between arabic words for arabic text mining applications

Domain:

natural language processing

Record type:

paperdataset
Creator:
FroLacOua
Host:avatar
Representation of semantic information contained in the words is needed for any Arabic Text Mining applications. More precisely, the purpose is to better take into account the semantic dependencies between words expressed by the co-occurrence frequencies of these words. There have been many proposals to compute similarities between words based on their distributions in contexts. In this paper, we compare and contrast the effect of two preprocessing techniques applied to Arabic corpus: Rootbased (Stemming), and Stem-based (Light Stemming) approaches for measuring the similarity between Arabic words with the well known abstractive model -Latent Semantic Analysis (LSA)- with a wide variety of distance functions and similarity measures, such as the Euclidean Distance, Cosine Similarity, Jaccard Coefficient, and the Pearson Correlation Coefficient. The obtained results show that, on the one hand, the variety of the corpus produces more accurate results; on the other hand, the Stem-based approach outperformed the Root-based one because this latter affects the words meanings.

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and LanguageInformation Retrieval

Similar

Word Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English SentencesA Comparative Analysis of CNN and RNN Architectures for Deep Learning-Based Arabic Text ClassificationShort Text Topic Modeling for Moroccan Darija Online Content: A Comparative Study Between Classical And Embedding-based ApproachesPre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative StudyDistance-Based Meta-Features for Arabic Text ClassificationQuery Expansion Based-on Similarity of Terms for Improving Arabic Information Retrieval

Word Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English Sentences

International audience Semantic Textual Similarity (STS) is an important component in

A Comparative Analysis of CNN and RNN Architectures for Deep Learning-Based Arabic Text Classification

The proliferation of digital Arabic content has created a pressing need for efficient text classifi

Short Text Topic Modeling for Moroccan Darija Online Content: A Comparative Study Between Classical And Embedding-based Approaches

Pre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative Study

Question answering(QA) is one of the most challenging yet widely investigated problems in Natural La

Distance-Based Meta-Features for Arabic Text Classification

Query Expansion Based-on Similarity of Terms for Improving Arabic Information Retrieval

Part 6: Information Retrieval International audience This research suggests a method