Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Evaluation of feature-embedding methods for word spotting in historical arabic documents

Domaine:

natural language processing

Type de record:

paper
Créateur:
FatIbnEl Ess
Éditeur:
InsDépInsARM
Éditeur:
CCSDIEE
Hôte:avatar
International audience Retrieving and indexing historical Arabic documents remain a very significant challenge. The purpose of this paper is to compare the feature representation spaces for word spotting in historical Arabic documents. Our goal is to create embedding spaces using the characteristics of different machine learning methods: i) linear such as principal component analysis and linear discriminant analysis, and ii) non-linear including convolutional neural networks for triplets and Siamese. Subsequently, each word image is represented by a dense vector. Thus, to match feature representations, a Euclidean distance is used. An evaluation of various representation space models is presented. The embedding word models are evaluated on the VML-HD dataset, and the experiments show the effectiveness of non-linear methods compared to linear ones.

Visit

hal.science

Tags

Historical Arabic documentsWord spottingFeature embedding[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI][INFO]Computer Science [cs]