Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study

Domain:

natural language processing

Record type:

paper
Creator:
FarAlaMaaMaa
Publisher:
arXiv
Host:avatar
Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remains underexplored. To address such a gap in this paper, we evaluate traditional OCR, general-purpose VLMs, Arabic-specialized VLMs, and OCR-conditioned VLM correction across eight Arabic text datasets spanning historical manuscripts, aged printed books, clean print, multi-domain documents, and handwriting. The results show that no single approach dominates across setups. On line-level historical manuscripts, VLMs are close to Tesseract; on page-level manuscript images, they perform better; and in several settings, an OCR-conditioned corrector improves over both standalone OCR and standalone VLMs. The central finding is an OCR-prior recoverability principle: OCR conditioning helps when the OCR output remains visually and textually recoverable, providing anchors that the VLM can refine against the image. It improves recognition on aged print, clean print, mixed-domain Arabic, and some Naskh manuscripts, but degrades performance when the prior is script-mismatched or systematically misleading, as in Maghribi manuscripts and realistic student handwriting. Additional diagnostics show that Arabic VLM-OCR is sensitive to diacritics, preprocessing, generation budget, and repetition loops. These findings support an adaptive OCR-VLM workflow that routes pages according to script, OCR-prior recoverability, length diagnostics, and failure-mode indicators.

Visit

doi.org

Languages

Arabic, Moroccan Spoken

Tags

Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text RecognitionWhen do they seek help and where do they go? Health-related help-seeking behaviors and health service utilization among older adults in rural GhanaWhen Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism DetectionTahaLamhandi/Arabic-Darija-OCR-Systemtaqacuct/kab-ocr-datasetmarconilabmak/luganda-ocr-dataset

SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print

When do they seek help and where do they go? Health-related help-seeking behaviors and health service utilization among older adults in rural Ghana

Although many scholars have researched topics related to older Ghanaians, few recent studies have in

When Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism Detection

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

TahaLamhandi/Arabic-Darija-OCR-System

Ce projet propose une solution complète d'OCR (Reconnaissance Optique de Caractères) spécialement co

taqacuct/kab-ocr-dataset

Dataset synthétique pour entraîner un modèle OCR sur le kabyle. images/ : contiendra les images gén

marconilabmak/luganda-ocr-dataset

This dataset contains segmented line-level images and corresponding transcriptions in Luganda, a low