Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Is a Generic Dataset and Foundation VLM for Arabic HTR Worth It? Lessons from AMIDDA

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
VidLucSalDec
Editor:
CenCalLabIns
Publisher:
CCSD
Host:avatar
Which handwritten text recognition (HTR) strategy works best for Arabic scripts, for which use case, and at what cost? We address this question with AMIDDA, an open dataset that compiles eight Arabic-script corpora into 66,779 real line images spanning Maghribi, Oriental, lithographed and modern scripts, with no synthetic data. On its test split we compare six strategies, from training-free in-context learning (ICL) with generalist vision-language models (VLMs) to specialised CRNN networks. Specialised CRNNs remain the most accurate option for production, while ICL reaches competitive character error rates (below 8 % on modern KHATT, 10.6 % on historical RASM) from minimal input, a few line examples or the plain transcription of one reference page. Backed by a cost evaluation, ICL stands as a credible bootstrapping mechanism: it produces first-pass annotations at scale that a human corrector can turn into the training data of a specialised model, and a Mixture-of-Experts post-correction step adds a low-cost lift when available models generalise poorly. We release the dataset, a Qwen3.5-4B LoRA foundation model covering the eight corpora, the ICL and MoE code, and a web application exposing these pipelines.

Visit

enc.hal.science

Tasks

computer visionoptical character recognition

Languages

Arabic, Moroccan Spoken

Tags

Arabic HTRVision-Language ModelsIn-Context LearningData BootstrappingDISTAM[INFO.INFO-CV]Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV]

Licenses

https://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/OpenAccess

Similar

Is linguistically-motivated data augmentation worth it?Ajami HTR DatasetA Dataset of Wolof Ajami Manuscripts for HTR and OCRCross-Lingual Learning within Arabic Script for Low-Resource HTRKBayoud/Darija-VLM-GQA-DatasetKBayoud/Darija-VLM-Dataset-Chat

Is linguistically-motivated data augmentation worth it?

Data augmentation, a widely-employed technique for addressing data scarcity, involves generating syn

Ajami HTR Dataset

This dataset contains images of Fulfulde and Hausa Ajami manuscripts and each page's polygon coordinates for segmentation (region and line), and each line's transcription.

A Dataset of Wolof Ajami Manuscripts for HTR and OCR

Cross-Lingual Learning within Arabic Script for Low-Resource HTR

Handwritten Text Recognition (HTR) with limited labeled data remains a challenging problem, particul

KBayoud/Darija-VLM-GQA-Dataset

Original dataset : vikhyatk/gqa-val

KBayoud/Darija-VLM-Dataset-Chat