Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Hierarchical Dataset of Ancient Arabic Manuscripts (HAAM)

Domain:

natural language processing

Record type:

dataset
Creator:
Bou
Publisher:
Zenodo
Host:avatar
A hierarchically annotated dataset for the segmentation of ancient Arabic manuscripts at four levels: lines, words, pseudo-words (Pieces of Arabic Word), and characters. The dataset supports research on the automatic analysis and segmentation of historical Arabic handwriting, addressing the scarcity of multi-level, openly available resources for ancient (as opposed to modern) Arabic script. Annotations are provided in COCO format (bounding boxes and polygon segmentation masks). Lines are annotated on full manuscript pages; words, pseudo-words, and characters are annotated on line crops. The character level covers 29 classes corresponding to the letters of the Arabic alphabet. The train/validation/test split is performed at the page level (approximately 70/15/15) so that no page appears in more than one split, preventing information leakage. Manuscript images were collected from openly accessible digital libraries, including Gallica (Bibliothèque nationale de France), the Qatar Digital Library, the National Library of the Kingdom of Morocco, and the King Abdulaziz Public Library, covering several calligraphic styles (Naskh, Maghribi, Kufi) and document genres.

Visit

doi.org

Tasks

optical character recognitioncomputer vision

Languages

Arabic, Moroccan Spoken

Tags

Arabic manuscriptshandwritten text recognitiondocument image analysishierarchical segmentationinstance segmentationCOCO datasetdeep learningOCRhistorical documentsancient Arabic script

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

A Text Recognition Dataset from Sahidic Coptic Ancient ManuscriptsSCAM – Sahidic Coptic Ancient ManuscriptsTHE PRESERVATION OF ANCIENT ARABIC MANUSCRIPTS: A REFLECTION ON SOME SELECTED PUBLIC REPOSITORIES IN NORTHERN NIGERIAback-kh/SADA-Ancient-Palm-Leaf-Manuscripts-RecognitionsCharacter recognition of ancient ethiopic Ge'ez manuscripts using deep convolutional neural networksTASHELHIYT BERBER MANUSCRIPTS IN ARABIC CHARACTERS

A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

In this work, we target Handwritten Text Recognition (HTR) in low-resource scenarios, which arise fr

SCAM – Sahidic Coptic Ancient Manuscripts

Project page: scam.silviacascianelli.com SCAM is a dataset of handwritten Sahidic Coptic manuscript

THE PRESERVATION OF ANCIENT ARABIC MANUSCRIPTS: A REFLECTION ON SOME SELECTED PUBLIC REPOSITORIES IN NORTHERN NIGERIA

The three repositories covered by this paper, the Arewa House Arabic Manuscript repository, the Nati

back-kh/SADA-Ancient-Palm-Leaf-Manuscripts-Recognitions

[PRL 2025, APSIPA 2022] Syllable Analysis Data Augmentation (SADA), This project introduces a glyph

Character recognition of ancient ethiopic Ge'ez manuscripts using deep convolutional neural networks

TASHELHIYT BERBER MANUSCRIPTS IN ARABIC CHARACTERS