A hierarchically annotated dataset for the segmentation of ancient Arabic manuscripts at four levels: lines, words, pseudo-words (Pieces of Arabic Word), and characters. The dataset supports research on the automatic analysis and segmentation of historical Arabic handwriting, addressing the scarcity of multi-level, openly available resources for ancient (as opposed to modern) Arabic script.
Annotations are provided in COCO format (bounding boxes and polygon segmentation masks). Lines are annotated on full manuscript pages; words, pseudo-words, and characters are annotated on line crops. The character level covers 29 classes corresponding to the letters of the Arabic alphabet. The train/validation/test split is performed at the page level (approximately 70/15/15) so that no page appears in more than one split, preventing information leakage.
Manuscript images were collected from openly accessible digital libraries, including Gallica (Bibliothèque nationale de France), the Qatar Digital Library, the National Library of the Kingdom of Morocco, and the King Abdulaziz Public Library, covering several calligraphic styles (Naskh, Maghribi, Kufi) and document genres.