A diagnostic dataset of ancient Egyptian hieratic script from Deir el-Medina with a focus on paleography and document analysis.
Description
This dataset was created as part of the Crossing Boundaries work on the Deir el-Medina papyri and fragments, which are now housed at the Museo Egizio in Turin, Italy.
The dataset was primarily compiled for the annotation of selected characters and groups of characters. The emphasis was placed on visual units that frequently occur in Egyptian texts and can therefore be used for palaeographic analysis, as well as units indicative of specific text genres. For this reason, this dataset does not contain complete character annotations.
The annotations for this dataset were generated using LabelMe by Kentaro Wada. A total of 159 images from 50 different papyrus documents were analysed and 504 categories of characters or character groups were outlined using polygon boxes. This dataset contains the resulting image croppings and complete annotations for each papyrus. The annotations are available as either JSON or YOLO files.
Content and Structure
The classes.json, samples.json and papyri.json files contain detailed information about the various class labels, every sample in the dataset (named after a running sample number) and the annotated papyri (named after a running image number). All files are structured JSON files containing more information on exact documents, copyrights and reference links, as well as additional information that might help with the algorithmic processing and grouping of the data.
DDD_annotations.zip contains annotation data for the papyri. There is one folder which contains the original annotations in LabelMe format, each connected to a corresponding papyrus image. And another folder which contains manual maskings to remove unannotated areas from character detection tasks.
DDD_images.zip contains the full set of images produced with this dataset: a folder containing all cut-out samples as rectangles; a folder containing all samples as cut-out polygons; a folder containing all (publishable) full papyri images; and a folder containing all (publishable) masked papyri images.
Finally, the Split X-X.zip files contain all data necessary to work with the proposed splits. Each folder has a structured JSON file which lists all documents/images and samples for training, validation and test sets. The naming convention is as follows:
C-A: Closed-set Recognition, All Documents, Document-aware.
C-B: Closed-set Recognition, Published Documents, Document-aware.
C-C: Closed-set Recognition, All Documents, Cross-document.
C-D: Closed-set Recognition, Published Documents, Cross-document.
D-B: Character Detection, Published Documents, Document-aware.
O-A: Open-set Recognition, All Documents, Document-aware.
O-B: Open-set Recognition, Published Documents, Document-aware.
We do not provide data for the D-A: Character Detection on all documents case, as we are not able to provide all necessary source files.
There are two additional PDF files (DDD - Dataset Numbers.pdf and DDD - Label Overview.pdf) for detailed dataset numbers and an overview of class labels with hieratic examples and hieroglyphic transcriptions to make the dataset more accessible.
Retrospective Changes
At the time of publication, this dataset was already serving as the basis for some academic papers. Shortly before publication, a further round of corrections resulted in changes including the merging of classes and the redefinition of visual complexities. Although only a very small number of samples are affected, we would like to transparently disclose these changes here:
The labels hAj and hAj_name were merged to hAj_name.
The labels Hwtj and Hwtjw were merged to Hwtjw.
The label TAw has been renamed to P5; it has also been reclassified from Small Group to Character/Number.
Copyright and Licence
All full papyrus images are courtesy of the Museo Egizio in Turin, Italy. Please note that in the papyri.json file you will find exact copyright information per image. You will also find direct links to the documents in the Turin Papyrus Online Platform (TPOP) (free registration required).
The dataset is published under a CC 4.0 BY-SA-NC licence.