This dataset contains Ethiopian legal-document page images paired with OCR ground-truth text.
It is prepared for training and evaluating OCR and vision-language transcription models.
Total examples: 1,534
Train split: 1,381
Validation split: 153
Storage format: Parquet (viewer-safe image struct)
train: 1,381 rows
validation: 153 rows
Columns