A corpus of 12,890 Arabic text-line images cropped from scanned issues of theMoroccan Official Gazette (Bulletin Officiel), each paired with a manuallyverified ground-truth transcription (UTF-8 JSONL). Scan quality spans sharpprint to heavily blurred reproductions, making the corpus suitable for studyingArabic printed-text recognition under real-world degradation.
The record includes a predefined held-out test split of 1,212 lines(real_holdout_test.jsonl); the recommended protocol excludes its filenames fromtraining (leaving 11,678 lines) so that results remain comparable acrossstudies. See README.md for the record format, character inventory, and details.
Files: images.zip (12,890 PNG line crops), merged.jsonl (all transcriptions),real_holdout_test.jsonl (test split), README.md, checksums.txt.
License: CC BY 4.0.