Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Sindhi OCR Dataset: Multi-Font Machine-Printed Text Line Images with Ground Truth for Low-Resource Language Recognition

Domain:

natural language processing

Record type:

dataset
Creator:
GoeLeh
Editor:
LehHak
Publisher:
Zenodo
Host:avatar
This dataset has 60,122 binary TIF text line images across different Sindhi fonts. Ground truth provided as UTF-8 encoded .txt files and  .xml metadata files. Train/test splits  provided as split files for reproducibility. For academic and non-commercial research use only.

Visit

doi.orgzenodo.org

Tasks

computer visionoptical character recognition

Tags

Sindhi OCROptical Character RecognitionCursive ScriptArabic ScriptLow-resource LanguageText Line DatasetGround TruthMulti-FontMDLSTMDocument Digitization

Licenses

Creative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcodeCopyright © 2025 Shanky Goel and Gurpreet Singh Lehal, Chitkara University and IIIT Hyderabad. This dataset is made available for academic and non-commercial research use only. Any use of this dataset must cite the associated paper: "Comprehensive Multi-Font Sindhi OCR: A Fully Integrated Machine-Printed System", ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 2025.http://rightsstatements.org/vocab/InC/1.0/

Similar

amharic-ocr-ground-truthMulti-font Printed Amharic Character Image Recognition: Deep Learning TechniquesGLOCR: GeezLab OCR Dataset Tigrinya Text Recognition DatasetMoroccan Official Gazette (Bulletin Officiel) — Transcribed Arabic Text-Line Images for OCRjoearid/Amharic-OCR-CRNN-CTC-for-Printed-Ge-ez-Script-RecognitionIndexing by recognition approach for printed Arabic documents images

amharic-ocr-ground-truth

Multi-font Printed Amharic Character Image Recognition: Deep Learning Techniques

GLOCR: GeezLab OCR Dataset Tigrinya Text Recognition Dataset

A Text Recognition (TR) and Optical Character Recognition (OCR) dataset for the Tigrinya language. T

Moroccan Official Gazette (Bulletin Officiel) — Transcribed Arabic Text-Line Images for OCR

A corpus of 12,890 Arabic text-line images cropped from scanned issues of theMoroccan Official Gazet

joearid/Amharic-OCR-CRNN-CTC-for-Printed-Ge-ez-Script-Recognition

Amharic OCR: CRNN-CTC pipeline for printed Ge'ez script. 96.3% accuracy, 0.64% CER. CNN+BiLSTM+CTC w

Indexing by recognition approach for printed Arabic documents images