Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sindhi OCR Dataset: Multi-Font Machine-Printed Text Line Images with Ground Truth for Low-Resource Language Recognition

Domaine:

natural language processing

Type de record:

dataset
Créateur:
GoeLeh
Éditeur:
LehHak
Éditeur:
Zenodo
Hôte:avatar
This dataset has 60,122 binary TIF text line images across different Sindhi fonts. Ground truth provided as UTF-8 encoded .txt files and  .xml metadata files. Train/test splits  provided as split files for reproducibility. For academic and non-commercial research use only.

Visit

doi.orgzenodo.org

Tasks

computer visionoptical character recognition

Tags

Sindhi OCROptical Character RecognitionCursive ScriptArabic ScriptLow-resource LanguageText Line DatasetGround TruthMulti-FontMDLSTMDocument Digitization

Licenses

Creative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcodeCopyright © 2025 Shanky Goel and Gurpreet Singh Lehal, Chitkara University and IIIT Hyderabad. This dataset is made available for academic and non-commercial research use only. Any use of this dataset must cite the associated paper: "Comprehensive Multi-Font Sindhi OCR: A Fully Integrated Machine-Printed System", ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 2025.http://rightsstatements.org/vocab/InC/1.0/

Similaires

amharic-ocr-ground-truthMulti-font Printed Amharic Character Image Recognition: Deep Learning TechniquesGLOCR: GeezLab OCR Dataset Tigrinya Text Recognition DatasetMoroccan Official Gazette (Bulletin Officiel) — Transcribed Arabic Text-Line Images for OCRjoearid/Amharic-OCR-CRNN-CTC-for-Printed-Ge-ez-Script-RecognitionIndexing by recognition approach for printed Arabic documents images

amharic-ocr-ground-truth

Multi-font Printed Amharic Character Image Recognition: Deep Learning Techniques

GLOCR: GeezLab OCR Dataset Tigrinya Text Recognition Dataset

A Text Recognition (TR) and Optical Character Recognition (OCR) dataset for the Tigrinya language. T

Moroccan Official Gazette (Bulletin Officiel) — Transcribed Arabic Text-Line Images for OCR

A corpus of 12,890 Arabic text-line images cropped from scanned issues of theMoroccan Official Gazet

joearid/Amharic-OCR-CRNN-CTC-for-Printed-Ge-ez-Script-Recognition

Amharic OCR: CRNN-CTC pipeline for printed Ge'ez script. 96.3% accuracy, 0.64% CER. CNN+BiLSTM+CTC w

Indexing by recognition approach for printed Arabic documents images