Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script

Domain:

natural language processing

Record type:

paperdataset
Creator:
Mal
Host:avatar
This technical report presents the 600K-KS-OCR Dataset, a large-scale synthetic corpus comprising approximately 602,000 word-level segmented images designed for training and evaluating optical character recognition systems targeting Kashmiri script. The dataset addresses a critical resource gap for Kashmiri, an endangered Dardic language utilizing a modified Perso-Arabic writing system spoken by approximately seven million people. Each image is rendered at 256x64 pixels with corresponding ground-truth transcriptions provided in multiple formats compatible with CRNN, TrOCR, and generalpurpose machine learning pipelines. The generation methodology incorporates three traditional Kashmiri typefaces, comprehensive data augmentation simulating real-world document degradation, and diverse background textures to enhance model robustness. The dataset is distributed across ten partitioned archives totaling approximately 10.6 GB and is released under the CC-BY-4.0 license to facilitate research in low-resource language optical character recognition.

Visit

arxiv.org

Tasks

computer visionoptical character recognition

Tags

Computer Vision and Pattern RecognitionComputation and Language

Similar

SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognitionisti-sys/Optical-Character-Recognition-OCR-Ancient but Digitized: Developing Handwritten Optical Character Recognition for East Syriac Script Through Creating KHAMIS DatasetFidel: A Large-Scale Sentence Level Amharic OCR DatasetDeep learning for Ethiopian Ge'ez script optical character recognisionPsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language

SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print

isti-sys/Optical-Character-Recognition-OCR-

Optical Character Recognition (OCR) untuk membaca teks plat nomor kendaraan Indonesia dari gambar me

Ancient but Digitized: Developing Handwritten Optical Character Recognition for East Syriac Script Through Creating KHAMIS Dataset

Many languages have vast amounts of handwritten texts, such as ancient scripts about folktale storie

Fidel: A Large-Scale Sentence Level Amharic OCR Dataset

Abstract The Ethiopic script used in the Amharic Language presents persistent chal

Deep learning for Ethiopian Ge'ez script optical character recognision

PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language

This paper evaluates the performance of Large Multimodal Models (LMMs) on Optical Character Recognit