Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Exploiting Script Similarities to Compensate for the Large Amount of Data in Training Tesseract LSTM: Towards Kurdish OCR

Domaine:

natural language processing

Type de record:

paper
Créateur:
SamHos
Éditeur:
MDP
Hôte:
Applications based on Long-Short-Term Memory (LSTM) require large amounts of data for their training. Tesseract LSTM is a popular Optical Character Recognition (OCR) engine that has been trained and used in various languages. However, its training becomes obstructed when the target language is not resourceful. This research suggests a remedy for the problem of scant data in training Tesseract LSTM for a new language by exploiting a training dataset for a language with a similar script. The target of the experiment is Kurdish. It is a multi-dialect language and is considered less-resourced. We choose Sorani, one of the Kurdish dialects, that is mostly written in Persian-Arabic script. We train Tesseract using an Arabic dataset, and then we use a considerably small amount of texts in Persian-Arabic to train the engine to recognize Sorani texts. Our dataset is based on a series of court case documents in the Kurdistan Region of Iraq. We also fine-tune the engine using 10 Unikurd fonts. We use Lstmeval and Ocreval to evaluate the outputs. The result indicates the achievement of 95.45% accuracy. We also test the engine using texts outside the context of court cases. The accuracy of the system remains close to what was found earlier indicating that the script similarity could be used to overcome the lack of large-scale data.

Visit

doi.org

Tasks

optical character recognitioncomputer vision

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

AitBAD/kab-Taqbaylit-Tesseract-ocr600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri scriptzacharyb02/OCR-tifinagh-scriptKHLD: A Large-Scale Benchmark of the Kurdish Handwritten Lines Dataset for Low-Resource Central Kurdish (Sorani)Exploiting cross-linguistic similarities in Zulu and Xhosa computational morphologyKurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech

AitBAD/kab-Taqbaylit-Tesseract-ocr

600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script

This technical report presents the 600K-KS-OCR Dataset, a large-scale synthetic corpus comprising ap

zacharyb02/OCR-tifinagh-script

CNN-based Optical Character Recognition system specifically designed for the Tifinagh script, used f

KHLD: A Large-Scale Benchmark of the Kurdish Handwritten Lines Dataset for Low-Resource Central Kurdish (Sorani)

The Kurdish Handwritten Lines Dataset (KHLD) is a large-scale image dataset aiming to facilitate han

Exploiting cross-linguistic similarities in Zulu and Xhosa computational morphology

KurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech

KurFemTTS is a collaborative initiative between the University of Kurdistan Hewlêr (UKH) and Kurdish