Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR

Domain:

natural language processing

Record type:

paperdataset
Creator:
GagGagKir
Host:avatar
Kazakh is a Turkic language using the Arabic, Cyrillic, and Latin scripts, making it unique in terms of optical character recognition (OCR). Work on OCR for low-resource Kazakh scripts is very scarce, and no OCR benchmarks or images exist for the Arabic and Latin scripts. We construct a synthetic OCR dataset of 7,219 images for all three scripts with font, color, and noise variations to imitate real OCR tasks. We evaluated three multimodal large language models (MLLMs) on a subset of the benchmark for OCR and language identification: Gemma-3-12B-it, Qwen2.5-VL-7B-Instruct, and Llama-3.2-11B-Vision-Instruct. All models are unsuccessful with Latin and Arabic script OCR, and fail to recognize the Arabic script as Kazakh text, misclassifying it as Arabic, Farsi, and Kurdish. We further compare MLLMs with a classical OCR baseline and find that while traditional OCR has lower character error rates, MLLMs fail to match this performance. These findings show significant gaps in current MLLM capabilities to process low-resource Abjad-based scripts and demonstrate the need for inclusive models and benchmarks supporting low-resource scripts and languages. Accepted to AbjadNLP @ EACL 2026

Visit

arxiv.org

Tasks

computer visionoptical character recognition

Tags

Computer Vision and Pattern RecognitionComputation and Language

Similar

ganesh045078-create/Synthetic-Document-for-Low-Resource-OCRLarge Multimodal Models for Low-Resource Languages: A SurveyMangaUB: A Manga Understanding Benchmark for Large Multimodal Models600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri scriptUhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African LanguagesMultimodal Teacher Models for Cross-Lingual NER Robustness in Low-Resource Languages

ganesh045078-create/Synthetic-Document-for-Low-Resource-OCR

**Title:** Generative AI for Synthetic Document Creation for Low-Resource OCR **Description:** In t

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

MangaUB: A Manga Understanding Benchmark for Large Multimodal Models

Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panel

600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script

This technical report presents the 600K-KS-OCR Dataset, a large-scale synthetic corpus comprising ap

Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages

Evaluations of Large Language Models (LLMs) on knowledge-intensive tasks and factual accuracy often

Multimodal Teacher Models for Cross-Lingual NER Robustness in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident