Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Concise Survey of OCR for Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
AssAgaAnastasopoulos, Antonios
Publisher:
Und
Host:avatar
Modern natural language processing (NLP) techniques increasingly require substantial amounts of data to train robust algorithms. Building such technologies for low-resource languages requires focusing on data creation efforts and data-efficient algorithms. For a large number of low-resource languages, especially Indigenous languages of the Americas, this data exists in image-based non-machine-readable documents. This includes scanned copies of comprehensive dictionaries, linguistic field notes, children's stories, and other textual material. To digitize these resources, Optical Character Recognition (OCR) has played a major role but it comes with certain challenges in low-resource settings. In this paper, we share the first survey of OCR techniques specific to low-resource data creation settings and outline several open challenges, with a special focus on Indigenous Languages of the Americas. Based on experiences and results from previous research, we conclude with recommendations on utilizing and improving OCR for the benefit of computational researchers, linguists, and language communities.

Visit

doi.orgunderline.io

Tasks

computer visionoptical character recognition

Tags

Computational LinguisticsNatural Language ProcessingMachine translationArtificial Intelligence

Similar

Neural Machine Translation for Low-Resource Languages: A SurveyLarge Multimodal Models for Low-Resource Languages: A SurveyNaijaNLP: A Survey of Nigerian Low-Resource LanguagesYayQare/low-resource-ocr-profilesganesh045078-create/Synthetic-Document-for-Low-Resource-OCRMonolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey

Neural Machine Translation for Low-Resource Languages: A Survey

Neural Machine Translation (NMT) has seen a tremendous spurt of growth in less than ten years, and h

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

NaijaNLP: A Survey of Nigerian Low-Resource Languages

With over 500 languages in Nigeria, three languages -- Hausa, Yorùbá and Igbo -- spoken by over 175

YayQare/low-resource-ocr-profiles

Methodology for building custom ABBYY FineReader recognition languages for low-resource scripts # C

ganesh045078-create/Synthetic-Document-for-Low-Resource-OCR

**Title:** Generative AI for Synthetic Document Creation for Low-Resource OCR **Description:** In t

Monolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a signi