Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Concise Survey of OCR for Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssAgaAnastasopoulos, Antonios
Éditeur:
Und
Hôte:avatar
Modern natural language processing (NLP) techniques increasingly require substantial amounts of data to train robust algorithms. Building such technologies for low-resource languages requires focusing on data creation efforts and data-efficient algorithms. For a large number of low-resource languages, especially Indigenous languages of the Americas, this data exists in image-based non-machine-readable documents. This includes scanned copies of comprehensive dictionaries, linguistic field notes, children's stories, and other textual material. To digitize these resources, Optical Character Recognition (OCR) has played a major role but it comes with certain challenges in low-resource settings. In this paper, we share the first survey of OCR techniques specific to low-resource data creation settings and outline several open challenges, with a special focus on Indigenous Languages of the Americas. Based on experiences and results from previous research, we conclude with recommendations on utilizing and improving OCR for the benefit of computational researchers, linguists, and language communities.

Visit

doi.orgunderline.io

Tasks

computer visionoptical character recognition

Tags

Computational LinguisticsNatural Language ProcessingMachine translationArtificial Intelligence

Similaires

Neural Machine Translation for Low-Resource Languages: A SurveyLarge Multimodal Models for Low-Resource Languages: A SurveyNaijaNLP: A Survey of Nigerian Low-Resource LanguagesYayQare/low-resource-ocr-profilesganesh045078-create/Synthetic-Document-for-Low-Resource-OCRMonolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey

Neural Machine Translation for Low-Resource Languages: A Survey

Neural Machine Translation (NMT) has seen a tremendous spurt of growth in less than ten years, and h

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

NaijaNLP: A Survey of Nigerian Low-Resource Languages

With over 500 languages in Nigeria, three languages -- Hausa, Yorùbá and Igbo -- spoken by over 175

YayQare/low-resource-ocr-profiles

Methodology for building custom ABBYY FineReader recognition languages for low-resource scripts # C

ganesh045078-create/Synthetic-Document-for-Low-Resource-OCR

**Title:** Generative AI for Synthetic Document Creation for Low-Resource OCR **Description:** In t

Monolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a signi