Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

LIACC/Emakhuwa-Portuguese-OCR-post-correction

Domain:

natural language processing

Record type:

dataset
Creator:
LIA
Host:
BibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung",

Visit

huggingface.co

Tasks

computer visionoptical character recognition

Languages

MakhuwaMakhuwa-Meetto

Licenses

cc-by-4.0

Similar

LIACC/Emakhuwa-MonolingualLIACC/Emakhuwa-FLORESLIACC/Emakhuwa-loanwords-detectionLIACC/Emakhuwa-News-Topic-ClassificationOCR Post Correction for Endangered Language TextsAdvancing Post-OCR Correction: A Comparative Study of Synthetic Data

LIACC/Emakhuwa-Monolingual

BibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-

LIACC/Emakhuwa-FLORES

FLORES+ dev and devtest set in Emakhuwa CC-BY-SA-4.0 @inproceedings{ali-etal-2024-expanding, title

LIACC/Emakhuwa-loanwords-detection

Paper: Detecting Loanwords in Emakhuwa: An Extremely Low-Resource {B}antu Language Exhibiting Signif

LIACC/Emakhuwa-News-Topic-Classification

BibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-

OCR Post Correction for Endangered Language Texts

There is little to no data available to build natural language processing models for most endangered languages. However, textual data in these languages often exists in formats that are not machine-readable, such as paper books and scanned images. In this work, we

Advancing Post-OCR Correction: A Comparative Study of Synthetic Data

This paper explores the application of synthetic data in the post-OCR domain on multiple fronts by c