Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LIACC/Emakhuwa-Portuguese-OCR-post-correction

Domaine:

natural language processing

Type de record:

dataset
Créateur:
LIA
Hôte:
BibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung",

Visit

huggingface.co

Tasks

computer visionoptical character recognition

Languages

MakhuwaMakhuwa-Meetto

Licenses

cc-by-4.0

Similaires

LIACC/Emakhuwa-FLORESLIACC/Emakhuwa-MonolingualLIACC/Emakhuwa-loanwords-detectionLIACC/Emakhuwa-News-Topic-ClassificationOCR Post Correction for Endangered Language TextsAdvancing Post-OCR Correction: A Comparative Study of Synthetic Data

LIACC/Emakhuwa-FLORES

FLORES+ dev and devtest set in Emakhuwa CC-BY-SA-4.0 @inproceedings{ali-etal-2024-expanding, title

LIACC/Emakhuwa-Monolingual

BibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-

LIACC/Emakhuwa-loanwords-detection

Paper: Detecting Loanwords in Emakhuwa: An Extremely Low-Resource {B}antu Language Exhibiting Signif

LIACC/Emakhuwa-News-Topic-Classification

BibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-

OCR Post Correction for Endangered Language Texts

There is little to no data available to build natural language processing models for most endangered languages. However, textual data in these languages often exists in formats that are not machine-readable, such as paper books and scanned images. In this work, we

Advancing Post-OCR Correction: A Comparative Study of Synthetic Data

This paper explores the application of synthetic data in the post-OCR domain on multiple fronts by c