Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Optical Character Recognition and text cleaning in the indigenous South African languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
DanElsMic
Éditeur:
Stellenbosch University
Hôte:
This article represents follow-up work on unpublished presentations by the authors of text and corpus cleaning strategies for the African languages. In this article we provide a comparative description of cleaning of web-sourced and text-sourced material to be used for the compilation of corpora with specific attention to cleaning of text-based material, since this is particularly relevant for the indigenous South African languages. For the purposes of this study, we use the term “web-sourced material” to refer to digital data sourced from the internet, whereas “text-based material” refers to hard copy textual material. We identify the different types of errors found in such texts, looking specifically at typical scanning errors in these languages, followed by an evaluation of three commercially available Optical Character Recognition (OCR) tools. We argue that the cleanness of texts is a matter of granularity, depending on the envisaged application of the corpus comprised by the texts. Text corpora which are to be utilized for e.g. lexicographic purposes can tolerate a higher level of ‘noise’ than those used for the compilation of e.g. spelling and grammar checkers. We conclude with some suggestions for text cleaning for the indigenous languages of South Africa.

Visit

doi.org

Tasks

computer visionoptical character recognitiontext normalization

Similaires

Optical character recognition for South African languagesNCHLT Optical Character Recognition for South African LanguagesOptical character recognition of arabic printed textOptical character recognition of typeset Coptic text with neural networksOptical Character Recognition of Amharic Documentsisti-sys/Optical-Character-Recognition-OCR-

Optical character recognition for South African languages

NCHLT Optical Character Recognition for South African Languages

An OCR system is an application that enables one to convert scanned paper documents into editable an

Optical character recognition of arabic printed text

Optical character recognition of typeset Coptic text with neural networks

Abstract Digital Humanities (DH) within Coptic Studies, an emerging field of develo

Optical Character Recognition of Amharic Documents

In Africa around 2,500 languages are spoken. Some of these languages have their own indigenous scrip

isti-sys/Optical-Character-Recognition-OCR-

Optical Character Recognition (OCR) untuk membaca teks plat nomor kendaraan Indonesia dari gambar me