Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NCHLT Optical Character Recognition for South African Languages

Domain:

natural language processing

Record type:

software
Creator:
Martin PuttkammerJustin HockingRoald Eiselen
Publisher:
North-West UniversityCentre for Text Technology (CTexT)
Host:avatar
An OCR system is an application that enables one to convert scanned paper documents into editable and searchable texts. The engine analyses the structure of document image and divides the page into elements such as blocks of texts, tables and images. These blocks are used to identify character image patterns which are used to advance several hypotheses about the character possibilities. These hypotheses are used to produce different character, word and line level variations and associated probabilities. The set of probability hypotheses are then searched to find the most likely combination of characters, words and lines to produce a textual representation of the image.

Visit

hdl.handle.net

Tasks

computer visionoptical character recognition

Languages

AfrikaansNdebeleSetswanaSotho, NorthernSotho, SouthernSwatiTsongaVendaXhosaZulu

Licenses

Creative Commons Attribution 3.0 Unported License (CC BY 3.0): https://creativecommons.org/licenses/by/3.0/za/

Similar

Optical character recognition for South African languagesOptical Character Recognition and text cleaning in the indigenous South African languagesOptical character recognition for multilingual documents: Amazigh-FrenchGeneralized hough transform for arabic optical character recognitionOptical Character Recognition of Amharic Documentsisti-sys/Optical-Character-Recognition-OCR-

Optical character recognition for South African languages

Optical Character Recognition and text cleaning in the indigenous South African languages

This article represents follow-up work on unpublished presentations by the authors of text and corpu

Optical character recognition for multilingual documents: Amazigh-French

Generalized hough transform for arabic optical character recognition

Optical Character Recognition of Amharic Documents

In Africa around 2,500 languages are spoken. Some of these languages have their own indigenous scrip

isti-sys/Optical-Character-Recognition-OCR-

Optical Character Recognition (OCR) untuk membaca teks plat nomor kendaraan Indonesia dari gambar me