Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

marconilabmak/luganda-ocr-dataset

Domain:

natural language processing

Record type:

dataset
Creator:
mar
Host:
This dataset contains segmented line-level images and corresponding transcriptions in Luganda, a low-resource Bantu language spoken primarily in Uganda. It was created to support research in optical character recognition (OCR) and handwritten/printed text recognition for under-resourced African languages.

Visit

huggingface.co

Tasks

computer visionoptical character recognition

Languages

Ganda

Similar

Beijuka/luganda-ocr-datasetfidel-amharic-ocr-datasettaqacuct/kab-ocr-datasetHassaniya Stories OCR DatasetGebremaryam/Tigrigna-OCR-TTS-DatasetMoroccan Cultural Books OCR Dataset

Beijuka/luganda-ocr-dataset

fidel-amharic-ocr-dataset

Large-Scale Sentence Level Amharic OCR Dataset

taqacuct/kab-ocr-dataset

Dataset synthétique pour entraîner un modèle OCR sur le kabyle. images/ : contiendra les images gén

Hassaniya Stories OCR Dataset

Image-to-text OCR pairs extracted from a Hassaniya Arabic stories corpus. Each row couples an image

Gebremaryam/Tigrigna-OCR-TTS-Dataset

Moroccan Cultural Books OCR Dataset

This dataset contains 85,600+ page-level records extracted from PDF books covering the multifaceted