Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kiswahili-Tz-Hansard

Domain:

natural language processing

Record type:

dataset
This is a dataset of publically available Tanzania Hansard documents, in Kiswahili. It contains 2735 png images of pages from pdf documents, and text files containing transcripts obtained from the OCR tool tesseract-ocr. The images are obtained via scanning pdf files using imagemagick. Its intended use is in how improvements to language/word sequence modeling can improve OCR in a low-resource setting, and as a record of the accuracy of pre-existing OCR tools that use language models before any other methods are applied.

Visit

zenodo.org

Tasks

optical character recognitioncomputer vision

Languages

SwahiliSwahili, CoastalSwahili, Congo

Tags

ocrfair forwardai4dzindi

Similar

Open Source Kiswahili Spell Checker (SW-TZ)mysociety/za-hansardDicksonGT/fdi-tzmwemanoor/tz-locationsjovinvicent10/hydroclimate-tzopenhie/openinfoman-tz

Open Source Kiswahili Spell Checker (SW-TZ)

Different English software products are localized into many native languages spoken around the world

mysociety/za-hansard

A parser for South African Hansards, as published at http://www.parliament.gov.za/live/content.php?

DicksonGT/fdi-tz

site with infographics on FDI in Tanzania for the Data-journalism Bootcamp fdi-tz ====== site with

mwemanoor/tz-locations

Tanzania locations API and npm package— 31 regions, 168 districts, 4,054 wards, 75,061 streets with

jovinvicent10/hydroclimate-tz

End-to-end hydrometeorological data engineering pipeline for agricultural and climate decision suppo

openhie/openinfoman-tz

OpenInfoMan suport for Tanzania openinfoman-tz ==================== OpenInfoMan TZ Converting an