Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

GLOCR: GeezLab OCR Dataset Tigrinya Text Recognition Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Gai
Editor:
GaiGai
Publisher:
Har
Host:avatar
A Text Recognition (TR) and Optical Character Recognition (OCR) dataset for the Tigrinya language. The dataset contains a total of 710k image-label pairs from multiple data sources. In addition to the characters-only data, the major part of the dataset is a collection of multi-word text images with labels from three categories: News (from Haddas Ertra newspaper), the Bible, and random-trigrams of the 150k most common words in Tigrinya.

Visit

doi.orgdataverse.harvard.edu

Tasks

computer visionoptical character recognition

Languages

Tigrigna

Tags

Computer and Information ScienceTigrinya, Geez, OCR, Optical Character Recognition, Text Recognition

Licenses

info:eu-repo/semantics/openAccessCreative Commons Zero v1.0 Universalhttps://creativecommons.org/publicdomain/zero/1.0/legalcode