Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Hassaniya Stories OCR Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
has
Host:
Image-to-text OCR pairs extracted from a Hassaniya Arabic stories corpus. Each row couples an image crop from the source book with its manually cleaned text. The primary columns are image and text. The original material was published in 1994 under the title 41 Short Stories About Life in Mauritania, Especially Life in Nouakchott. It was written by K.

Visit

huggingface.co

Tasks

computer visionoptical character recognition

Languages

Hassaniyya

Tags

hassaniyamauritaniaocrpost-ocr-correction

Similar

HASSANIYA DatasetDAH (Dataset Hassaniya)Emin009/hassaniya-punctuation-datasetMamadou-Aw/Hassaniya-speech-datasettaqacuct/kab-ocr-datasetmarconilabmak/luganda-ocr-dataset

HASSANIYA Dataset

The attached file is the first Mauritanian dialect dataset called “HASSANIYA” containing two thousan

DAH (Dataset Hassaniya)

DAH is a bilingual dataset created to support translation between Hassaniya dialect and English, wit

Emin009/hassaniya-punctuation-dataset

Mamadou-Aw/Hassaniya-speech-dataset

taqacuct/kab-ocr-dataset

Dataset synthétique pour entraîner un modèle OCR sur le kabyle. images/ : contiendra les images gén

marconilabmak/luganda-ocr-dataset

This dataset contains segmented line-level images and corresponding transcriptions in Luganda, a low