Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Old Nogay Turkish OCR Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
fat
Host:
This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis.

Visit

huggingface.co

Tasks

computer visionoptical character recognition

Languages

Burak

Tags

texthistoricalocrnogaiturkic-languageslow-resource-languagehistorical-nlpcorpus-linguisticsparquet

Licenses

cc-by-4.0

Similar

Old Catalan Morphosyntax: Developing an Annotated CorpusDevelopment of Lexicon of Turkish-French bilingual and Turkish and French monolingual children at primary schoolamharic-ocrdeepcopy/Ethiopic-OCRTifinagh OCR 39kTifinagh OCR 39k

Old Catalan Morphosyntax: Developing an Annotated Corpus

This paper presents a full procedure for the development of a Part-of-Speech (POS) tagged corpus of

Development of Lexicon of Turkish-French bilingual and Turkish and French monolingual children at primary school

International audience Lexical development of monolingual (Bassano, 2000, Kern, 2005)

amharic-ocr

deepcopy/Ethiopic-OCR

Tifinagh OCR 39k

This dataset is a comprehensive collection of 39,104 synthetic images designed for training and eval

Tifinagh OCR 39k

This dataset is a comprehensive collection of 39,101 synthetic images designed for training and eval