Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Old Nogay Turkish OCR Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
fat
Hôte:
This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis.

Visit

huggingface.co

Tasks

computer visionoptical character recognition

Languages

Burak

Tags

texthistoricalocrnogaiturkic-languageslow-resource-languagehistorical-nlpcorpus-linguisticsparquet

Licenses

cc-by-4.0

Similaires

Old Catalan Morphosyntax: Developing an Annotated CorpusDevelopment of Lexicon of Turkish-French bilingual and Turkish and French monolingual children at primary schoolamharic-ocrdeepcopy/Ethiopic-OCRTifinagh OCR 39kTifinagh OCR 39k

Old Catalan Morphosyntax: Developing an Annotated Corpus

This paper presents a full procedure for the development of a Part-of-Speech (POS) tagged corpus of

Development of Lexicon of Turkish-French bilingual and Turkish and French monolingual children at primary school

International audience Lexical development of monolingual (Bassano, 2000, Kern, 2005)

amharic-ocr

deepcopy/Ethiopic-OCR

Tifinagh OCR 39k

This dataset is a comprehensive collection of 39,104 synthetic images designed for training and eval

Tifinagh OCR 39k

This dataset is a comprehensive collection of 39,101 synthetic images designed for training and eval