Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

CAMIO Transcription Languages

Domaine:

natural language processing

Type de record:

dataset
Éditeur:
ArrStrCar
Éditeur:
Linhtt
Hôte:avatar
*Introduction* CAMIO Transcription Languages was developed by the Linguistic Data Consortium and contains nearly 70,000 images of machine printed text with corresponding annotations and transcripts in the following 13 languages: Arabic, Chinese, English, Farsi, Hindi, Japanese, Kannada, Korean, Russian, Tamil, Thai, Urdu, and Vietnamese. This corpus is a subset of data created for a broader effort to support the development and evaluation of optical character recognition (OCR) and related technologies for 35 languages across 24 unique script types. The CAMIO (Corpus of Annotated Multilingual Images for OCR) collection was designed to address gaps in language and script coverage from existing corpora and to support future evaluation of OCR capabilities through a systematically constructed data set. *Data* Most images were annotated for text localization, resulting in over 2.3M line-level bounding boxes. For the 13 languages represented in this release, 1250 images per language were also annotated with orthographic transcriptions of each line plus specification of reading order, yielding over 2.4M tokens of transcribed text. The resulting annotations are represented in a comprehensive XML output format defined for this corpus. The script for each language is indicated in parentheses: Arabic (Arabic), Chinese (Simplified), English (Latin), Farsi (Arabic), Hindi (Devanagari), Japanese (Japanese), Kannada (Kannada), Korean (Hangul), Russian (Cyrillic), Tamil (Tamil), Thai (Thai), Urdu (Arabic), and Vietnamese (Latin). Data for each language is partitioned into test, train or validation sets. *Samples* Please view these samples: * Image Sample (png.ldcc) * Annotation Sample (xml) *Updates* None at this time.

Visit

catalog.ldc.upenn.edu

Tasks

computer visionoptical character recognition

Licenses

Licensing Instructions for Subscription & Standard Members, and Non-Members: http://www.ldc.upenn.edu/language-resources/data/obtainingLDC User Agreement for Non-Members: https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf

Similaires

samolubukun/Nigerian-Languages-WAZOBIA-Speech-TranscriptionFast transcription of speech in low-resource languagesCost Analysis of Human-corrected Transcription for Predominately Oral LanguagesEdge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili LanguagesUser-friendly automatic transcription of low-resource languages: Plugging ESPnet into ElpisAutomatic speech recognition for less-represented languages Transcription automatique de langues peu dotées

samolubukun/Nigerian-Languages-WAZOBIA-Speech-Transcription

This Gradio application provides a user-friendly interface for transcribing spoken audio in three ma

Fast transcription of speech in low-resource languages

We present software that, in only a few hours, transcribes forty hours of recorded speech in a surpr

Cost Analysis of Human-corrected Transcription for Predominately Oral Languages

Creating speech datasets for low-resource languages is a critical yet poorly understood challenge, p

Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages

This paper presents a novel framework for speech transcription and synthesis, leveraging edge-cloud

User-friendly automatic transcription of low-resource languages: Plugging ESPnet into Elpis

International audience This paper reports on progress integrating the speech recognit

Automatic speech recognition for less-represented languages Transcription automatique de langues peu dotées

With the development of technologies operating in a multilingual context, portability