Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Image-to-Image Translation Approach for Page Layout Analysis and Artificial Generation of Historical Manuscripts

Domaine:

natural language processing

Type de record:

paper
Créateur:
VidCam
Éditeur:
CenCalÉcoThi
Éditeur:
CCSDSpr
Hôte:avatar
Preprint version International audience Document layout analysis is essential in Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR), especially for historical and low-resource scripts. This study explores a novel data augmentation technique using Generative Adversarial Networks (GANs) to generate realistic document layouts from semantic masks, enhancing layout analysis without increasing human annotation effort.Our lightweight pipeline, tested on historical manuscripts (Latin, Arabic, Armenian, Hebrew), newspapers, and complex document layouts, shows that GAN-generated layouts are convincing and difficult to distinguish from real ones, even for paleographers. This method significantly boosts data augmentation, yielding a 3% point improvement in layout analysis metrics (precision, recall, mAP), and a 12 point increase in precision and recall for damaged documents. Additionally, masks with character information enhance image quality, boosting text recognition performance.

Visit

enc.hal.science

Tasks

computer visionoptical character recognition

Tags

Historical manuscriptsData augmentationHandwritten text recognitionGAN Generative Adversarial Network[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI][INFO.INFO-CV]Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV]

Licenses

http://creativecommons.org/licenses/by-nc-nd/info:eu-repo/semantics/OpenAccess

Similaires

Deposit page image for the deposit "Bogoŋ: a cultural, historical and linguistic documentation" depositpage_imageText Image Generation for Low-Resource Languages with Dual Translation LearningViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image GenerationDeposit page image for the deposit "Damakawa wordlist" depositpage_imageDeposit page image for the deposit "Description and Documentation of Avatime" depositpage_imageDeposit page image for the deposit "Documentation of Betta Kurumba" depositpage_image

Deposit page image for the deposit "Bogoŋ: a cultural, historical and linguistic documentation" depositpage_image

Text Image Generation for Low-Resource Languages with Dual Translation Learning

Scene text recognition in low-resource languages frequently faces challenges due to the limited avai

ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation

Recent studies have shown that Text-to-Image (T2I) model generations can reflect social stereotypes

Deposit page image for the deposit "Damakawa wordlist" depositpage_image

Deposit page image for the deposit "Description and Documentation of Avatime" depositpage_image

Deposit page image for the deposit "Documentation of Betta Kurumba" depositpage_image