Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Lacuna Reconstruction: Self-supervised Pre-training for Low-Resource Historical Document Transcription

Domain:

natural language processing

Record type:

paper
Creator:
VogAllMilBer
Host:avatar
We present a self-supervised pre-training approach for learning rich visual language representations for both handwritten and printed historical document transcription. After supervised fine-tuning of our pre-trained encoder representations for low-resource document transcription on two languages, (1) a heterogeneous set of handwritten Islamicate manuscript images and (2) early modern English printed documents, we show a meaningful improvement in recognition accuracy over the same supervised model trained from scratch with as few as 30 line image transcriptions for training. Our masked language model-style pre-training strategy, where the model is trained to be able to identify the true masked visual representation from distractors sampled from within the same line, encourages learning robust contextualized language representations invariant to scribal writing style and printing noise present across documents.

Visit

arxiv.org

Tasks

computer visionoptical character recognition

Tags

Computer Vision and Pattern RecognitionComputation and LanguageMachine Learning

Similar

Open-Domain Response Generation in Low-Resource Settings using Self-Supervised Pre-Training of Warm-Started TransformersComparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language ModelsImproving Low-Resource Morphological Inflection via Self-Supervised ObjectivesSelf-Supervised Learning for Low-Resource Voice Recognition in Regional Television ChannelsTabuLM: Morphology-Aware Tabular Pre-training for Low-Resource LanguagesAfrica-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context

Open-Domain Response Generation in Low-Resource Settings using Self-Supervised Pre-Training of Warm-Started Transformers

Learning response generation models constitute the main component of building open-domain dialogue s

Comparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language Models

International audience This paper investigates the potential of improving a hybrid au

Improving Low-Resource Morphological Inflection via Self-Supervised Objectives

Self-supervised objectives have driven major advances in NLP by leveraging large-scale unlabeled dat

Self-Supervised Learning for Low-Resource Voice Recognition in Regional Television Channels

The development of an automatic speech recognition (ASR) system for regional television channels is

TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages

We present TabuLM, the first language model pre-trained on Kinyarwanda tabular data. Kinyarwanda is

Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context

We present the first self-supervised multilingual speech model trained exclusively on African speech