Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation

Domaine:

natural language processing

Type de record:

paper
Créateur:
Zhade Poe
Éditeur:
arXiv
Hôte:avatar
Pixel-based language models (LMs) replace traditional tokenizers by processing rendered images of text, making cross-lingual transfer heavily dependent on the visual and structural properties of writing systems. However, the dynamics of adapting these models to low-resource languages with complex morphology and written in unique scripts are not yet explored. Using Tibetan as a case study, we analyze how continued pre-training of pixel-based LMs is influenced by data scale, initial script exposure, and cross-lingual transfer from languages written in other Brahmic scripts. We introduce four rendering-level metrics to quantify visual script similarity. We evaluate downstream performance across three tasks. Our results show that higher orthographic proximity enhances semantic transfer, even under severe data constraints. Additionally, we find a performance asymmetry based on the pre-training starting point: while multilingual pre-training PIXEL-M4 has stronger initial performance, its capacity for subsequent adaptation seems to be constrained, whereas adapting a monolingual model PIXEL with mixed scripts yields more gains on sentence-level tasks. Our metrics and case study offer empirical observations that could help inform data selection and script adaptation choices when working with pixel-based models in similar low-resource settings. EMNLP 2026: Main Conference

Visit

doi.org

Tasks

transfer learninglanguage modeling

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language ModelsPixel segmentation model for Algerian grapevine varietiesUnsupervised Language Model Adaptation for Low-Resource LanguagesSimiliscopy, leveraging large language model hallucinations to verify contextual similarity‘Seeing’ is ‘trying’: The relation of visual perception to attemptive modality in the world's languagesOn the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation

Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models

While pixel-based language modeling aims to bypass the sub-word tokenization bottleneck by rendering

Pixel segmentation model for Algerian grapevine varieties

0_main_data_preprocessing.py Summary This script standardizes and processes a raw dataset of plant

Unsupervised Language Model Adaptation for Low-Resource Languages

This paper introduces a two-way neural machine translation system from Bengali to English and vice v

Similiscopy, leveraging large language model hallucinations to verify contextual similarity

Abstract: This paper proposes "similiscopy" a novel binary (presence/absence) scientific technique t

‘Seeing’ is ‘trying’: The relation of visual perception to attemptive modality in the world's languages

Abstract This paper examines the relationship between the concepts of ‘seeing’ and ‘attempting/tryi

On the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation

Adapter-based tuning has recently arisen as an alternative to fine-tuning. It works by adding light-