Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Romanization-based Large-scale Adaptation of Multilingual Language Models

Domaine:

natural language processing

Type de record:

paper
Créateur:
PurRudPfeGur
Hôte:avatar
Large multilingual pretrained language models (mPLMs) have become the de facto state of the art for cross-lingual transfer in NLP. However, their large-scale deployment to many languages, besides pretraining data scarcity, is also hindered by the increase in vocabulary size and limitations in their parameter budget. In order to boost the capacity of mPLMs to deal with low-resource and unseen languages, we explore the potential of leveraging transliteration on a massive scale. In particular, we explore the UROMAN transliteration tool, which provides mappings from UTF-8 to Latin characters for all the writing systems, enabling inexpensive romanization for virtually any language. We first focus on establishing how UROMAN compares against other language-specific and manually curated transliterators for adapting multilingual PLMs. We then study and compare a plethora of data- and parameter-efficient strategies for adapting the mPLMs to romanized and non-romanized corpora of 14 diverse low-resource languages. Our results reveal that UROMAN-based transliteration can offer strong performance for many languages, with particular gains achieved in the most challenging setups: on languages with unseen scripts and with limited training data without any vocabulary augmentation. Further analyses reveal that an improved tokenizer based on romanized data can even outperform non-transliteration-based methods in the majority of languages. 9 pages, 5 figures

Visit

arxiv.org

Tasks

language modelingtransfer learning

Tags

Computation and LanguageMachine Learning

Similaires

EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language ModelsMassively Multilingual Adaptation of Large Language Models Using Bilingual Translation DataEvaluating the capability of base and large-scale language models for multilingual sarcasm detectionMaLA-500: Massive Language Adaptation of Large Language ModelsQuantifying Language Disparities in Multilingual Large Language ModelsBridging language gaps in multilingual large language models

EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models

In this work, we introduce EMMA-500, a large-scale multilingual language model continue-trained on t

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

This paper investigates a critical design decision in the practice of massively multilingual continu

Evaluating the capability of base and large-scale language models for multilingual sarcasm detection

Even though natural language understanding has made significant progress, language models still stru

MaLA-500: Massive Language Adaptation of Large Language Models

Large language models (LLMs) have advanced the state of the art in natural language processing. Howe

Quantifying Language Disparities in Multilingual Large Language Models

Results reported in large-scale multilingual evaluations are often fragmented and confounded by fact

Bridging language gaps in multilingual large language models

Large language models (LLMs) have revolutionized natural language processing, yet significant perfor