Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tokenizer-Aware Cross-Lingual Adaptation of Decoder-Only LLMs through Embedding Relearning and Swapping

Domain:

natural language processing

Record type:

paper
Creator:
AssChuCohJiang, Fan
Publisher:
Und
Host:avatar
Extending Large Language Models (LLMs) to new languages is challenging, with most methods proposed suffering from high computational cost and catastrophic forgetting of original model capabilities. Embedding relearning~\citep{artetxe-etal-2020-cross}, a technique that creates new tokenizers and tunes embeddings on fixed model weights for target language adaptation, is both light-weight and performant. However, it has only been shown to work for older generation encoder-only models and for high resource languages. In this paper, we extend this framework to decoder-only LLMs focusing on joint adaptation to many languages, including low-resource ones. We experiment in three language groups over 100 languages each. We adapt a pre-trained LLM via switching to a customized tokenizer, and relearning the embedding layer. Across 96 diverse languages spanning both classification and generation tasks, we show embedding relearning improves \texttt{Gemma2} models by up to 20%, being highly competitive with full-weight updating baselines while vastly more computationally efficient and mitigating catastrophic forgetting. This translates into better results in transferring the improved multilingual performance to tasks that build on core English abilities (e.g., multilingual math reasoning), compared to various baselines. Further analysis reveals the critical role of customizing tokenizers in achieving effective language transfer, particularly for non-Latin script languages.

Visit

doi.orgunderline.io

Tasks

language modelingtransfer learning

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similar

Improving Cross-Lingual Transfer through Subtree-Aware Word ReorderingTrans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLPPidginUNMT: Cross-lingual embedding between Pidgin and EnglishUnderstanding Linearity of Cross-Lingual Word Embedding MappingsLLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual FeedbackDo language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens

Improving Cross-Lingual Transfer through Subtree-Aware Word Reordering

Despite the impressive growth of the abilities of multilingual language models, such as XLM-R and mT

Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP

The development of monolingual language models for low and mid-resource languages continues to be hi

PidginUNMT: Cross-lingual embedding between Pidgin and English

This repository contains the implementation of an Unsupervised NMT model from West African Pidgin (Creole) to English without using a single parallel sentence during training. Link to paper - https://arxiv.org/abs/1912.03444 (Accepted at NeurIPS 2019 Workshop on M

Understanding Linearity of Cross-Lingual Word Embedding Mappings

The technique of Cross-Lingual Word Embedding (CLWE) plays a fundamental role in tackling Natural La

LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback

To democratize large language models (LLMs) to most natural languages, it is imperative to make thes

Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens

Social media sentiment analysis has become one of the most significant instruments for understanding