Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Vocabulary Compression and Perplexity Degradation in Low-Resource Language Adaptation

Domain:

natural language processing

Record type:

paper
Creator:
Ass
Publisher:
Zenodo
Host:avatar
This report synthesises findings from 13 peer-reviewed papers addressing the following research question: What is the correlation between language-specific vocabulary compression ratios and perplexity degradation in low-resource language adaptation experiments. 10 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.2/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the correlation between language-specific vocabulary compression ratios and perplexity degradation in low-resource language adaptation experiments? Autonomous literature synthesis. Automated review score: 8.2/10. Full text and citation available at Assignee Research. Machine-generated literature synthesis. Content is derived from peer-reviewed papers; see individual sources for authoritative data. Automated review score: 8.2/10. Published by Assignee Research (assignee.net).

Visit

doi.orgzenodo.org

Tags

correlationlanguage-specificvocabularycompressionratiosperplexitydegradationlow-resource

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLPTheRootOf3/low-resource-language-model-adaptationUnsupervised Language Model Adaptation for Low-Resource LanguagesDeltaMerge-LowRes: Composing Language and Task Deltas for Low-Resource AdaptationEffective vocabulary expansion of multilingual language models for extremely low-resource languagesCross-lingual NER Transfer with Pretrained Language Models: Accuracy Degradation in Low-Resource Languages

Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP

The development of monolingual language models for low and mid-resource languages continues to be hi

TheRootOf3/low-resource-language-model-adaptation

Adapting pre-trained large language models to new languages in a low-resource regime 🌍 # Language M

Unsupervised Language Model Adaptation for Low-Resource Languages

This paper introduces a two-way neural machine translation system from Bengali to English and vice v

DeltaMerge-LowRes: Composing Language and Task Deltas for Low-Resource Adaptation

Adapting a multilingual encoder to a new language \emph{and} a new task with only a few hundred gold

Effective vocabulary expansion of multilingual language models for extremely low-resource languages

Multilingual pre-trained language models(mPLMs) offer significant benefits for many low-resource lan

Cross-lingual NER Transfer with Pretrained Language Models: Accuracy Degradation in Low-Resource Languages

Multilingual Language Models (MLLMs) exhibit robust cross-lingual transfer capabilities, or the abil