Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
KimJunRohKo,
Hôte:avatar
Low-resource language varieties used by specific groups remain neglected in the development of Multilingual Language Models. A great deal of cross-lingual research focuses on inter-lingual language transfer which strives to align allied varieties and minimize differences between them. However, for low-resource varieties, linguistic dissimilarity is also an important cue allowing generalization to unseen varieties. Unlike prior approaches, we propose a two-stage Language Generalization framework that focuses on capturing variety-specific cues while also exploiting rich overlap offered by high-resource source variety. First, we propose TOPPing, a source-selection method specifically designed for low-resource varieties. Second, we suggest a lightweight VACAI-Bowl architecture that learns variety-specific attributes with one branch while a parallel branch captures variety-invariant attributes using adversarial training. We evaluate our framework on structural prediction tasks, which are among the few tasks available, as proxy for performance on other downstream tasks. Using VACAI-Bowl with TOPPing yields an average 54.62% improvement in the dependency parsing task, which serves as a proxy for performance on other downstream tasks across 10 low-resource varieties. Accepted to CoNLL 2026

Visit

arxiv.org

Tasks

dependency parsingparsingtransfer learning

Tags

Computation and LanguageArtificial Intelligence

Similaires

LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language GeneralizationA Transformer-Based Language Model for Sentiment Classification and Cross-Linguistic Generalization: Empowering Low-Resource African LanguagesMulti-source Intermediate-task Training for Low-resource XTREME Language GeneralizationScaling Performance of Models Trained on Artificially Code-Switched Data for Unseen Low-Resource Language PairsMachine Translation into Low-resource Language VarietiesEfficient Test Time Adapter Ensembling for Low-resource Language Varieties

LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization

Pretrained language models (PLMs) have become remarkably adept at task and language generalization.

A Transformer-Based Language Model for Sentiment Classification and Cross-Linguistic Generalization: Empowering Low-Resource African Languages

Multi-source Intermediate-task Training for Low-resource XTREME Language Generalization

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Scaling Performance of Models Trained on Artificially Code-Switched Data for Unseen Low-Resource Language Pairs

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Machine Translation into Low-resource Language Varieties

State-of-the-art machine translation (MT) systems are typically trained to generate the "standard" t

Efficient Test Time Adapter Ensembling for Low-resource Language Varieties

Adapters are light-weight modules that allow parameter-efficient fine-tuning of pretrained models. S