Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Ensembling of Distilled Models from Multi-task Teachers for Constrained Resource Language Pairs

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
HenGadAbdElM
Hôte:avatar
This paper describes our submission to the constrained track of WMT21 shared news translation task. We focus on the three relatively low resource language pairs Bengali to and from Hindi, English to and from Hausa, and Xhosa to and from Zulu. To overcome the limitation of relatively low parallel data we train a multilingual model using a multitask objective employing both parallel and monolingual data. In addition, we augment the data using back translation. We also train a bilingual model incorporating back translation and knowledge distillation then combine the two models using sequence-to-sequence mapping. We see around 70% relative gain in BLEU point for English to and from Hausa, and around 25% relative improvements for both Bengali to and from Hindi, and Xhosa to and from Zulu compared to bilingual baselines.

Visit

arxiv.org

Tasks

machine translation

Languages

HausaXhosaZulu

Tags

Computation and Language

Similaires

Efficient Test Time Adapter Ensembling for Low-resource Language VarietiesMulti-source Intermediate-task Training for Low-resource XTREME Language GeneralizationA Multi-Task Benchmark for Abusive Language Detection in Low-Resource SettingsA multi-task learning framework for sentiment analysis and news classification for low-resource languageDeep Conditional Census-Constrained Clustering (DeepC4) for Large-scale Multi-task Disaggregation of Urban MorphologyImpact of Intermediate-Task Training on Low-Resource Languages in Multilingual Language Models

Efficient Test Time Adapter Ensembling for Low-resource Language Varieties

Adapters are light-weight modules that allow parameter-efficient fine-tuning of pretrained models. S

Multi-source Intermediate-task Training for Low-resource XTREME Language Generalization

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

Content moderation research has recently made significant advances, but remains limited in serving t

A multi-task learning framework for sentiment analysis and news classification for low-resource language

Despite the growing progress in Natural Language Processing (NLP), low-resource languages such as Ha

Deep Conditional Census-Constrained Clustering (DeepC4) for Large-scale Multi-task Disaggregation of Urban Morphology

This Zenodo record contains the datasets of our research (Spatial Disaggregation of Rwandan Building

Impact of Intermediate-Task Training on Low-Resource Languages in Multilingual Language Models

Accuracy of English-language Question Answering (QA) systems has improved significantly in recent ye