Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Predicting Machine Translation Performance on Low-Resource Languages: The Role of Domain Similarity

Domaine:

natural language processing

Type de record:

paper
Créateur:
KhiTooAnuLiu
Hôte:avatar
Fine-tuning and testing a multilingual large language model is expensive and challenging for low-resource languages (LRLs). While previous studies have predicted the performance of natural language processing (NLP) tasks using machine learning methods, they primarily focus on high-resource languages, overlooking LRLs and shifts across domains. Focusing on LRLs, we investigate three factors: the size of the fine-tuning corpus, the domain similarity between fine-tuning and testing corpora, and the language similarity between source and target languages. We employ classical regression models to assess how these factors impact the model's performance. Our results indicate that domain similarity has the most critical impact on predicting the performance of Machine Translation models. 13 pages, 5 figures, accepted to EACL 2024, findings

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageMachine Learning

Similaires

Domain Similarity Impact on Zero-Shot Cross-Lingual Transfer Performance in Low-Resource LanguagesSelecting data for multilingual multi-domain neural machine translation on low resource languagesAfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African LanguagesCharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource LanguagesLesan -- Machine Translation for Low Resource LanguagesLesan: Machine Translation for Low Resource Languages

Domain Similarity Impact on Zero-Shot Cross-Lingual Transfer Performance in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Selecting data for multilingual multi-domain neural machine translation on low resource languages

[ACCESS RESTRICTED TO THE UNIVERSITY OF MISSOURI AT REQUEST OF AUTHOR.] While machine translation ha

AfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African Languages

CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages

We address the task of machine translation (MT) from extremely low-resource language (ELRL) to Engli

Lesan -- Machine Translation for Low Resource Languages

Millions of people around the world can not access content on the Web because most of the content is not readily available in their language. Machine translation (MT) systems have the potential to change this for many languages. Current MT systems provide very accu

Lesan: Machine Translation for Low Resource Languages

Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.