Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation

Domain:

natural language processing

Record type:

paper
A “bigger is better” explosion in the number of parameters in deep neural networks has made it increasingly challenging to make state-of-the-art networks accessible in compute-restricted environments. Compression techniques have taken on renewed importance as a way to bridge the gap. However, evaluation of the trade-offs incurred by popular compression techniques has been centered on high-resource datasets. In this work, we instead consider the impact of compression in a data-limited regime. We introduce the term low-resource double bind to refer to the co-occurrence of data limitations and compute resource constraints. This is a common setting for NLP for low-resource languages, yet the trade-offs in performance are poorly studied. Our work offers surprising insights into the relationship between capacity and generalization in data-limited regimes for the task of machine translation. Our experiments on magnitude pruning for translations from English into Yoruba, Hausa, Igbo and German show that in low-resource regimes, sparsity preserves performance on frequent sentences but has a disparate impact on infrequent ones. However, it improves robustness to out-of-distribution shifts, especially for datasets that are very distinct from the training distribution. Our findings suggest that sparsity can play a beneficial role at curbing memorization of low frequency attributes, and therefore offers a promising solution to the low-resource double bind.

Visit

aclanthology.orgarxiv.org

Tasks

machine translation

Languages

HausaIgboYoruba

Similar

Adapting to the Low-Resource Double-Bind: Investigating Low-Compute Methods on Low-Resource African LanguagesAn Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource LanguagesLow-Resource Machine Translation Training Curriculum Fit for Low-Resource LanguagesLesan -- Machine Translation for Low Resource LanguagesLesan: Machine Translation for Low Resource LanguagesLow-Resource Machine Translation for Moroccan Arabic

Adapting to the Low-Resource Double-Bind: Investigating Low-Compute Methods on Low-Resource African Languages

Many natural language processing (NLP) tasks make use of massively pre-trained language models, whic

An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages

In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks from a few examp

Low-Resource Machine Translation Training Curriculum Fit for Low-Resource Languages

We conduct an empirical study of neural machine translation (NMT) for truly low-resource languages,

Lesan -- Machine Translation for Low Resource Languages

Millions of people around the world can not access content on the Web because most of the content is not readily available in their language. Machine translation (MT) systems have the potential to change this for many languages. Current MT systems provide very accu

Lesan: Machine Translation for Low Resource Languages

Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.

Low-Resource Machine Translation for Moroccan Arabic