Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Is linguistically-motivated data augmentation worth it?

Domain:

natural language processing

Record type:

paper
Creator:
GroGinPal
Host:avatar
Data augmentation, a widely-employed technique for addressing data scarcity, involves generating synthetic data examples which are then used to augment available training data. Researchers have seen surprising success from simple methods, such as random perturbations from natural examples, where models seem to benefit even from data with nonsense words, or data that doesn't conform to the rules of the language. A second line of research produces synthetic data that does in fact follow all linguistic constraints; these methods require some linguistic expertise and are generally more challenging to implement. No previous work has done a systematic, empirical comparison of both linguistically-naive and linguistically-motivated data augmentation strategies, leaving uncertainty about whether the additional time and effort of linguistically-motivated data augmentation work in fact yields better downstream performance. In this work, we conduct a careful and comprehensive comparison of augmentation strategies (both linguistically-naive and linguistically-motivated) for two low-resource languages with different morphological properties, Uspanteko and Arapaho. We evaluate the effectiveness of many different strategies and their combinations across two important sequence-to-sequence tasks for low-resource languages: machine translation and interlinear glossing. We find that linguistically-motivated strategies can have benefits over naive approaches, but only when the new examples they produce are not significantly unlike the training data distribution. Accepted to ACL 2025 Main. First two authors contributed equally

Visit

arxiv.org

Tags

Computation and Language

Similar

Linguistically-Motivated Yorùbá-English Machine TranslationIs a Generic Dataset and Foundation VLM for Arabic HTR Worth It? Lessons from AMIDDAWhen Unseen Domain Generalization is Unnecessary? Rethinking Data AugmentationGlide formation is not motivated by onset requirement in Malawian Tonga, glide epenthesis isData and Code for: Worth Your WeightWhat is the impact of synthetic data augmentation on low-resource machine translation quality

Linguistically-Motivated Yorùbá-English Machine Translation

Translating between languages where certain features are marked morphologically in one but absent or

Is a Generic Dataset and Foundation VLM for Arabic HTR Worth It? Lessons from AMIDDA

Which handwritten text recognition (HTR) strategy works best for Arabic scripts, for

When Unseen Domain Generalization is Unnecessary? Rethinking Data Augmentation

Recent advances in deep learning for medical image segmentation demonstrate expert-level accuracy. H

Glide formation is not motivated by onset requirement in Malawian Tonga, glide epenthesis is

Data and Code for: Worth Your Weight

I study the economic value of obesity---a status symbol in poor countries associated with raised hea

What is the impact of synthetic data augmentation on low-resource machine translation quality

One important issue that affects the performance of neural machine translation is the scale of avail