Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Datasets for Low-Resource Machine Translation of Arabic Dialects

Domain:

natural language processing

Record type:

datasetpaper
Creator:
IntAbi
Publisher:
Und
Host:avatar
Low-resource Machine Translation recently gained a lot of popularity, and for certain languages, it has made great strides. However, it is still difficult to track progress in other languages for which there is no publicly available evaluation data. In this paper, we introduce benchmark datasets for Arabic and its dialects. We describe our design process and motivations and analyze the datasets to understand their resulting properties. Numerous successful attempts use large monolingual corpora to augment low-resource pairs. We try to approach augmentation differently and investigate whether it is possible to improve MT models without any external sources of data. We accomplish this by bootstrapping existing parallel sentences and complement this with multilingual training to achieve strong baselines.

Visit

doi.orgunderline.io

Tasks

machine translation

Tags

Computer and Information ScienceInformation and Knowledge EngineeringIntelligent SystemNatural Language ProcessingNeural Network

Similar

Low-Resource Machine Translation for Moroccan ArabicBack-Translation and Unsupervised Domain Adaptation for Machine Translation of Arabic DialectsLesan -- Machine Translation for Low Resource LanguagesLesan: Machine Translation for Low Resource LanguagesLow-Resource Machine Translation Training Curriculum Fit for Low-Resource LanguagesData Augmentation for Low-Resource Neural Machine Translation

Low-Resource Machine Translation for Moroccan Arabic

Back-Translation and Unsupervised Domain Adaptation for Machine Translation of Arabic Dialects

Despite widespread use of Dialectal Arabic, research and resources for machine translation of the va

Lesan -- Machine Translation for Low Resource Languages

Millions of people around the world can not access content on the Web because most of the content is not readily available in their language. Machine translation (MT) systems have the potential to change this for many languages. Current MT systems provide very accu

Lesan: Machine Translation for Low Resource Languages

Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.

Low-Resource Machine Translation Training Curriculum Fit for Low-Resource Languages

We conduct an empirical study of neural machine translation (NMT) for truly low-resource languages,

Data Augmentation for Low-Resource Neural Machine Translation

The quality of a Neural Machine Translation system depends substantially on the availability of siza