Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Separating Grains from the Chaff: Using Data Filtering to Improve Multilingual Translation for Low-Resourced African Languages

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
Abdulmumin, IdrisBeuAlaEme
Éditeur:
arXiv
Hôte:avatar
We participated in the WMT 2022 Large-Scale Machine Translation Evaluation for the African Languages Shared Task. This work describes our approach, which is based on filtering the given noisy data using a sentence-pair classifier that was built by fine-tuning a pre-trained language model. To train the classifier, we obtain positive samples (i.e. high-quality parallel sentences) from a gold-standard curated dataset and extract negative samples (i.e. low-quality parallel sentences) from automatically aligned parallel data by choosing sentences with low alignment scores. Our final machine translation model was then trained on filtered data, instead of the entire noisy dataset. We empirically validate our approach by evaluating on two common datasets and show that data filtering generally improves overall translation quality, in some cases even significantly. Accepted at the Seventh Conference on Machine Translation (WMT22)

Visit

doi.orgarxiv.org

Tasks

machine translation

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Machine Translation for Morphologically Rich Low-Resourced South African LanguagesLow Resourced Multilingual Neural Machine Translation for Ometo-EnglishBuilding Text-to-Speech Models for Low-Resourced Languages from Crowdsourced DataParticipatory Research for Low-resourced Machine Translation: A Case Study in African LanguagesGAIfE: Using GenAI to Improve Literacy in Low-resourced SettingsSmall Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages

Machine Translation for Morphologically Rich Low-Resourced South African Languages

Low Resourced Multilingual Neural Machine Translation for Ometo-English

In this paper, we present a new approach to overcome the problem of language resources that share si

Building Text-to-Speech Models for Low-Resourced Languages from Crowdsourced Data

Text-to-speech (TTS) models have expanded the scope of digital inclusivity by becoming a basis for a

Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages

Research in NLP lacks geographic diversity, and the question of how NLP can be scaled to low-resourced languages has not yet been adequately solved. {`}Low-resourced{'}-ness is a complex problem going beyond data availability and reflects systemic problems in socie

GAIfE: Using GenAI to Improve Literacy in Low-resourced Settings

Illiteracy is a predictor of many negative social and personal outcomes. Illiteracy rates are partic

Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages

Pretrained multilingual language models have been shown to work well on many languages for a variety of downstream NLP tasks. However, these models are known to require a lot of training data. This consequently leaves out a huge percentage of the world{'}s language