Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Improving the quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics

Domaine:

natural language processing

Type de record:

paperdatasetmodelsoftware
Créateur:
Ferde VelRat
Hôte:avatar
Parallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from web-mined corpora. Ranking sentence pairs using similarity scores on sentence embeddings derived from Pre-trained Multilingual Language Models (multiPLMs) is the most common PDC technique. However, previous research has shown that the choice of the multiPLM significantly impacts the quality of the filtered parallel corpus, and the Neural Machine Translation (NMT) models trained using such data show a disparity across multiPLMs. This paper shows that this disparity is due to different multiPLMs being biased towards certain types of sentence pairs, which are treated as noise from an NMT point of view. We show that such noisy parallel sentences can be removed to a certain extent by employing a series of heuristics. The NMT models, trained using the curated corpus, lead to producing better results while minimizing the disparity across multiPLMs. We publicly release the source code and the curated datasets. EMNLP 2025 Camera-ready version

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language

Similaires

Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel CorporaHeeLeeOss/low-resource-parallel-corporaGhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian LanguagesLearning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel CorporaImproving neural machine translation for low resource languages through non-parallel corpora: a case study of Egyptian dialect to modern standard Arabic translationPARME: Parallel Corpora for Low-Resourced Middle Eastern Languages

Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora

We conducted a detailed analysis on the quality of web-mined corpora for two low-resource languages

HeeLeeOss/low-resource-parallel-corpora

Openly-licensed, consented parallel corpora for under-served languages — with a clear schema, consen

GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages

Low resource languages present unique challenges for natural language processing due to the limited

Learning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel Corpora

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small

Improving neural machine translation for low resource languages through non-parallel corpora: a case study of Egyptian dialect to modern standard Arabic translation

Abstract Machine translation for low-resource languages poses significant challenges, primarily due

PARME: Parallel Corpora for Low-Resourced Middle Eastern Languages

The Middle East is characterized by remarkable linguistic diversity, with over 400 million inhabitan