Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
TouHarBas
Hôte:avatar
Ancient and endangered languages pose a unique challenge for NLP: their datasets are inherently scarce, difficult to expand, and built from formulaic corpora -- making data-quality issues especially consequential yet rarely audited. Motivated by the need to understand what current NMT can realistically achieve for such languages, we investigate hieroglyphic-to-German translation, where a recent study reported 61.5 BLEU using fine-tuned M2M-100. Our reproduction yields only 37.0 BLEU with the released model. Investigating this gap, we find 2\% of test targets appear identically in training (16/50; 50\% under 8-gram overlap at 70\% threshold). This contamination inflates scores dramatically: contaminated samples achieve up to 83.8 BLEU / 0.924 COMET-22 versus 30.9--39.2 BLEU / 0.622--0.676 COMET-22 on clean samples across five model configurations spanning two architectures. Document-level decontamination reduces contaminated BLEU by only 4.6 points because 8/16 targets persist via other source documents -- target-level deduplication is required. We release a decontaminated 34-sample test set and establish corrected baselines (30.9--39.2 BLEU), providing a realistic assessment of NMT capability for this endangered writing system. Accepted to NLP4DH 2026 Conference

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language

Similaires

Revisiting Low-Resource Neural Machine Translation: A Case StudyA Diverse Data Augmentation Strategy for Low-Resource Neural Machine TranslationData Augmentation for Low-Resource Neural Machine TranslationBacemDataScience/africa-climate-finance-translation-gap: Initial reproducibility releaseNeural Machine Translation for Amharic-English TranslationPre-Training on Mixed Data for Low-Resource Neural Machine Translation

Revisiting Low-Resource Neural Machine Translation: A Case Study

It has been shown that the performance of neural machine translation (NMT) drops starkly in low-reso

A Diverse Data Augmentation Strategy for Low-Resource Neural Machine Translation

One important issue that affects the performance of neural machine translation is the scale of avail

Data Augmentation for Low-Resource Neural Machine Translation

The quality of a Neural Machine Translation system depends substantially on the availability of siza

BacemDataScience/africa-climate-finance-translation-gap: Initial reproducibility release

Reproducibility package for: Digital Adoption, AI Readiness, and Climate Finance Mobilisati

Neural Machine Translation for Amharic-English Translation

This paper describes neural machine translation between orthographically and morphologically divergent languages. Amharic has a rich morphology; it uses the syllabic Ethiopic script. We used a new transliteration technique for Amharic to facilitate vocabulary shari

Pre-Training on Mixed Data for Low-Resource Neural Machine Translation

The pre-training fine-tuning mode has been shown to be effective for low resource neural machine tra