Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation

Domain:

natural language processing

Record type:

paperdataset
Creator:
MutMug
Host:avatar
Cross-attention patterns in neural machine translation (NMT) are widely used to study how multilingual models align linguistic structure. We report a systematic artifact in cross-attention analysis of NLLB-200 (600M): non-content tokens - primarily end-of-sequence tokens, language tags, and punctuation - capture 83 percent to 91 percent of total cross-attention mass. We term these "attention sinks," extending findings from LLMs [Xiao et al., 2023] to NMT cross-attention and identifying a causal mechanism rooted in vocabulary design rather than position bias. This artifact causes raw metrics to underestimate content-level similarity by nearly half (36.7 percent raw vs. 70.7 percent filtered), rendering uncorrected analyses unreliable. To address this, we validate a content-only filtering methodology that removes non-content tokens and renormalizes the distribution. Applying this to 1,000 parallel sentences across African languages (Swahili, Kikuyu, Somali, Luo) and non-African benchmarks (German, Turkish, Chinese, Hindi), we confirm the artifact is universal and recover masked linguistic signals: a 16.9 percentage-point gap between teacher-forcing and generation modes, clear language-family clustering in attention entropy, and a hidden Somali paradox linking SOV word order to monotonic alignment. We release our filtering toolkit and corrected datasets to support reproducible interpretability research on multilingual NMT.

Visit

arxiv.org

Tasks

machine translation

Languages

GikuyuSomaliSwahili

Tags

Machine LearningComputation and Language

Similar

Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translationnileshbansal/Neural-Machine-Translation-with-AttentionMassively Multilingual Word EmbeddingsATTENTION BASED ENGLISH-AFAAN OROMO NEURAL MACHINE TRANSLATIONSyntax-Based Attention Masking for Neural Machine Translation BibleMMS: massively multilingual TTS dataset

Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation

Massively multilingual models for neural machine translation (NMT) are theoretically attractive, but often underperform bilingual models and deliver poor zero-shot translations. In this paper, we explore ways to improve them. We argue that multilingual NMT requires

nileshbansal/Neural-Machine-Translation-with-Attention

NMT with Attention to translate French into Fongbe and Ewe (two African languages - collectively cal

Massively Multilingual Word Embeddings

We introduce new methods for estimating and evaluating embeddings of words in more than fifty langua

ATTENTION BASED ENGLISH-AFAAN OROMO NEURAL MACHINE TRANSLATION

Neural Machine Translation (NMT) is a method for learning automatic translation using a single large

Syntax-Based Attention Masking for Neural Machine Translation

We present a simple method for extending transformers to source-side trees. We define a number of ma

BibleMMS: massively multilingual TTS dataset

The Dataset associated with the Paper "Meta Learning Text-to-Speech Synthesis in over 7000 Languages" by Florian Lux, Sarina Meyer, Lyonel Behringer, Frank Zalkow, Phat Do, Matt Coler, Emanuël A. P. Habets and Ngoc Thang Vu (Interspeech 2024). We generate 2000 spo