Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Massively Multilingual Document Alignment with Cross-lingual Sentence-Mover's Distance

Domain:

natural language processing

Record type:

papersoftware
Creator:
El-Guz
Host:avatar
Document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. Such aligned data can be used for a variety of NLP tasks from training cross-lingual representations to mining parallel data for machine translation. In this paper we develop an unsupervised scoring function that leverages cross-lingual sentence embeddings to compute the semantic distance between documents in different languages. These semantic distances are then used to guide a document alignment algorithm to properly pair cross-lingual web documents across a variety of low, mid, and high-resource language pairs. Recognizing that our proposed scoring function and other state of the art methods are computationally intractable for long web documents, we utilize a more tractable greedy algorithm that performs comparably. We experimentally demonstrate that our distance metric performs better alignment than current baselines outperforming them by 7% on high-resource language pairs, 15% on mid-resource language pairs, and 22% on low-resource language pairs. In Proceedings of AACL-IJCNLP, 2020

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageInformation RetrievalMachine Learning

Similar

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and SpeechExploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence AlignmentEnhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word AlignmentDocHPLT: A Massively Multilingual Document-Level Translation DatasetXTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralizationCross-lingual Contrastive Loss Effects on Multilingual Embedding Alignment Quality

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech

Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstr

Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment

Multilingual sentence representations pose a great advantage for low-resource languages that do not

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu

DocHPLT: A Massively Multilingual Document-Level Translation Dataset

Existing document-level machine translation resources are only available for a handful of languages,

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

The Cross-lingual Natural Language Inference (XNLI) corpus is a crowd-sourced collection of 5,000 test and 2,500 dev pairs for the MultiNLI corpus. The pairs are annotated with textual entailment and translated into 14 languages: French, Spanish, German, Greek, Bu

Cross-lingual Contrastive Loss Effects on Multilingual Embedding Alignment Quality

Existing zero-shot cross-lingual transfer methods rely on parallel corpora or bilingual dictionaries