Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
SchChaSunGon
Hôte:avatar
We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages. We do not limit the the extraction process to alignments with English, but systematically consider all possible language pairs. In total, we are able to extract 135M parallel sentences for 1620 different language pairs, out of which only 34M are aligned with English. This corpus of parallel sentences is freely available at github.com. To get an indication on the quality of the extracted bitexts, we train neural MT baseline systems on the mined data only for 1886 languages pairs, and evaluate them on the TED corpus, achieving strong BLEU scores for many language pairs. The WikiMatrix bitexts seem to be particularly interesting to train MT systems between distant languages without the need to pivot through English. 13 pages, 3 figures, 6 tables

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language

Similaires

Extracting Parallel Sentences from Low-Resource Language Pairs with Minimal Supervisionmichsethowusu/African-Language-Parallel-Sentences-CollectionMining Large-Scale Low-Resource Pronunciation Data From Wikipediaghananlpcommunity/pristine-twi-english-parallel-sentencesNekonardo/lrl-parallel-sentence-miningParallel Corpora for bi-Directional Statistical Machine Translation for Seven Ethiopian Language Pairs

Extracting Parallel Sentences from Low-Resource Language Pairs with Minimal Supervision

Abstract At present, machine translation in the market depends on parallel sentence

michsethowusu/African-Language-Parallel-Sentences-Collection

This dataset collection includes sentence pairs for African languages along with similarity scores.

Mining Large-Scale Low-Resource Pronunciation Data From Wikipedia

Pronunciation modeling is a key task for building speech technology in new languages, and while soli

ghananlpcommunity/pristine-twi-english-parallel-sentences

A massive, sentence-level, deduplicated Twi ↔ English parallel dataset optimized specifically for tr

Nekonardo/lrl-parallel-sentence-mining

Codebase for benchmarking and enhancing multilingual sentence embeddings for parallel sentence minin

Parallel Corpora for bi-Directional Statistical Machine Translation for Seven Ethiopian Language Pairs

In this paper, we describe the development of parallel corpora for Ethiopian Languages: Amharic, Tigrigna, Afan-Oromo, Wolaytta and Geez. To check the usability of all the corpora we conducted baseline bi-directional statistical machine translation (SMT) experiment