Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Machine Assisted Translation of Wikipedia Articles into Low Resource Languages

Domaine:

natural language processing

Type de record:

paper

Wikipedia is the largest encyclopedia ever assembled with the vision of enabling every human being to freely share in the sum of all knowledge. Wikipedia currently has a total of more than six million articles and over 17 billion words in its English edition. Unfortunately, millions of people cannot access this resource because it’s not available in their language. For instance, at the moment there are only 218 Tigrinya Wikipedia and 15,018 Amharic Wikipedia articles.

In this project, we investigate the problem of translating Wikipedia articles from a high resource language into low resource languages using human-in-the-loop MT systems. In particular, we investigate different approaches to translate a sample of English Wikipedia articles into Tigrinya and Amharic. Currently, this repository contains 100k English Wikipeida abstracts translated using Lesan (lesan.ai) into Amharic and Tigrinya.

Visit

zenodo.orgProposal for Wikimedia Research Grant

Tasks

machine translation

Languages

AmharicTigrigna

Similaires

Machine Translation into Low-resource Language VarietiesLesan -- Machine Translation for Low Resource LanguagesLesan: Machine Translation for Low Resource LanguagesLow-Resource Machine Translation Training Curriculum Fit for Low-Resource LanguagesIntegrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource LanguagesMultilingual Neural Machine Translation for Low Resource Languages

Machine Translation into Low-resource Language Varieties

State-of-the-art machine translation (MT) systems are typically trained to generate the "standard" t

Lesan -- Machine Translation for Low Resource Languages

Millions of people around the world can not access content on the Web because most of the content is not readily available in their language. Machine translation (MT) systems have the potential to change this for many languages. Current MT systems provide very accu

Lesan: Machine Translation for Low Resource Languages

Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.

Low-Resource Machine Translation Training Curriculum Fit for Low-Resource Languages

We conduct an empirical study of neural machine translation (NMT) for truly low-resource languages,

Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages

For most language combinations, parallel data is either scarce or simply unavailable. To address thi

Multilingual Neural Machine Translation for Low Resource Languages

Neural Machine Translation (NMT) has been shown to be more effective in translation tasks compared t