Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
AkeBaiDasHer
Hôte:avatar
Efficiently and accurately translating a corpus into a low-resource language remains a challenge, regardless of the strategies employed, whether manual, automated, or a combination of the two. Many Christian organizations are dedicated to the task of translating the Holy Bible into languages that lack a modern translation. Bible translation (BT) work is currently underway for over 3000 extremely low resource languages. We introduce the eBible corpus: a dataset containing 1009 translations of portions of the Bible with data in 833 different languages across 75 language families. In addition to a BT benchmarking dataset, we introduce model performance benchmarks built on the No Language Left Behind (NLLB) neural machine translation (NMT) models. Finally, we describe several problems specific to the domain of BT and consider how the established data and model benchmarks might be used for future translation efforts. For a BT task trained with NLLB, Austronesian and Trans-New Guinea language families achieve 35.1 and 31.6 BLEU scores respectively, which spurs future innovations for NMT for low-resource languages in Papua New Guinea.

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageArtificial Intelligence

Similaires

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation BenchmarksData Augmentation for Low Resource Neural Machine Translation for Sotho-Tswana LanguagesOpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource LanguagesEnabling Medical Translation for Low-Resource LanguagesLesan -- Machine Translation for Low Resource LanguagesLesan: Machine Translation for Low Resource Languages

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks

Data Augmentation for Low Resource Neural Machine Translation for Sotho-Tswana Languages

Neural Machine Translation (NMT) models have achieved remarkable performance on translating

OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages

In machine translation (MT), health is a high-stakes domain characterised by widespread deployment a

Enabling Medical Translation for Low-Resource Languages

We present research towards bridging the language gap between migrant workers in Qatar and medical s

Lesan -- Machine Translation for Low Resource Languages

Millions of people around the world can not access content on the Web because most of the content is not readily available in their language. Machine translation (MT) systems have the potential to change this for many languages. Current MT systems provide very accu

Lesan: Machine Translation for Low Resource Languages

Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.