Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages

Domaine:

natural language processinghealthcare

Type de record:

paperdataset
Créateur:
MerSuoCohVyl
Hôte:avatar
In machine translation (MT), health is a high-stakes domain characterised by widespread deployment and domain-specific vocabulary. However, there is a lack of MT evaluation datasets for low-resource languages in this domain. To address this gap, we introduce OpenWHO, a document-level parallel corpus of 2,978 documents and 26,824 sentences from the World Health Organization's e-learning platform. Sourced from expert-authored, professionally translated materials shielded from web-crawling, OpenWHO spans a diverse range of over 20 languages, of which nine are low-resource. Leveraging this new resource, we evaluate modern large language models (LLMs) against traditional MT models. Our findings reveal that LLMs consistently outperform traditional MT models, with Gemini 2.5 Flash achieving a +4.79 ChrF point improvement over NLLB-54B on our low-resource test set. Further, we investigate how LLM context utilisation affects accuracy, finding that the benefits of document-level translation are most pronounced in specialised domains like health. We release the OpenWHO corpus to encourage further research into low-resource MT in the health domain. Accepted at WMT 2025

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageArtificial Intelligence

Similaires

EthioMT: Parallel Corpus for Low-resource Ethiopian LanguagesJW300: A Wide-Coverage Parallel Corpus for Low-Resource LanguagesIWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel CorpusFiltered Pseudo-parallel Corpus Improves Low-resource Neural Machine TranslationA Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation BenchmarksParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

EthioMT: Parallel Corpus for Low-resource Ethiopian Languages

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT leaves much to be desired for

JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages

IWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel Corpus

Repository for sharing the data in the Tamasheq language, one of the languages for the low-resource speech translation track at IWSLT 2022.

Filtered Pseudo-parallel Corpus Improves Low-resource Neural Machine Translation

Large-scale parallel corpora are essential for training high-quality machine translation systems; ho

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks

ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

Description of the Dataset This dataset consists of three parallel corpora involving the Kabyle lan