Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MENYO-20k: A Multi-domain English - Yorùbá Corpus for Machine Translation

Domain:

natural language processing

Record type:

dataset
Creator:
David Ifeoluwa AdelaniJesujoba O. AlabiDamilola AdebonojoAdesina Ayeni
Editor:
Adebayo O. AdeojoBabunde O. PopoolaOlumide AwokoyaModupe Olaniyi
Publisher:
Zenodo
Host:avatar

MENYO-20k is a multi-domain parallel dataset with texts obtained from news articles, ted talks, movie transcripts, radio transcripts, science and technology texts, and other short articles curated from the web and professional translators. The dataset has 20,100 parallel sentences split into 10,070 training sentences, 3,397 development sentences, and 6,633 test sentences (3,419 multi-domain, 1,714 news domain, and 1,500 ted talks speech transcript domain)

The dataset is open but for non-commercial use because some of the data sources like Ted talks and JW news requires permission for commercial use.

Acknowledgement: This project was supported by the AI4D language dataset fello… through K4All and Zindi Africa

Visit

doi.org

Tasks

machine translation

Languages

Yoruba

Tags

machine translation, yoruba, multi-domain

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcode

Similar

TWIENG: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of Twi, a Low-Resource African LanguageThe Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation تأثير المجال والتشكيل في الترجمة الآلية العصبية اليوروبية- الإنجليزية The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation The Effect of Domain and Diacritics in Yorùbá-English Neural Machine TranslationThe Effect of Domain and Diacritics in Yorùbá-English Neural Machine TranslationTwieng: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of the Twi Language, A Low-Resource African LanguageLinguistically-Motivated Yorùbá-English Machine TranslationA Corpus for Amharic-English Speech Translation: The Case of Tourism Domain

TWIENG: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of Twi, a Low-Resource African Language

A Twi-English parallel corpus is certainly an important resource for Machine Translation of Twi (ISO

The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation تأثير المجال والتشكيل في الترجمة الآلية العصبية اليوروبية- الإنجليزية The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation

Massively multilingual machine translation (MT) has shown impressive capabilities, including zero an

The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation

Massively multilingual machine translation (MT) has shown impressive capabilities, including zero and few-shot translation between low-resource language pairs. However, these models are often evaluated on high-resource languages with the assumption that they genera

Twieng: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of the Twi Language, A Low-Resource African Language

A Twi-English parallel corpus is certainly an important resource for Machine Translation of Twi (ISO

Linguistically-Motivated Yorùbá-English Machine Translation

Translating between languages where certain features are marked morphologically in one but absent or

A Corpus for Amharic-English Speech Translation: The Case of Tourism Domain

Speech translation research for the major languages like English, Japanese and Spanish has been conducted since the 1980’s. But no attempt were made in speech translation to/from the under-resourced language like Amharic. These activities suffered from the lack of