Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Akan–English Maternal Health Parallel Text Corpus for Machine Translation

Domaine:

natural language processinghealthcare

Type de record:

dataset
Créateur:
Wia
Éditeur:
WiaEkpWinKwa
Éditeur:
Men
Hôte:avatar
This dataset contains a curated bilingual parallel corpus developed to support domain-specific neural machine translation (NMT) for maternal health communication between Akan and English. The corpus was constructed to address the scarcity of healthcare-specific parallel data for low-resource African languages, particularly Akan. The dataset comprises 20,101 cleaned English–Akan parallel sentence pairs, of which 12,100 pairs (60.2%) originate from maternal health content covering prenatal and postnatal care domains, and 8,006 pairs (39.8%) are drawn from general-domain sources to enhance linguistic diversity and model robustness. Maternal health topics represented include antenatal care, childbirth preparation, maternal mental health, nutrition, vaccination, medication use, lifestyle behaviours, preventive medicine, personal hygiene, and common pregnancy-related conditions. Value of the Dataset This dataset provides one of the first domain-specific maternal health resources for Akan that integrates both parallel text and aligned speech data, enabling research in neural machine translation, speech recognition, text-to-speech, and multimodal health communication systems. It supports the development of inclusive digital health tools such as maternal health chatbots and voice-based systems designed for Akan-speaking communities and other low-resource language contexts. This corpus was developed as part of the Ɔbaa Panin Project, which seeks to build a conversational maternal health chatbot in Akan. The Ɔbaa Panin Project is funded by Google Research.

Visit

data.mendeley.com

Tasks

machine translation

Languages

Akan

Tags

LinguisticsArtificial IntelligenceMachine TranslationText-to-SpeechLow-Resource LLM

Licenses

Creative Commons Attribution 4.0 Internationalhttp://creativecommons.org/licenses/by/4.0

Similaires

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel CorpusEnglish-Twi Parallel Corpus for Machine TranslationExtended Parallel Corpus for Amharic-English Machine TranslationAmharic-English Parallel Corpus for Neural Machine TranslationAmhEn: Amharic-English Large Parallel Corpus for Machine TranslationAmEn: Amharic-English Large Parallel Corpus for Machine Translation

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel Corpus

This dataset contains a curated bilingual parallel corpus developed to support domain-specific neura

English-Twi Parallel Corpus for Machine Translation

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected where necessary by native

Extended Parallel Corpus for Amharic-English Machine Translation

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for research purposes. Furthermore,

Amharic-English Parallel Corpus for Neural Machine Translation

Amharic is the working language of Ethiopia and, owing to its Semitic characteristics, the language

AmhEn: Amharic-English Large Parallel Corpus for Machine Translation

Recently, using deep neural networks for machine translation (MT) tasks has received great attention. In order for these networks to learn abstract representations of the input and store them as continuous vectors, they need a lot of data. However, very few researc

AmEn: Amharic-English Large Parallel Corpus for Machine Translation

Recently, using deep neural networks for machine translation (MT) tasks has received great attention