Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BHEPRO-C: Bemba Health Promotional Corpus

Domaine:

natural language processinghealthcare

Type de record:

dataset
Créateur:
NeeKunNjo
Éditeur:
Zenodo
Hôte:avatar

This dataset contains a curated English–Bemba parallel corpus of 4,715 sentence pairs in the health-promotion domain. The corpus was collected using a custom mobile application that enabled teacher-in-the-loop crowdsourcing, offline/online synchronization, and real-time translation validation.

The dataset is designed to address the scarcity of digital resources for low-resource African languages, with specific focus on public health communication. Corpus analysis shows balanced sentence lengths, high lexical diversity, and domain alignment with medical texts.

Visit

doi.org

Tasks

machine translation

Languages

BembaNdasa

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

michsethowusu/bemba-sentiments-corpusBIG-C: Bemba Image Grounded Conversationsmichsethowusu/bemba-english-emotions-corpusBIG-C: a Multimodal Multi-Purpose Dataset for BembaBembaSpeech: A Speech Recognition Corpus for the Bemba LanguageBembaSpeech: A Speech Recognition Corpus for the Bemba Language

michsethowusu/bemba-sentiments-corpus

This dataset contains sentiment-labeled text data in Bemba for binary sentiment classification (Posi

BIG-C: Bemba Image Grounded Conversations

Training ASR (Automatic Speech Recognition) for Bemba - Building MT (Machine Translation) from Bemba

michsethowusu/bemba-english-emotions-corpus

This dataset contains emotion-labeled text data in Bemba-english for emotion classification (joy, sa

BIG-C: a Multimodal Multi-Purpose Dataset for Bemba

We present BIG-C (Bemba Image Grounded Conversations), a large multimodal dataset for Bemba. While Bemba is the most populous language of Zambia, it exhibits a dearth of resources which render the development of language technologies or language processing resea

BembaSpeech: A Speech Recognition Corpus for the Bemba Language

We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness fo

BembaSpeech: A Speech Recognition Corpus for the Bemba Language

We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, c