
This dataset contains a curated English–Bemba parallel corpus of 4,715 sentence pairs in the health-promotion domain. The corpus was collected using a custom mobile application that enabled teacher-in-the-loop crowdsourcing, offline/online synchronization, and real-time translation validation.
The dataset is designed to address the scarcity of digital resources for low-resource African languages, with specific focus on public health communication. Corpus analysis shows balanced sentence lengths, high lexical diversity, and domain alignment with medical texts.