Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Mburisano Covid-19 multilingual corpus

Domaine:

natural language processinghealthcare

Type de record:

dataset
Créateur:
Marais, Laurette
Éditeur:
Wilken, IlanaVan Niekerk, NinaCalteaux, Karen
Éditeur:
CSIR Voice Computing
Hôte:avatar
This corpus was created to aid development of the AwezaMed Covid-19 speech-to-speech mobile application. The project within which it was created, Mburisano, was funded by the Department of Sport, Arts and Culture (DSAC). A selection of English sentences was generated in consultation with medical domain experts, and these sentences were manually translated into all official South African languages. The sentences formed the basis of the rapid development of Grammatical Framework (GF) application grammars for all the languages, to aid spoken communication about Covid-19 with a particular focus on screening and triage. The corpus is presented as a limited domain, manually translated parallel corpus in all 11 official South African languages. The AwezaMed Covid-19 application can be found [here](play.google.com).

Visit

hdl.handle.net

Tasks

machine translationspeech processingspeech translation

Tags

Covid-19

Licenses

Creative Commons Attribution 3.0 Unported (CC BY 3.0): https://www.creativecommons.org/licenses/by/3.0/