Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Mburisano Covid-19 multilingual corpus

Domain:

natural language processinghealthcare

Record type:

dataset
Creator:
Marais, Laurette
Editor:
Wilken, IlanaVan Niekerk, NinaCalteaux, Karen
Publisher:
CSIR Voice Computing
Host:avatar
This corpus was created to aid development of the AwezaMed Covid-19 speech-to-speech mobile application. The project within which it was created, Mburisano, was funded by the Department of Sport, Arts and Culture (DSAC). A selection of English sentences was generated in consultation with medical domain experts, and these sentences were manually translated into all official South African languages. The sentences formed the basis of the rapid development of Grammatical Framework (GF) application grammars for all the languages, to aid spoken communication about Covid-19 with a particular focus on screening and triage. The corpus is presented as a limited domain, manually translated parallel corpus in all 11 official South African languages. The AwezaMed Covid-19 application can be found [here](play.google.com).

Visit

hdl.handle.net

Tasks

machine translationspeech processingspeech translation

Tags

Covid-19

Licenses

Creative Commons Attribution 3.0 Unported (CC BY 3.0): https://www.creativecommons.org/licenses/by/3.0/

Similar

COVID-19 Multilingual TerminologyUNISA Multilingual CorpusMultilingual Spoken Words CorpusAfrican Multilingual Text CorpusNCHLT Auxiliary Speech Corpus - MultilingualLughaGen Multilingual African Language Corpus

COVID-19 Multilingual Terminology

COVID-19 multilingual terminology list document in all the South African languages. The development

UNISA Multilingual Corpus

The resource comprises a diverse selection of TEI P5 marked up documents from institutional origin,

Multilingual Spoken Words Corpus

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.

African Multilingual Text Corpus

NLP model training, translation Notes / challenges: Unequal language representation

NCHLT Auxiliary Speech Corpus - Multilingual

This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data S

LughaGen Multilingual African Language Corpus

LughaGen is a curated multilingual corpus for four Kenyan and East African languages: Swahili (sw),