Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Lwazi ASR Corpus: Low Resourced Languages

Domain:

natural language processing

Record type:

dataset
Creator:
dsf
Host:
This repository contains a curated collection of the Lwazi Automatic Speech Recognition (ASR) Corpus for several low-resourced South African languages. These datasets are designed for use in speech recognition research and development, particularly for underrepresented languages. Corpus Overview

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

NdebeleNdebeleSetswanaSotho, NorthernSotho, SouthernSwatiTsongaVendaXhosaZulu

Licenses

cc-by-3.0

Similar

Lwazi ASR Corpus: Low Resourced LanguagesLwazi Afrikaans ASR corpusLwazi English ASR corpusLwazi isiZulu ASR corpusLwazi isiNdebele ASR corpusLwazi Sepedi ASR corpus

Lwazi ASR Corpus: Low Resourced Languages

This repository contains a curated collection of the Lwazi Automatic Speech Recognition (ASR) Corpus

Lwazi Afrikaans ASR corpus

Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.

Lwazi English ASR corpus

Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.

Lwazi isiZulu ASR corpus

Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.

Lwazi isiNdebele ASR corpus

Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.

Lwazi Sepedi ASR corpus

Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.