Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Shona Speech Dataset (SNA)

Domain:

natural language processing

Record type:

dataset
Creator:
man
Host:
A cleaned, metadata-rich Shona (sna) speech dataset prepared through a reproducible data engineering pipeline for downstream ASR and TTS workflows. This release is intended as a general-purpose standard corpus: quality metadata is provided, but aggressive opinionated filtering is avoided so users can apply task-specific thresholds. Source dataset: google/WaxalNLP

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Shona

Tags

audiospeechshonaasrttsafrican-languagelarge datasets from Lanfrica Insights

Licenses

cc-by-4.0

Similar

Shona Speech Dataset (SNA) - Annotatedshunyalabs/shona-speech-datasetBalanced Shona Hate Speech Dataset

Shona Speech Dataset (SNA) - Annotated

An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared throug

shunyalabs/shona-speech-dataset

Balanced Shona Hate Speech Dataset

This dataset contains 2,000 balanced examples of Shona text classified into four categories: NEUTRAL