A cleaned, metadata-rich Shona (sna) speech dataset prepared through a reproducible data engineering pipeline for downstream ASR and TTS workflows.
This release is intended as a general-purpose standard corpus: quality metadata is provided, but aggressive opinionated filtering is avoided so users can apply task-specific thresholds.
Source dataset: google/WaxalNLP