Speech dataset for Turkana (tuv) — Eastern Nilotic language, ~1M speakers, Kenya.
Property
Value
Format
WAV, 16 kHz, mono / UTF-8 transcripts
Clips
5,151 segments
Splits
Train: 3,090 (60%) · Validation: 1,030 (20%) · Test: 1,031 (20%) — seed 42
Source
GRN Bible narratives (Global Recordings Network, LLL series 1–8), segmented via silence detection
Transcription
Auto-generated via facebook/mms-1b-all(Teso adapter)