First public release, accompanying the preprint "Open Kalenjin Automatic Speech Recognition: Adapting Parakeet-TDT to a Low-Resource Nilotic Language" (Kipkemboi, 2026).
kalebench/: scoring harness (CER/WER, paired bootstrap B=2000 seed=1234, significance), L0-L2 orthographic normalizer (normalize_kln.py), tier rescorer (rescore_tiers.py), 198-clip held-out evaluation manifest (gold text + filename pointers, speaker IDs pseudonymized), archived evaluation artifacts. CC-BY-SA-4.0.
asr-train/: Parakeet-TDT-0.6B-v3 fine-tuning pipeline for Modal. MIT.
No audio is redistributed; the source corpus (Anv-ke/Kalenjin, AfriVoices-KE) remains gated under its own access and consent terms.
Every number in the paper regenerates from the files in this release; see the paper's Appendix A and kalebench/paper/artifacts/asr_eval_report.json (frozen protocol, hashed evaluation set).