Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

kln001/kalenjin-asr-data

Domain:

natural language processing

Record type:

dataset
Creator:
kln
Host:
This dataset contains audio clips of Kalenjin speech and their corresponding transcriptions. It is intended for use in training and evaluating Automatic Speech Recognition (ASR) models for the Kalenjin language. This dataset was created from the Mozilla Common Voice project. It contains a total of [NUMBER] hours of audio, split into train, test, and validated sets. Languages

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

KalenjinKipsigis

Similar

rekody/kalenjin-asrRareElf/kalenjin-asrrekody/kalenjin-asr: v1.0.0 — paper releasebadrex/anv-data-ke-kalenjin-evalFrom a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and KalenjinpyFongbe ASR data

rekody/kalenjin-asr

Open Kalenjin ASR: evaluation harness, orthographic normalizer, and Parakeet-TDT fine-tuning pipelin

RareElf/kalenjin-asr

rekody/kalenjin-asr: v1.0.0 — paper release

First public release, accompanying the preprint "Open Kalenjin Automatic Speech Recognition: Adaptin

badrex/anv-data-ke-kalenjin-eval

From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin

Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistenc

pyFongbe ASR data

Python scripts and data for building Fongbe ASR by Fréjus Laleye