Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kirundi Open Speech & Text Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ijw
Host:
Building the first large-scale, open-source speech and text dataset for Kirundi 🚀 Get Started • 📊 Dataset • 🎯 Roadmap • 🫱🏿‍🫲🏾 Community Kirundi is spoken by over 12 million people, yet it remains a low-resource language largely ignored by modern AI systems. We're changing that.

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Rundi

Tags

kirundilow-resourceaudiospeech

Licenses

cc-by-4.0

Similar

Tamazight Open Speech DatasetTamazight Open Speech Dataset{language_name} Speech-Text DatasetKasem Speech-Text Parallel DatasetVai Speech-Text Parallel DatasetGa Speech-Text Parallel Dataset

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice

{language_name} Speech-Text Dataset

This dataset is a collection of aligned audio-text pairs in Lomwe, extracted from the CMU Wilderness

Kasem Speech-Text Parallel Dataset

This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Gha

Vai Speech-Text Parallel Dataset

This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana

Ga Speech-Text Parallel Dataset

This dataset is made available because of Ghana NLP's volunteer driven research work. Please conside