Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

stem-content-ai-project/swahili-speech

Domain:

natural language processing

Record type:

dataset
Creator:
ste
Host:
This dataset contains paired audio and text data for training and evaluating speech-to-text models in Swahili. The audio files have been processed to remove silence, converted to 44.1kHz mono FLAC format, and are paired with corresponding transcriptions. audio_*.flac: Audio files in FLAC format, named by their corresponding text corpus ID.

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Swahili

Licenses

mit

Similar

stem-content-ai-project/swahili-text-corpus

stem-content-ai-project/swahili-text-corpus

This dataset contains a synthetic Swahili text corpus designed for training Text-to-Speech (TTS) mod