Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

stem-content-ai-project/swahili-text-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
ste
Host:
This dataset contains a synthetic Swahili text corpus designed for training Text-to-Speech (TTS) models. The dataset includes a variety of Swahili phonemes to ensure phonetic diversity and high-quality TTS training. Format: JSONL (JSON Lines) Data Creation

Visit

huggingface.co

Tasks

speech processingtext to speech

Languages

Swahili

Similar

stem-content-ai-project/swahili-speechAdeptschneider/CiviVox-Swahili-text-corpusAdeptschneider/CiviVox-Swahili-text-corpus-v2.0niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0Adeptschneider/CiviVox-English-Swahili-text-translation-corpus

stem-content-ai-project/swahili-speech

This dataset contains paired audio and text data for training and evaluating speech-to-text models i

Adeptschneider/CiviVox-Swahili-text-corpus

This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Co

Adeptschneider/CiviVox-Swahili-text-corpus-v2.0

This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Co

niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0

This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Co

Adeptschneider/CiviVox-English-Swahili-text-translation-corpus