Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Swahili Small Clean Corpus (Demo)

Domain:

natural language processing

Record type:

dataset
Creator:
mar
Host:
One-line summaryA small high-quality Swahili text corpus (≈58k documents, ~52.8 MB) extracted from Swahili Wikipedia and public news sources, cleaned, deduplicated, and split into train / validation / test. Suitable for fine-tuning and small-scale experiments. Languages: Swahili (sw) Total rows: ≈58,386 Size: ≈52.8 MB (parquet) Splits: train (≈46.7k rows), validation (≈5.84k rows), test (≈5.84k rows) Fields:

Visit

huggingface.co

Languages

Swahili

Tags

swahiliwikipediatextcleaneddeduplicated

Licenses

cc-by-sa-4.0

Similar

KamelTouati/whisper-small-darja-demongill-blip/asha-saathi-demo-english-swahiliKenyan Swahili ASR (clean, Nemotron-ready)Swahili CorpusSwahili CorpusSwahili Corpus

KamelTouati/whisper-small-darja-demo

--- language: - ar language_details: Algerian Arabic (Darja / الدارجة الجزائرية) license: mit tags:

ngill-blip/asha-saathi-demo-english-swahili

# Afya Rafiki — WhatsApp Refresher Demo (English / Kiswahili) A clickable, single-file mock-up of a

Kenyan Swahili ASR (clean, Nemotron-ready)

A cleaned, validated Kenyan-Swahili ASR corpus prepared for fine-tuning streaming ASR models (e.g. N

Swahili Corpus

The repository contains several text files each corresponding to categories of Swahili textual conte

Swahili Corpus

This is a Swahili corpus obtained from CC-100: Monolingual Datasets from Web Crawl Data

Swahili Corpus