Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Swahili Corpus Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
ngu
Host:
A large-scale Swahili text corpus for language model pretraining and NLP research. The Swahili Corpus Dataset is a large-scale collection of Swahili (Kiswahili) text designed to support Natural Language Processing (NLP) research and the development of large language models (LLMs) for a low-resource African language.

Visit

huggingface.co

Tasks

language modeling

Languages

SwahiliSwahili, CoastalSwahili, Congo

Tags

swahilikiswahiliafricalow-resource-languagellm-pretrainingtext-corpus

Licenses

apache-2.0

Similar

Swahili Corpus DatasetSwahili CorpusSwahili CorpusSwahili CorpusSwahili News Corpusnavilindo/swahili-corpus

Swahili Corpus Dataset

Swahili Corpus

The repository contains several text files each corresponding to categories of Swahili textual conte

Swahili Corpus

This is a Swahili corpus obtained from CC-100: Monolingual Datasets from Web Crawl Data

Swahili Corpus

Swahili News Corpus

Language modeling, topic classification, AI training for Swahili NLP, digital literacy tools Notes

navilindo/swahili-corpus