Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Adeptschneider/CiviVox-Swahili-text-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Ade
Host:
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: [Insert total number of samples here] Format: Hugging Face Dataset The dataset consists of two main columns:

Visit

huggingface.co

Languages

Swahili

Tags

legal

Licenses

apache-2.0

Similar

Adeptschneider/CiviVox-Swahili-text-corpus-v2.0Adeptschneider/CiviVox-English-Swahili-text-translation-corpusniqqyniqqy/CiviVox-Swahili-text-corpus-v2.0Adeptschneider/llama3-civivox-english-to-swahili-translation

Adeptschneider/CiviVox-Swahili-text-corpus-v2.0

This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Co

Adeptschneider/CiviVox-English-Swahili-text-translation-corpus

niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0

This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Co

Adeptschneider/llama3-civivox-english-to-swahili-translation