Logo Lanfrica

niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0

Domain:

natural language processing

Record type:

dataset
Creator:
niq
Host:
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: 1.54M Format: Hugging Face Dataset The dataset consists of two main columns: id: A unique identifier for each text entry