This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language.
Source: AfriBERTa Corpus (Swahili subset)
Language: Swahili
Size: [Insert total number of samples here]
Format: Hugging Face Dataset
The dataset consists of two main columns: