Logo Lanfrica

niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0

Domaine:

natural language processing

Type de record:

dataset
Créateur:
niq
Hôte:
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: 1.54M Format: Hugging Face Dataset The dataset consists of two main columns: id: A unique identifier for each text entry