Logo Lanfrica

misingo255/african-languages-conversation-datasets

Domaine:

natural language processing

Type de record:

dataset
Créateur:
mis
Hôte:
African-languages-conversation-datasets # African Languages Conversation Datasets This is the collection of parallel dialogue datasets for 5 African languages for training language models for conversational bots. The datasets are for training and evaluating open-domain dialogue models. #### Languages * Yoruba * Swahili * Hausa * Nigerian Pidin * Kinyarwanda * Wolof #### NB: The Yoruba data was sourced from 2 blogs. Total samples per language: 1,500 Training set per language: 1,000 Validation set per language: 250 Test set per language: 250 The licence for using this dataset comes under CC-BY 4.0. #### Credits Github : Masakhane