Logo Lanfrica

VOLTRIX-01/yoruba-sentiment-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
VOL
Host:
# Yoruba Sentiment Corpus A 5,000-sentence sentiment-annotated dataset for the Yoruba language, created to address the shortage of labelled data for low-resource African NLP. ## Overview - 5,000 sentences sourced from news, social media, and conversational text - Labels: Positive, Negative, Neutral - Annotation done by 3 collaborators using custom guidelines - Inter-annotator agreement: Cohen's Kappa = 0.82 ## Annotation Process 1. Sentences were collected and deduplicated 2. Guidelines drafted covering edge cases and cultural context 3. Each sentence independently labelled by 2 annotators 4. Disagreements resolved through consensus discussion ## Purpose Built to support sentiment analysis research in Yoruba and contribute to the broader low-resource NLP ecosystem.

Languages