# Yoruba Sentiment Corpus
A 5,000-sentence sentiment-annotated dataset for the Yoruba language,
created to address the shortage of labelled data for low-resource African NLP.
## Overview
- 5,000 sentences sourced from news, social media, and conversational text
- Labels: Positive, Negative, Neutral
- Annotation done by 3 collaborators using custom guidelines
- Inter-annotator agreement: Cohen's Kappa = 0.82
## Annotation Process
1. Sentences were collected and deduplicated
2. Guidelines drafted covering edge cases and cultural context
3. Each sentence independently labelled by 2 annotators
4. Disagreements resolved through consensus discussion
## Purpose
Built to support sentiment analysis research in Yoruba and contribute
to the broader low-resource NLP ecosystem.