NaijaSenti
An MTEB dataset
Massive Text Embedding Benchmark
NaijaSenti is the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria — Hausa, Igbo, Nigerian-Pidgin, and Yorùbá — consisting of around 30,000 annotated tweets per language, including a significant fraction of code-mixed tweets.
Task category
t2c
Domains
Social, Written
Reference
NaijaSenti: A Nigerian Twit…