Logo Lanfrica

Detecting sentiments and combatting hate speech in Hausa, Igbo, Nigerian-Pidgin and Yorùbá - NaijaSenti: a Nigerian Corpus for Multilingual Sentiment Analysis

Domain:

natural language processing

Record type:

dataset
The NaijaSenti dataset is the first large-scale human-annotated Twitter sentiment dataset for Hausa, Igbo, Nigerian-Pidgin, and Yorùbá, the four most widely spoken languages in Nigeria. It consists of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), in… Notes / challenges: Fair Forward portfolio. Dataset > Model > Pilot Lacuna Fund data.

Similar