Logo Lanfrica

shmuhammadd/NaijaSenti

Domain:

natural language processing

Record type:

dataset
Creator:
shm
Host:
This is a Lacuna Funded Project to develop sentiment and emotion corpus for three Nigerian languages: Igbo, Hausa, and Yoruba. Sannu da zuwa !!! Kaabo !!! Nnọọ!!! ⚠️ This README has been generated from the file(s) "blueprint.md" ⚠ ️--> -------------------------------------------------------------------------------- # Note: This repo is moved to: NaijaSenti: A Nigerian Twit…. NaijaSenti dataset can also be found on HugginFace: NaijaSenti `NaijaSenti` is an open-source sentiment and emotion corpora for four major Nigerian languages. This project was supported by lacuna-fund initiatives. Jump straight to one of the sections below, or just scroll down to find out more. ## Update (05/09/2022): We are running a SemEval competition and we release more sentiment dataset from African languages including NaiJaSenti Dataset. Visit the AfriSenti SemEval page for more information : AfriSenti-SemEval Task 12 ## Update (05/09/2022): Send me email (shamsuddeen2004@gmail.com) if you need NaijaSenti Dataset. We can send you anonymized dataset. anonomize Table of Contents ## Table of Contents - Paper - Abstract - Language Resource Developed - papers from this project - Contact us ## Paper and Datasheet for Dataset - Read the `NaijaSenti` paper: NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis - Read the `NaijaSenti Datasheet` coming soon... ## Abstract Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria—Hausa, Igbo, Nigerian-Pidgin, and Yorùbá—consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets. We propose text collection, filtering, processing, and labelling methods that enable us to create datasets for these low-resource languages. We evaluate a range of pre-trained mo …