Logo Lanfrica

iCompass-ai/Bambara-Sentiment-Analysis

Domaine:

natural language processing

Type de record:

dataset
Créateur:
iCo
Hôte:
# Context and Topics For easier communication, posting, or commenting on each others posts, people use their dialects. In Africa, various languages and dialects exist. One of the African languages is Bambara, used by citizens in different countries. Our dataset is the first Bamabara Dataset including more than 3K sentences, covering different topics, preprocessed and annotated as positive, negative, and neutral. # Collection Process our common-crawl-based dataset is composed of 1663 positive, 579 negative, and 804 neutral sentences. Data was collected by the iCompass team (icompass.tn). # Preprocessing and annotation BAMBARA was preprocessed by removing links, emoji symbols and punctuation. Annotation was then performed by TWO Malian native speakers, who are engineering students. Sentences are annotated as positive (1), negative(-1), or neutral (0). Find more in our paper: arxiv.org # Paper citation @article{BambaraSA, title={Bambara Language Dataset for Sentiment Analysis}, author={Mountaga Diallo and Chayma Fourati and Hatem Haddad}, year={2021}, eprint={2108.02524}, archivePrefix={arXiv}, primaryClass={cs.CL} } # Contact information * Website: icompass.tn * Twitter: @iCompass_ * Email: team@icompass.digital