Logo Lanfrica

iCompass-ai/Bambara-Sentiment-Analysis

Domain:

natural language processing

Record type:

dataset
Creator:
iCo
Host:
# Context and Topics For easier communication, posting, or commenting on each others posts, people use their dialects. In Africa, various languages and dialects exist. One of the African languages is Bambara, used by citizens in different countries. Our dataset is the first Bamabara Dataset including more than 3K sentences, covering different topics, preprocessed and annotated as positive, negative, and neutral. # Collection Process our common-crawl-based dataset is composed of 1663 positive, 579 negative, and 804 neutral sentences. Data was collected by the iCompass team (icompass.tn). # Preprocessing and annotation BAMBARA was preprocessed by removing links, emoji symbols and punctuation. Annotation was then performed by TWO Malian native speakers, who are engineering students. Sentences are annotated as positive (1), negative(-1), or neutral (0). Find more in our paper: arxiv.org # Paper citation @article{BambaraSA, title={Bambara Language Dataset for Sentiment Analysis}, author={Mountaga Diallo and Chayma Fourati and Hatem Haddad}, year={2021}, eprint={2108.02524}, archivePrefix={arXiv}, primaryClass={cs.CL} } # Contact information * Website: icompass.tn * Twitter: @iCompass_ * Email: team@icompass.digital