Amharic text dataset extracted from memes for hate speech detection or classification: Amharic Language Hate Speech Detection System from Facebook Image Post Using Deep Learning System
Domain:
natural language processing
Record type:
dataset
Creator:
meqmeq
Editor:
meq
Publisher:
Men
Host:
the dataset is collected from social media such as facebook and telegram.
the dataset is further processed. the collection are D1_org: this dataset is neither stemed nor stopword are remove: D1_sf: in this dataset stopwords are removed but not stemmed and in D3_stemed datset is stemmed and stopwords are removed. stemming is done using hornmorpho developed by Michael Gesser( available at HornMorpho)
all datasets are normalized and free from noise such as punctuation marks and emojs.