Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Amharic Social Media Dataset for Hate Speech Detection and Classification in Amharic Text with Deep Learning

Domain:

natural language processing

Record type:

dataset
Creator:
Sam
Editor:
Sam
Publisher:
Men
Host:avatar
This dataset is prepared for hate speech detection and classification into four categories of speech. Namely, Normal speech, Racial Hate speech, Religious Hate speech, Gender Hate speech and Disability Hate speech. This dataset is collected from three social media sites: Facebook, Twitter, and YouTube. The collection is done automatically and the data is annotated by human annotators. The dataset is collected only for Amharic Language. To make a clear annotation process we have developed and prepared an annotation guideline. We have made the annotation process a twofold round. The first round annotation is done by 100 annotators who have different demographic and sociocultural backgrounds. Before the annotation process is started, besides giving the developed guideline, a brief introduction is given to the annotators which includes: ● What hate speech is ● Social media and hate speech ● Impact of hate speech ● Types of hate speech ● How to control hate speech and also ● How to use the annotation website system to annotate the hate speech dataset To start the annotation, process the annotators have to sign up and login to our custom built annotators tool (annotate.shegerapps.com) called “Amharic Hate Speech Annotation Tool”. As the schema shows in Figure 5.2 the annotation tool which is the website has a database with ten tables in it, eight of the tables hold an annotated or labeled dataset. The rest two tables are to hold users (annotators, curators, and admin) for authentication purposes, and finally, the tenth table holds the raw data. Raw data table is a container where the to be annotated dataset is dumped then the annotators fetch the data from this raw data table, when data is annotated it is inserted into the respective eight tables. On this annotation part, we have annotated texts in eight categories but for this research, we need only the four categories. We included the other four hate speech categories for future studies so any interested researcher or ourselves can continue researching without the need for a new annotation. This annotation tool database is MySQL, the backend is developed using PHP and the frontend is done using HTML, JavaScript, and jQuery. Some of the advantages of this annotation tool are to create an efficient team-based annotation experience, it maintains control for data preparation, it is used to manage annotators’ tasks and their progress, and also it makes exporting the annotated dataset easier. After finalizing the annotation, the dataset is given to the respective model as input in CSV format. During training time this data is split into three with an 80:10:10 ratio for training, validation, and testing purposes

Visit

doi.orgdata.mendeley.com

Tasks

hate speech detectiontext classification

Languages

Amharic

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Hate Speech Detection and Classification in Amharic Text with Deep LearningHATE SPEECH DETECTION ON SOCIAL MEDIA FOR AMHARIC TEXT USING DEEP LEARNING APPROACHAmharic text dataset extracted from memes for hate speech detection or classification: Amharic Language Hate Speech Detection System from Facebook Image Post Using Deep Learning SystemAmharic text dataset extracted from memes for hate speech detection or classification Hate Speech Detection from Transliterated Amharic Social Media Comments Using Machine Learning and Deep Learning ApproachesEmotion Classification for Amharic Social Media Text Comments Using Deep Learning

Hate Speech Detection and Classification in Amharic Text with Deep Learning

Hate speech is a growing problem on social media. It can seriously impact society, especially in cou

HATE SPEECH DETECTION ON SOCIAL MEDIA FOR AMHARIC TEXT USING DEEP LEARNING APPROACH

HATE SPEECH DETECTION ON SOCIAL MEDIA FOR AMHARIC TEXT USING DEEP LEARNING APPROACH

Amharic text dataset extracted from memes for hate speech detection or classification: Amharic Language Hate Speech Detection System from Facebook Image Post Using Deep Learning System

the dataset is collected from social media such as facebook and telegram. the dataset is further pro

Amharic text dataset extracted from memes for hate speech detection or classification

the dataset is collected from social media such as facebook and telegram. the dataset is further pro

Hate Speech Detection from Transliterated Amharic Social Media Comments Using Machine Learning and Deep Learning Approaches

The rise of transliterated script usage on social media has presented significant challenges to hate

Emotion Classification for Amharic Social Media Text Comments Using Deep Learning