Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Challenges of Amharic Hate Speech Data Annotation Using Yandex Toloka Crowdsourcing Platform

Domain:

natural language processing

Record type:

dataset
Creator:
AssAsfAyeBel
Publisher:
Und
Host:avatar
This paper presents an Amharic hate speech annotation using the Yandex Toloka crowdsourcing platform. The dataset is collected from 5 consecutive years 2018-2022, following some controversial events in Ethiopia that put the country in violence. We consider one-month data starting from the event date in each year. Accordingly, we annotate 5,400 tweets, nearly 1k tweets for each year. We explore the main challenges of crowdsourcing annotation for Amharic hate speech data collection using Toloka. We attain a Fleiss kappa score of 0.34 using three independent annotators that annotate the tweets where the gold label is determined using majority voting. Using the datasets, we build classification models using deep learning (LSTM and BiLSTM) and achieved a 0.49 F1-score for both models. We will publicly release the dataset, source code, and models with a permissive license.

Visit

doi.orgunderline.io

Tasks

hate speech detectiontext classification

Languages

Amharic

Tags

Natural Language ProcessingLanguage Models