Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Challenges of Amharic Hate Speech Data Annotation Using Yandex Toloka Crowdsourcing Platform

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AssAsfAyeBel
Éditeur:
Und
Hôte:avatar
This paper presents an Amharic hate speech annotation using the Yandex Toloka crowdsourcing platform. The dataset is collected from 5 consecutive years 2018-2022, following some controversial events in Ethiopia that put the country in violence. We consider one-month data starting from the event date in each year. Accordingly, we annotate 5,400 tweets, nearly 1k tweets for each year. We explore the main challenges of crowdsourcing annotation for Amharic hate speech data collection using Toloka. We attain a Fleiss kappa score of 0.34 using three independent annotators that annotate the tweets where the gold label is determined using majority voting. Using the datasets, we build classification models using deep learning (LSTM and BiLSTM) and achieved a 0.49 F1-score for both models. We will publicly release the dataset, source code, and models with a permissive license.

Visit

doi.orgunderline.io

Tasks

hate speech detectiontext classification

Languages

Amharic

Tags

Natural Language ProcessingLanguage Models

Similaires

The 5Js in Ethiopia: Amharic Hate Speech Data Annotation Using Toloka Crowdsourcing Platform

The 5Js in Ethiopia: Amharic Hate Speech Data Annotation Using Toloka Crowdsourcing Platform