Logo Lanfrica

Error Analysis in Amharic Hate Speech Detection Task

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssAsfAyeBel
Éditeur:
Und
Hôte:avatar
This paper examines the main challenges in Amharic hate speech classification tasks. The potential sources of errors in Amharic hate speech classification are explored in detail. Out of a total of more than 18 million tweets in our Twitter corpus, we annotate over 15k tweets with two independent Amharic native-speaker annotators. A Cohen's kappa score of 48% for inter-annotator agreement is achieved. An adjudicator is employed to decide on disputed annotations and determine the final gold labels. Logistic regression, Linear SVM, and the two fine-tuned contextual embedding transformer models: AmFLAIR and AmRoBERTa are employed. Among all the models, AmFLAIR achieves the best performance result with an F1-score of 72%.