Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

shama-llama/hate-speech-detection

Domaine:

natural language processing

Type de record:

model
Créateur:
sha
Hôte:
Adversarial and hierarchical transformer for Amharic hate speech detection # Hate Speech Detection for Amharic Language This project provides a pipeline for detecting hate speech in Amharic social media and online text. It combines multiple Amharic hate speech datasets, and then applies preprocessing and normalization. The detection system uses a state-of-the-art transformer model and applies adversarial and hierarchical architectures for classification. ## Architecture ## Data Sources The unified dataset is constructed from the following public Amharic hate speech datasets: - **SG2020** (Getachew, 2020): Dataset collected from Facebook pages of activists who write their posts using Geez script and comments of their followers. It is collected manually by going through each post and comment based on predefined rules. - **ZAK2021** (Zeleke, 2021): Extracted comments/posts pertaining to race, religion, and ethnicity using the Facepager API, resulting in a set of 30,000 comments between April 15, 2019 and December 15, 2019. A total of 5,000 comments/posts chosen at random for annotation. Three annotators (two candidate PhD. in Linguistics and one MSc. in Law) manually annotated the selected samples as “Hate” or “not-Hate” resulting 2,000 (1000 hate and 1000 non-hate) labeled comments. - **SM2022** (Minale, 2022): Dataset is prepared for hate speech detection and classification into four categories of speech. Namely, Normal speech, Racial Hate speech, Religious Hate speech, Gender Hate speech and Disability Hate speech. This dataset is collected from three social media sites: Facebook, Twitter, and YouTube. The collection is done automatically and the data is annotated by human annotators. The dataset is collected only for Amharic Language. - **MD2023** (Degu, 2023): Amharic text dataset extracted from memes in social media posts on Facebook and Telegram for hate speech detection / classification. - **RANLP2023** (Ayele et al., 2023): Collected using the Twitter API spanning from October 1, 2020 - November 30, 2022, considering t …

Visit

github.com

Tasks

hate speech detectiontext classification

Languages

AmharicShama-Sambuga

Tags

computer-sciencecosc-6252natural-language-processingxlm-roberta

Licenses

MIT

Similaires

shama-llama/miscellaneous-nlp-projectsLusanji/Hate-Speech-DetectionLuckilyeee/Hate-Speech-DetectionMekdim12/Amharic-Hate-speech-DetectionAllanOtieno254/Kiswahili-Hate-Speech-Detectionzakaria-lagouader/hate-speech-detection

shama-llama/miscellaneous-nlp-projects

Repository for NLP projects as part of CoSc 6252 # Miscellaneous NLP Projects This repository i

Lusanji/Hate-Speech-Detection

Hate speech detection in audio for English and Kiswahili languages # Automatic hate speech detectio

Luckilyeee/Hate-Speech-Detection

Achieving Hate Speech Detection in a Low Resource Setting # Achieving Hate Speech Detection in a Lo

Mekdim12/Amharic-Hate-speech-Detection

# Amharic Hate Speech Detection ### Using : ##### Tf-Idf ##### word2Vec ##### N-gras[ uni-gram , bi

AllanOtieno254/Kiswahili-Hate-Speech-Detection

A comprehensive project for detecting hate speech in Kiswahili text using machine learning technique

zakaria-lagouader/hate-speech-detection

a python app for detection hate speech in darija # Hate Speech Detection using NLP Created by: