Logo Lanfrica

Centre-for-Information-Resilience/ethiopia-hate-speech-lexicon

Domain:

natural language processing

Record type:

dataset
Creator:
Cen
Host:
CIR: An analysis of gendered hate speech on social media in Ethiopia # Inflammatory Keywords and Phrases **UPDATED LEXICON** Researching digital violence targeting women and girls — and understanding the intersectional hate they face — often starts with lexicons. Why? Because lexicons help researchers: * Detect harmful content at scale, even in massive data sets * Spot patterns and trends in gendered, ethnic, and religious hate * Compare narratives across languages and communities * Build evidence for advocacy, safety interventions, and policy change \ In the spirit of collaboration and open knowledge, we’re releasing an expanded lexicon of inflammatory terms in **Amharic, Afaan Oromo, Tigrigna, and English** — `now 47% larger than before`. Sharing this openly means researchers, civil society, journalists, and technologists can all investigate digital violence more rigorously and transparently. This lexicon is the most comprehensive of its kind for analysing the Ethiopian social media landscape. It covers inflammatory keywords across four languages **(Amharic, Afaan Oromo, Tigrigna, and English)** which may be indicative of hate speech along gendered, ethnic, and religious lines. At CIR, we build and use lexicons to help us identify harmful content across digital spaces. The lexicon was developed for CIR's project on Tech-Facilitated Gender-Based Violence in Ethiopia. For more information on the Lexicon development, see CIR's reports and academic publications: \ **Normalised and Invisible: An analysis of gendered hate speech on social media in Ethiopia** **No Safe Scroll: Investigating gendered hate speech on TikTok and YouTube in Ethiopia** **Resources for Annotating Hate Speech in Social Media Platforms Used in Ethiopia: A Novel Lexicon and Labelling Scheme** \ It is important to note that terms on their own, may not constitute hate speech. The keywords were used to obtain content from social media which could contain hate speech; however, human annotators then analysed whether the content was/wasn't hate speech, as pe …