Logo Lanfrica

Hindi_Hate_Lexicon

Domain:

natural language processing

Record type:

dataset
Creator:
SawCatQur
Publisher:
Zenodo
Host:avatar
This repository presents a lexicon, the Hindi Hate Lexicon (HH-LEX), comprising lexical items associated with targets of hate speech. The lexicon predominantly consists of terms written in the Devanagari script, with a limited number of entries in Romanised Hindi. HH-LEX was constructed using a Hindi hate-speech dataset, namely the HASOC 2019 Hindi dataset, which contains 4,666 tweets. The lexicon was developed through a comprehensive manual analysis of individual tweets. Each tweet was systematically examined to identify lexical terms that explicitly or implicitly indicate the target of hate speech. These terms primarily include proper nouns as well as descriptive references to the targeted entities. The lexicon is organised into three categories to facilitate systematic analysis: Explicit Target Groups (n = 250), Implicit Targets (n = 250), and Explicitly Targeted Personalities (n = 346). Each category is further subdivided into relevant subcategories. In addition, a supplementary list of Hindi swear words (n=362) is compiled to support the study. This lexicon dataset is of particular relevance to the NLP research community, especially researchers working on hate speech detection, Indic languages, non-English text, multilingual and low-resource settings, annotation studies, and model explainability.