This dataset is a superset (N=48,076) of posts annotated as hateful or not. It results from the preprocessing and merge of all available hate speech datasets grounded geographically in Kenya in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in mid 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API