Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
RötNozBiaHov
Hôte:avatar
Hate speech is a global phenomenon, but most hate speech datasets so far focus on English-language content. This hinders the development of more effective hate speech detection models in hundreds of languages spoken by billions across the world. More data is needed, but annotating hateful content is expensive, time-consuming and potentially harmful to annotators. To mitigate these issues, we explore data-efficient strategies for expanding hate speech detection into under-resourced languages. In a series of experiments with mono- and multilingual models across five non-English languages, we find that 1) a small amount of target-language fine-tuning data is needed to achieve strong performance, 2) the benefits of using more such data decrease exponentially, and 3) initial fine-tuning on readily-available English data can partially substitute target-language data and improve model generalisability. Based on these findings, we formulate actionable recommendations for hate speech detection in low-resource language settings. Accepted at EMNLP 2022 (Main Conference)

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Tags

Computation and Language

Similaires

Improving Speech Recognition for Under-resourced Languages Utilizing Audio-codecs for Data AugmentationAutomatic speech recognition for under-resourced languages: A surveySpeech recognition for under-resourced languages: Data sharing in hidden Markov model systemsData-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled DataDataScience-ArtificialIntelligence/Hate-speech-detection-for-Low-resource-languagesStrategies for building wordnets for under-resourced languages: The case of African languages

Improving Speech Recognition for Under-resourced Languages Utilizing Audio-codecs for Data Augmentation

Presenter: Nirayo Hailu Gebreegziabher, Ingo Siegert, Andreas Nürnberger, MMSP 2020, Virtual Event,

Automatic speech recognition for under-resourced languages: A survey

(Impact-F 1.28 estim. in 2012) International audience no abstract

Speech recognition for under-resourced languages: Data sharing in hidden Markov model systems

For purposes of automated speech recognition in under-resourced environments, t

Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data

Considering the importance of detecting hateful language, labeled hate speech data is expensive and

DataScience-ArtificialIntelligence/Hate-speech-detection-for-Low-resource-languages

# Hate-speech-detection-for-Low-resource-languages ## Overview This project aims to develop machine

Strategies for building wordnets for under-resourced languages: The case of African languages

The African Wordnet Project (AWN) aims at building wordnets for five African languages: Setswana, is