Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Label modification and bootstrapping for zero-shot cross-lingual hate speech detection

Domaine:

natural language processing

Type de record:

paper
Créateur:
BigHanGurFra
Éditeur:
Spr
Hôte:avatar
The goal of hate speech detection is to filter negative online content aiming at certain groups of people. Due to the easy accessibility and multilinguality of social media platforms, it is crucial to protect everyone which requires building hate speech detection systems for a wide range of languages. However, the available labeled hate speech datasets are limited, making it difficult to build systems for many languages. In this paper we focus on cross-lingual transfer learning to support hate speech detection in low-resource languages, while highlighting label issues across application scenarios, such as inconsistent label sets of corpora or differing hate speech definitions, which hinder the application of such methods. We leverage cross-lingual word embeddings to train our neural network systems on the source language and apply them to the target language, which lacks labeled examples, and show that good performance can be achieved. We then incorporate unlabeled target language data for further model improvements by bootstrapping labels using an ensemble of different model architectures. Furthermore, we investigate the issue of label imbalance in hate speech datasets, since the high ratio of non-hate examples compared to hate examples often leads to low model performance. We test simple data undersampling and oversampling techniques and show their effectiveness.

Visit

doi.orgtuprints.ulb.tu-darmstadt.de

Tasks

hate speech detectiontransfer learningtext classification

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Improving Zero-Shot Cross-Lingual Hate Speech Detection with Pseudo-Label Fine-Tuning of Transformer Language ModelsDomain Similarity Impact on Multilingual Hate Speech Detection Generalization in Zero-Shot Cross-Lingual TransferMultilingual Auxiliary Task Scaling for Zero-Shot Hate Speech Detection in Low-Resource LanguagesMultimodal Pretraining and Sequential Fine-Tuning for Zero-Shot Cross-Lingual Euphemism DetectionCross-Lingual IPA Contrastive Learning for Zero-Shot NERMultilingual LLM-based Teacher-Student Frameworks vs. Label Projection in Zero-Shot Cross-Lingual NER for Low-Resource Languages

Improving Zero-Shot Cross-Lingual Hate Speech Detection with Pseudo-Label Fine-Tuning of Transformer Language Models

Hate speech has proliferated on social media platforms in recent years. While this has been the focu

Domain Similarity Impact on Multilingual Hate Speech Detection Generalization in Zero-Shot Cross-Lingual Transfer

Automatic detection of abusive online content such as hate speech, offensive language, threats, etc.

Multilingual Auxiliary Task Scaling for Zero-Shot Hate Speech Detection in Low-Resource Languages

The goal of hate speech detection is to filter negative online content aiming at certain groups of p

Multimodal Pretraining and Sequential Fine-Tuning for Zero-Shot Cross-Lingual Euphemism Detection

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

Cross-Lingual IPA Contrastive Learning for Zero-Shot NER

Existing approaches to zero-shot Named Entity Recognition (NER) for low-resource languages have prim

Multilingual LLM-based Teacher-Student Frameworks vs. Label Projection in Zero-Shot Cross-Lingual NER for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident