Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data

Domain:

natural language processing

Record type:

paper
Creator:
GhoDemFra
Host:avatar
Considering the importance of detecting hateful language, labeled hate speech data is expensive and time-consuming to collect, particularly for low-resource languages. Prior work has demonstrated the effectiveness of cross-lingual transfer learning and data augmentation in improving performance on tasks with limited labeled data. To develop an efficient and scalable cross-lingual transfer learning approach, we leverage nearest-neighbor retrieval to augment minimal labeled data in the target language, thereby enhancing detection performance. Specifically, we assume access to a small set of labeled training instances in the target language and use these to retrieve the most relevant labeled examples from a large multilingual hate speech detection pool. We evaluate our approach on eight languages and demonstrate that it consistently outperforms models trained solely on the target language data. Furthermore, in most cases, our method surpasses the current state-of-the-art. Notably, our approach is highly data-efficient, retrieving as small as 200 instances in some cases while maintaining superior performance. Moreover, it is scalable, as the retrieval pool can be easily expanded, and the method can be readily adapted to new languages and tasks. We also apply maximum marginal relevance to mitigate redundancy and filter out highly similar retrieved instances, resulting in improvements in some languages.

Visit

arxiv.org

Tasks

hate speech detectiontext classificationtransfer learning

Tags

Computation and LanguageComputers and SocietyMultimedia

Similar

Data Efficient Dense Cross-Lingual Information RetrievalOptimal Transport Distillation Performance in Low-Resource Cross-Lingual Retrieval with Varying Labeled DataData-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesRobustness of Dual-Contrastive Cross-Lingual NER Models Under Limited Source-Language Labeled Data ConstraintsCross-lingual Transfer Performance Variation with Intermediate Task Training and Labeled Data RatiosLabel modification and bootstrapping for zero-shot cross-lingual hate speech detection

Data Efficient Dense Cross-Lingual Information Retrieval

Cross-Lingual Information Retrieval (CIR) remains challenging due to limited annotated data and ling

Optimal Transport Distillation Performance in Low-Resource Cross-Lingual Retrieval with Varying Labeled Data

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages

Hate speech is a global phenomenon, but most hate speech datasets so far focus on English-language c

Robustness of Dual-Contrastive Cross-Lingual NER Models Under Limited Source-Language Labeled Data Constraints

Cross-lingual Named Entity Recognition (NER) has recently become a research hotspot because it can a

Cross-lingual Transfer Performance Variation with Intermediate Task Training and Labeled Data Ratios

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

Label modification and bootstrapping for zero-shot cross-lingual hate speech detection

The goal of hate speech detection is to filter negative online content aiming at certain groups of p