Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification

Domaine:

natural language processing

Type de record:

paper
Créateur:
BelSánDalSte
Hôte:avatar
Multilingual toxicity detection remains a significant challenge due to the scarcity of training data and resources for many languages. While prior work has leveraged the translate-test paradigm to support cross-lingual transfer across a range of classification tasks, the utility of translation in supporting toxicity detection at scale remains unclear. In this work, we conduct a comprehensive comparison of translation-based and language-specific/multilingual classification pipelines. We find that translation-based pipelines consistently outperform out-of-distribution classifiers in 81.3% of cases (13 of 16 languages), with translation benefits strongly correlated with both the resource level of the target language and the quality of the machine translation (MT) system. Our analysis reveals that traditional classifiers outperform large language model (LLM) judges, with this advantage being particularly pronounced for low-resource languages, where translate-classify methods dominate translate-judge approaches in 6 out of 7 cases. We additionally show that MT-specific fine-tuning on LLMs yields lower refusal rates compared to standard instruction-tuned models, but it can negatively impact toxicity detection accuracy for low-resource languages. These findings offer actionable guidance for practitioners developing scalable multilingual content moderation systems.

Visit

arxiv.org

Tasks

hate speech detectionmachine translationtext classification

Tags

Computation and Language

Similaires

Simplify-Then-Translate: Automatic Preprocessing for Black-Box TranslationNollySenti: Leveraging Transfer Learning and Machine Translation for Nigerian Movie Sentiment ClassificationWolaytta-English Cross-lingual Information Retrieval using Neural Machine TranslationUniversal Cross-Lingual Text ClassificationEfficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled DataCross-Lingual Transfer Learning for Bambara Leveraging Resources From Other Languages

Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation

Black-box machine translation systems have proven incredibly useful for a variety of applications ye

NollySenti: Leveraging Transfer Learning and Machine Translation for Nigerian Movie Sentiment Classification

Africa has over 2000 indigenous languages but they are under-represented in NLP research due to lack of datasets. In recent years, there have been progress in developing labeled corpora for African languages. However, they are often available in a single domain

Wolaytta-English Cross-lingual Information Retrieval using Neural Machine Translation

Universal Cross-Lingual Text Classification

Text classification, an integral task in natural language processing, involves the automatic categor

Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data

Automatic speech recognition for low-resource languages remains fundamentally constrained by the sca

Cross-Lingual Transfer Learning for Bambara Leveraging Resources From Other Languages

Bambara, a language spoken primarily in West Africa, faces resource limitations that hinder the deve