Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

GemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages

Domain:

natural language processing

Record type:

papersoftware
Creator:
DanD'E
Host:avatar
As social-media platforms emerge and evolve faster than the regulations meant to oversee them, automated detoxification might serve as a timely tool for moderators to enforce safe discourse at scale. We here describe our submission to the PAN 2025 Multilingual Text Detoxification Challenge, which rewrites toxic single-sentence inputs into neutral paraphrases across 15 typologically diverse languages. Building on a 12B-parameter Gemma-3 multilingual transformer, we apply parameter-efficient LoRA SFT fine-tuning and prompting techniques like few-shot and Chain-of-Thought. Our multilingual training corpus combines 3,600 human-authored parallel pairs, 21,600 machine-translated synthetic pairs, and model-generated pairs filtered by Jaccard thresholds. At inference, inputs are enriched with three LaBSE-retrieved neighbors and explicit toxic-span annotations. Evaluated via Style Transfer Accuracy, LaBSE-based semantic preservation, and xCOMET fluency, our system ranks first on high-resource and low-resource languages. Ablations show +0.081 joint score increase from few-shot examples and +0.088 from basic CoT prompting. ANOVA analysis identifies language resource status as the strongest predictor of performance ($η^2$ = 0.667, p < 0.01).

Visit

arxiv.org

Tags

Computation and Language

Similar

Massively Multilingual Text Translation For Low-Resource LanguagesText Detoxification in isiXhosa and Yorùbá: A Cross-Lingual Machine Learning Approach for Low-Resource African LanguagesEnhancing Multilingual Table-to-Text Generation with QA Blueprints: Overcoming Challenges in Low-Resource LanguagesText Normalization for Low Resource LanguagesEnhancing Conversational AI for Low-Resource Languages: A Case Study on SomaliEnhancing Pos Tagging For Low-Resource Languages: A Case Study On Dholuo

Massively Multilingual Text Translation For Low-Resource Languages

Translation into severely low-resource languages has both the cultural goal of saving and reviving t

Text Detoxification in isiXhosa and Yorùbá: A Cross-Lingual Machine Learning Approach for Low-Resource African Languages

Toxic language is one of the major barrier to safe online participation, yet robust mitigation tools

Enhancing Multilingual Table-to-Text Generation with QA Blueprints: Overcoming Challenges in Low-Resource Languages

Limiting training data in low-resource languages is a barrier to Natural Language Processing (NLP).

Text Normalization for Low Resource Languages

This repository contains code related to the Google open source internship project Text Normalization for Low Resource Languages.

Enhancing Conversational AI for Low-Resource Languages: A Case Study on Somali

Conversational AI has made huge strides in understanding and generating human language. However, the

Enhancing Pos Tagging For Low-Resource Languages: A Case Study On Dholuo