Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
Hu,HeeNakLee
Hôte:avatar
The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap. In particular, we introduce \textsf{SGToxicGuard}, a novel dataset and evaluation framework for benchmarking LLM safety in Singapore's diverse linguistic context, including Singlish, Chinese, Malay, and Tamil. SGToxicGuard adopts a red-teaming approach to systematically probe LLM vulnerabilities in three real-world scenarios: \textit{conversation}, \textit{question-answering}, and \textit{content composition}. We conduct extensive experiments with state-of-the-art multilingual LLMs, and the results uncover critical gaps in their safety guardrails. By offering actionable insights into cultural sensitivity and toxicity mitigation, we lay the foundation for safer and more inclusive AI systems in linguistically diverse environments.\footnote{Link to the dataset: github.com \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.} 9 pages, EMNLP 2025

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Tags

Computation and Language

Similaires

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature ReviewShadman19/llm-benchmark-low-resource-languagesaadilganigaie/LLM-for-Low-Resource-LanguagesLLM Probe: Evaluating LLMs for Low-Resource LanguagesGoing PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global SouthVLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safet

Shadman19/llm-benchmark-low-resource-languages

Evaluating open-source LLMs on Bengali, Swahili and Tamil vs English baseline # 🌍 LLM Benchmark for

aadilganigaie/LLM-for-Low-Resource-Languages

Extending the Vocabulary of Large Language Models for Low-Resource Languages In the realm of natural

LLM Probe: Evaluating LLMs for Low-Resource Languages

Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource a

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely cal

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evalu