Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

Domain:

natural language processing

Record type:

datasetpaper
Creator:
IfeKonMeh
Publisher:
arXiv
Host:avatar
Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each attack type produces a distinct vulnerability profile: transliteration vulnerability is mediated by script identity, code-switching maintains effectiveness through the lowest-resource tier, and a sharp safety regime transition between Tiers 2 and 3 is consistent across all models. Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometrically misaligned subspace that projects insufficiently onto the refusal directions, leaving the refusal mechanism intact but untriggered. These findings show that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage. The benchmark and analysis code is at github.com.

Visit

doi.org

Tags

Cryptography and Security (cs.CR)Artificial Intelligence (cs.AI)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Ranks without resolution: data and code for a measurement audit of a multilingual LLM benchmarkAlign Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety AlignmentKambaata LLM Cultural Benchmarkrifaasa/hassaniya-llm-benchmarkFatika01/nigeria-livestock-llm-benchmarkBehailuBerhanu/kambaata-llm-cultural-benchmark

Ranks without resolution: data and code for a measurement audit of a multilingual LLM benchmark

Data, code and results for a measurement audit of HELM's African-language MMLU and Winogrande suite,

Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment

The widespread deployment of large language models (LLMs) across linguistic communities necessitates

Kambaata LLM Cultural Benchmark

Maintenance release for Zenodo archival of the Kambaata LLM Cultural Benchmark. This release contain

rifaasa/hassaniya-llm-benchmark

benchmark of modern LLMs on Hassaniya Arabic dialect » # hassaniya-llm-benchmark Code and evaluati

Fatika01/nigeria-livestock-llm-benchmark

A 420-question benchmark evaluating LLM performance on Nigerian livestock management knowledge, desi

BehailuBerhanu/kambaata-llm-cultural-benchmark

A 77-item benchmark for evaluating cultural knowledge, hallucination, and epistemic behavior in larg