Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Lost in Translation: Safety Alignment Failures in Nepali and Code-Switched Variants of Instruction-Tuned Large Language Models

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Pok
Publisher:
Zenodo
Host:avatar
Large Language Models (LLMs) are increasingly deployed in multilingual settings, yet safety alignment research remains predominantly English-centric. This work investigates the generalization of safety guardrails across linguistic variants of Nepali, a low-resource language spoken by over 30 million people. We introduce the Nepali Adversarial Safety Benchmark (NASB), a structured two-phase evaluation framework spanning five harm categories and five linguistic registers, comprising 95 systematic queries (NASB 1.0) and 290 expanded stress-test queries (NASB 2.0), with over 1,200 total adversarial probes conducted. Evaluating state-of-the-art models including Qwen-2.5-7B, Gemma-4 E2B, and Llama-3.1-8B, we identify a severe and consistent Safety Divergence: while models exhibit 0% bypass rate in English, rates rise to 73.7% in Devanagari and Formal Nepali registers. We note that NASB 1.0 rates are based on N=19 per register and should be interpreted as indicative estimates pending larger-scale replication. We document a novel attack vector, Intra-sentential Multi-script Polyglot Morphing (Vajra Morphing), which exploits sub-tokenization gaps by fusing Devanagari and Latin characters within single harmful keywords. We further identify three distinct failure modes: Semantic Drift, Persona Collapse, and Politeness Override. Cloud API testing of the Gemini 3 and Gemini 3.1 model families reveals additional safety asymmetries correlated with model scale and optimization level. Our findings demonstrate that current safety alignment is token-dependent rather than concept-aware, resulting in asymmetric risk for non-English-speaking populations. All findings were disclosed to Google AI VRP and Meta Whitehat prior to publication.

Visit

doi.orgzenodo.org

Tasks

code switching

Tags

LLM safety, multilingual jailbreak, Nepali NLP, low-resource languages, adversarial prompting, safety alignment, Devanagari, code-switching, Vajra Morphing, NASB

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian LanguagesLarge language models (LLMs)-generated Afrikaans-English code-switched dataEvaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric ReliabilityInvestigating Cultural Alignment of Large Language ModelsFine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched SpeechCross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian Languages

Large language models (LLMs) show remarkable human-like capability in various domains and languages.

Large language models (LLMs)-generated Afrikaans-English code-switched data

As highlighted in recent surveys, one of the biggest barriers to progress in code-switc

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

We investigate the translation quality of current large language models (LLMs) for English-to-Hausa

Investigating Cultural Alignment of Large Language Models

The intricate relationship between language and culture has long been a subject of exploration withi

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguis

Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

Humanitarian organizations face a critical choice: invest in costly commercial APIs or rely on free