Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
AkiLi,HamZam
Hôte:avatar
Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, critically underexplored. We introduce TUKABENCH, a jailbreak benchmark for seven African languages that extends JailbreakBench (JBB) beyond direct translation through four settings: human translation of JBB prompts, English adaptation to African contexts followed by human translation, human-curated prompts validated through interactions with GPT-5.2, and code-switched prompts combining English and African languages, isolating the effect of language, cultural grounding, and prompt evasiveness on model safety. Across closed and open models, prompting in African languages reduces refusal relative to English, with culturally adapted prompts leading to least refusal. The evaluation also surfaces two structural limitations: model comprehension failures and reduced LLM-as-a-judge reliability in LRLs. To capture the first, we introduce Deflection alongside Refused and Jailbroken; to assess the second, we validate outputs with human annotations, showing that judge-human agreement drops in lower-resource languages and less commonly supported scripts. Under review

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similaires

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African LanguagesThe BRIDGE Model: A Culturally Grounded Emotional Intelligence Framework for African and Collectivist ContextsAfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesKrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural AdvisoryCulturally-Grounded Chain-of-Thought (CG-CoT):Enhancing LLM Performance on Culturally-Specific Tasks in Low-Resource LanguagesEverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

Current guardian models are predominantly Western-centric and optimized for high-resource languages,

The BRIDGE Model: A Culturally Grounded Emotional Intelligence Framework for African and Collectivist Contexts

The dominance of Western, individualistic frameworks in emotional intelligence (EI) theory and pract

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

Africa is home to over 2000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages

KrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural Advisory

We present KrishokChat, the first citation-grounded Bengali agricultural instruction-tuning dataset

Culturally-Grounded Chain-of-Thought (CG-CoT):Enhancing LLM Performance on Culturally-Specific Tasks in Low-Resource Languages

Large Language Models (LLMs) struggle with culturally-specific reasoning tasks, particularly in low-

EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA),