Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

Domain:

natural language processing

Record type:

paperdataset
Creator:
Dah
Host:avatar
Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally. We evaluate four open-weight instruction-tuned models on SomaliBench v0, a native-author-verified benchmark of 100 harmful-intent prompts paired across English and Somali. Each of Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B is run locally with temperature 0 and the same English "helpful, harmless, and honest" (HHH) system prompt. A pinned Claude Sonnet snapshot (claude-sonnet-4-5-20250929) classifies each response as refused, complied, or unclear; the native author spot-checks a stratified 80-row sample. We find large English-to-Somali refusal gaps for all four models: Llama-3.1-8B (0.90; 95% bootstrap CI [0.85, 0.96]), Aya-23-8B (0.75 [0.67, 0.83]), Qwen-2.5-7B (0.69 [0.59, 0.78]), and Gemma-2-9B (0.38 [0.27, 0.49]). For three models, the dominant Somali non-refusal mode is not fluent harmful compliance but unclear output: empty, wrong-language, or incoherent generations. The native verification spot-check achieves 100% agreement with the judge (Cohen's kappa = 1.00) on the 80 sampled rows. We report aggregate refusal rates, category gaps, and reliability statistics only; raw model generations are retained locally and are not released. 12 pages, 3 figures, 4 tables. Code: github.com Dataset: SomaliBench v0

Visit

arxiv.org

Tasks

text classification

Languages

Somali

Tags

Computation and LanguageArtificial IntelligenceComputers and SocietyI.2.7

Similar

Multilingual Biosecurity Safety Evaluation of Open-Weight Language Models: Evidence from African LanguagesRefusal Is Language-Dependent: A Cross-Lingual (English–Hindi) Refusal Benchmark and Evaluation of Llama-3.1-8BBridging language gaps in multilingual large language modelsMATERIAL Somali-English Language PackFuaadBashi/speech-to-text-AI-models-for-the-Somali-languageSomaliBench v0

Multilingual Biosecurity Safety Evaluation of Open-Weight Language Models: Evidence from African Languages

Refusal Is Language-Dependent: A Cross-Lingual (English–Hindi) Refusal Benchmark and Evaluation of Llama-3.1-8B

Safety alignment in large language models is trained and evaluated predominantly in English

Bridging language gaps in multilingual large language models

Large language models (LLMs) have revolutionized natural language processing, yet significant perfor

MATERIAL Somali-English Language Pack

MATERIAL Somali-English Language Pack was developed by Appen for the IARPA (Intelligence Advanced Re

FuaadBashi/speech-to-text-AI-models-for-the-Somali-language

Train or adapt a speech-to-text model capable of accurately transcribing Somali audio. # Somali Spe

SomaliBench v0

The first native-author-verified Somali safety evaluation benchmark. 100 harmful-intent prompts draw