Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

OpenAI's GPT-OSS-20B Model and Safety Alignment Issues in a Low-Resource Language

Domaine:

natural language processing

Type de record:

paper
Créateur:
Inu
Hôte:avatar
In response to the recent safety probing for OpenAI's GPT-OSS-20b model, we present a summary of a set of vulnerabilities uncovered in the model, focusing on its performance and safety alignment in a low-resource language setting. The core motivation for our work is to question the model's reliability for users from underrepresented communities. Using Hausa, a major African language, we uncover biases, inaccuracies, and cultural insensitivities in the model's behaviour. With a minimal prompting, our red-teaming efforts reveal that the model can be induced to generate harmful, culturally insensitive, and factually inaccurate content in the language. As a form of reward hacking, we note how the model's safety protocols appear to relax when prompted with polite or grateful language, leading to outputs that could facilitate misinformation and amplify hate speech. For instance, the model operates on the false assumption that common insecticide locally known as Fiya-Fiya (Cyphermethrin) and rodenticide like Shinkafar Bera (a form of Aluminium Phosphide) are safe for human consumption. To contextualise the severity of this error and popularity of the substances, we conducted a survey (n=61) in which 98% of participants identified them as toxic. Additional failures include an inability to distinguish between raw and processed foods and the incorporation of demeaning cultural proverbs to build inaccurate arguments. We surmise that these issues manifest through a form of linguistic reward hacking, where the model prioritises fluent, plausible-sounding output in the target language over safety and truthfulness. We attribute the uncovered flaws primarily to insufficient safety tuning in low-resource linguistic contexts. By concentrating on a low-resource setting, our approach highlights a significant gap in current red-teaming effort and offer some recommendations. 6 pages, 4 tables

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Languages

Hausa

Tags

Computation and LanguageArtificial Intelligence

Similaires

Nadhari/gpt-oss-swahili-20bNadhari/gpt-oss-swahili-20b-mxfp4mradermacher/gpt-oss-swahili-20b-GGUFLLM Safety Alignment in Low-Resource Languages: A Systematic Literature ReviewA Comparison of Three AI Tutoring Bots Communicating in isiZulu Using OpenAI's GPT-3.5-turbo, GPT-4-turbo, and GPT-4oLo-Renz-O/medical-oss-20b-malagasy-LoRA

Nadhari/gpt-oss-swahili-20b

Nadhari/gpt-oss-swahili-20b-mxfp4

mradermacher/gpt-oss-swahili-20b-GGUF

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safet

A Comparison of Three AI Tutoring Bots Communicating in isiZulu Using OpenAI's GPT-3.5-turbo, GPT-4-turbo, and GPT-4o

Lo-Renz-O/medical-oss-20b-malagasy-LoRA