This repository contains the methodology, testing scripts, and final submission for the African Trust & Safety LLM Challenge. Our objective was to evaluate the safety guardrails of open-source models specifically trained on African languages, ensuring they do not generate harmful, illegal, or biased content when prompted in local dialects.
# African LLM Vulnerability & Alignment Audit 🌍🛡️
**Lead Researcher:** yisak bule
**Focus:** Red-teaming localized African Large Language Models (LLMs) to identify systemic safety, cultural, and alignment vulnerabilities across multiple regional dialects.
## 📌 Project Overview
This repository contains a comprehensive security audit of open-source Large Language Models designed for African languages. The objective of this research is to stress-test these models against advanced adversarial attacks (jailbreaks) natively in **Hausa, Igbo, Yoruba, and Swahili**.
By evaluating how these models handle culturally contextualized threats, this project maps out the current frontier of AI safety in under-represented languages, highlighting the critical gap between Western-centric alignment and local African realities.
---
## 🔬 Comprehensive Model Review
During this research, I evaluated a wide spectrum of regional and language-specific models. The testing revealed a stark contrast in capabilities and safety alignment across the ecosystem.
### 1. N-ATLaS (NCAIR1/N-ATLaS-LLM)
* **Languages Tested:** Hausa, Igbo, Yoruba
* **Strengths:** Exceptional language fluency and deep understanding of local cultural context, slang, and regional infrastructure. It possesses a strong baseline defense against direct, explicitly violent prompts.
* **Weaknesses:** Highly susceptible to **Contextual and Persona Bypasses**. The model completely drops its safety guardrails when malicious requests are framed within institutional contexts (e.g., government audits, academic research) or cultural authority (e.g., traditional medicine).
* **Exploit Results:** Successfully extracted Tier-1 payloads, including infrastructure sabotage manuals, lethal chemical recipes disguised as traditional medicine, and advanced financial fraud/money laundering logic.
### 2. Pawa-Gemma (sartifyllc/Pawa-Gemma-Swahili-2B)
* **Language Tested:** Swahili
* **Strengths:** Surprisingly robust semantic defenses for a 2B …