Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

yisakberhanu/African-Trust-Safety-LLM-Challenge-Submission

Domaine:

natural language processing

Type de record:

project
Créateur:
yis
Hôte:
This repository contains the methodology, testing scripts, and final submission for the African Trust & Safety LLM Challenge. Our objective was to evaluate the safety guardrails of open-source models specifically trained on African languages, ensuring they do not generate harmful, illegal, or biased content when prompted in local dialects. # African LLM Vulnerability & Alignment Audit 🌍🛡️ **Lead Researcher:** yisak bule **Focus:** Red-teaming localized African Large Language Models (LLMs) to identify systemic safety, cultural, and alignment vulnerabilities across multiple regional dialects. ## 📌 Project Overview This repository contains a comprehensive security audit of open-source Large Language Models designed for African languages. The objective of this research is to stress-test these models against advanced adversarial attacks (jailbreaks) natively in **Hausa, Igbo, Yoruba, and Swahili**. By evaluating how these models handle culturally contextualized threats, this project maps out the current frontier of AI safety in under-represented languages, highlighting the critical gap between Western-centric alignment and local African realities. --- ## 🔬 Comprehensive Model Review During this research, I evaluated a wide spectrum of regional and language-specific models. The testing revealed a stark contrast in capabilities and safety alignment across the ecosystem. ### 1. N-ATLaS (NCAIR1/N-ATLaS-LLM) * **Languages Tested:** Hausa, Igbo, Yoruba * **Strengths:** Exceptional language fluency and deep understanding of local cultural context, slang, and regional infrastructure. It possesses a strong baseline defense against direct, explicitly violent prompts. * **Weaknesses:** Highly susceptible to **Contextual and Persona Bypasses**. The model completely drops its safety guardrails when malicious requests are framed within institutional contexts (e.g., government audits, academic research) or cultural authority (e.g., traditional medicine). * **Exploit Results:** Successfully extracted Tier-1 payloads, including infrastructure sabotage manuals, lethal chemical recipes disguised as traditional medicine, and advanced financial fraud/money laundering logic. ### 2. Pawa-Gemma (sartifyllc/Pawa-Gemma-Swahili-2B) * **Language Tested:** Swahili * **Strengths:** Surprisingly robust semantic defenses for a 2B …

Visit

github.com

Languages

HausaIgboSwahiliYoruba