A curated dataset demonstrating linguistic reward hacking & safety failures in LLMs for the Hausa la
Pairwise evaluation dataset and annotation guidelines for Oromo language LLM fine-tuning and safety
Systematic red-teaming framework for multilingual LLM safety evaluation # African LLM Safety Evalua
This dataset contains sample outputs and evaluation scores from the study “A Test of Meaning, Form,