Public benchmark and evaluation toolkit for clinical AI agents in African primary care settings. Tests evidence grounding, safe abstention, and escalation behavior across Shona, Ndebele, English, and expanding to Swahili/Amharic. Includes test scenarios, citation verification, runtime safety checks, and aligned with WHO/NICE protocols.