Background.
Multi-turn patient-persona evaluation is increasingly used to probe clinical-decision safety of large language models (LLMs) in low- and middle-income country (LMIC) primary care. The evaluation harnesses themselves, and the LLM-drafted test fixtures they consume, have received little methodological scrutiny.
Methods.
We evaluated the Bizusizo WhatsApp-native clinical triage system (aligned to the South African Triage Scale) against 14 multi-turn patient personas spanning three South African languages (English, isiZulu, Sesotho) under six LLM tiers across three vendors (Anthropic, Google, OpenAI). N=5 per cell; fourteen pre-specified ablation probes (N=5-30) against Sonnet 4; N=10-per-variant cross-vendor replication.
Findings.
Five classes of silent harness behaviour were identified at runtime. (1) Silent-constant fallback: 195/195 Opus calls returned constant YELLOW via API-compatibility fallback, detected by cross-language control personas. (2) Systematic parse-salvage fallback: 73.3% of Haiku turns emitted malformed JSON salvaged by token string-matching, detected by reasoning-field audit. (3) Structured-output internal inconsistency: Sonnet 4's reasoning and triage_level fields contradicted each other on 72/72 valid-JSON runs of an HIV+fever+meningism scenario; bidirectional clause-level ablation identified a concessive hedge clause as the trigger. Cross-vendor replication established Sonnet-4-family specificity (Gemini 2.5 Pro 59/60 coherent; GPT-4o 60/60 coherent). (4) Parse-salvage accidentally correct: harness's parse-salvage path produced correct ORANGE on probes E, F, masking the underlying inconsistency. (5) Coherent instructed-rule non-compliance: Gemini 2.5 Flash-Lite articulated and rejected the system prompt's Step-4 rule, emitting coherent-but-non-compliant YELLOW. A parallel native-speaker audit across eight South African languages identified a complementary sixth class - fixture-generation artefacts - yielding nine wrong-meaning tokens removed from production deterministic-rule keyword sets.
Interpretation.
Six classes of silent evaluator artefacts (five runtime, one fixture-generation) were identified in this existence-proof methodology study, not prevalence estimate; most would have led to incorrect clinical conclusions without the corresponding audit. We recommend six evaluation-practice additions: paired cross-language control personas as compatibility canaries, fallback-path instrumentation, parse-provenance disaggregation, coherence audits, instructed-rule-compliance auditing, and native-speaker test-fixture audit. A deterministic safety-rule layer provides model-agnostic floor behaviour where rule coverage holds.
Funding.
No external funding; pilot under consideration by J-PAL EVAH Pathway A.