
Deployed AI decision-support systems can satisfy every local performance metric while quietly eroding the conditions under which their outputs can be contested. Between 2024 and 2025, I led the rollout of an AI decision-support system across six counties in a sub-Saharan African country. By every measured indicator the deployment succeeded: faster processing, fewer errors, more consistent outputs. What the metrics did not capture was that officials had stopped contesting outputs they could not evaluate, and citizens had stopped contesting decisions they could not trace to any human decision-maker. Nothing collapsed. The loop had closed.
This paper argues that the failure mode is constitutive of opacity, not incidental to it, and that mechanistic interpretability — circuit analysis, feature attribution, sparse autoencoders — should be understood as contestability infrastructure, not research instruments alone. It maps specific failure modes to specific interpretability primitives, develops the consequences for how the field prioritizes work and what governance regimes can demand, states what interpretability cannot solve, and closes with four falsifiable predictions that would confirm or disconfirm the closed-loop dynamic across deployments. The contribution is a position grounded in deployment experience; the evidence is observational, intended to generate a testable hypothesis rather than prove one.