Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

When the Loop Closes Quietly: Mechanistic Interpretability as Contestability Infrastructure for Deployed AI

Domaine:

digital infrastructure

Type de record:

paper
Créateur:
Och
Éditeur:
Zenodo
Hôte:avatar

Deployed AI decision-support systems can satisfy every local performance metric while quietly eroding the conditions under which their outputs can be contested. Between 2024 and 2025, I led the rollout of an AI decision-support system across six counties in a sub-Saharan African country. By every measured indicator the deployment succeeded: faster processing, fewer errors, more consistent outputs. What the metrics did not capture was that officials had stopped contesting outputs they could not evaluate, and citizens had stopped contesting decisions they could not trace to any human decision-maker. Nothing collapsed. The loop had closed.

This paper argues that the failure mode is constitutive of opacity, not incidental to it, and that mechanistic interpretability — circuit analysis, feature attribution, sparse autoencoders — should be understood as contestability infrastructure, not research instruments alone. It maps specific failure modes to specific interpretability primitives, develops the consequences for how the field prioritizes work and what governance regimes can demand, states what interpretability cannot solve, and closes with four falsifiable predictions that would confirm or disconfirm the closed-loop dynamic across deployments. The contribution is a position grounded in deployment experience; the evidence is observational, intended to generate a testable hypothesis rather than prove one.

Visit

doi.org

Tags

mechanistic interpretabilitycontestabilityAI governancedeployed AI systemsalgorithmic accountabilityGlobal South AI deploymentinterpretability applicationsAI safetyfaithfulnesscausal interventions

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode