Biophysical core life systems in Kenya, including diagnostic imaging networks and physiological monitoring infrastructure, increasingly rely on machine learning models that assume training and deployment distributions coincide. We introduce a benchmark protocol for out-of-distribution generalisation that distinguishes covariate shift arising from sensor degradation, domain shift from population heterogeneity, and concept shift from evolving clinical protocols. The central contribution is a theoretical error decomposition theorem that separates representation error, causal mechanism error, and distributional extrapolation error, enabling principled diagnosis of model failures in deployment. We prove identifiability conditions for latent causal variables under a sparsity assumption on intervention targets, and we derive sample complexity bounds for estimating the causal graph from observational biophysical time series. An error analysis taxonomy is proposed to guide practitioners in attributing performance degradation to remediable causes rather than treating all distribution shift as equivalent. The framework is designed to support prospective evaluation in Kenyan referral hospitals without requiring access to private patient data, and it establishes the theoretical groundwork for subsequent empirical validation.