Federated learning offers a promising architecture for training clinical machine learning models across distributed Ugandan health facilities without centralising sensitive patient data. However, the practical viability of this approach depends on its robustness to two conditions that characterise real-world health informatics environments: distribution shift across heterogeneous facility populations and extreme label sparsity in routine clinical documentation. We establish convergence guarantees for the proposed algorithm under bounded distribution shift and demonstrate that the interaction between client heterogeneity and label scarcity produces a distinct failure mode—termed silent model divergence—that conventional federated averaging fails to detect. The analysis yields three principal contributions: a formal characterisation of the robustness-accuracy trade-off under shift and sparsity, a practical diagnostic criterion for identifying clients whose local data distributions have drifted beyond safe aggregation bounds, and a theoretical justification for prioritising certain clinical prediction tasks in low-resource settings. These findings provide a mathematical foundation for designing privacy-preserving clinical decision support systems that remain reliable as Ugandan health facilities scale their digital infrastructure.