The Area Under the Receiver Operating Characteristic Curve (AUC) performance of machine learning models trained and tested on data from single sites. Performance is compared between internal validations (training and testing on the same site) and external validations (training on one site and testing on others). Five distinct study sites in sub-Saharan Africa are evaluated: Agincourt (South Africa), Ifakara (Tanzania), Iganga (Uganda), Kilifi (Kenya), and Kintampo (Ghana). Boxplots describe the distribution of AUC values obtained through bootstrap resampling, indicating the variance within internal and external validations. At Agincourt, the internal AUC is 0.96 (interquartile range (IQR) 0.95–0.97), while the external AUC is 0.88 (IQR 0.86–0.88), demonstrating a statistically significant difference with a decrease of approximately 0.08 in performance when models are externally validated. Similarly, Kilifi shows an internal AUC of 0.97 (IQR 0.96–0.97) against an external AUC of 0.84 (IQR 0.78–0.89), indicating a significant decline in external validation performance. Iganga’s internal and external AUCs are 0.93 (0.91–0.96) and 0.88 (IQR 0.87–0.96), displaying a smaller yet significant discrepancy. In contrast, Ifakara and Kintampo exhibit a converse trend, where external AUCs 0.92 (IQR 0.89–0.96) and 0.92 (IQR 0.91–0.95) slightly exceed their internal counterparts 0.89 (0.89–0.91) and 0.90 (0.88–0.92), although these differences are also statistically significant. These findings underscore the variability in model generalizability and the importance of external validation when assessing the robustness of predictive models in healthcare settings. *** = p-value <0.001.