BACKGROUND
Comparative evaluation of machine-learning models for health surveillance data can be distorted by asymmetric random-seed evaluation, single holdout splits, weak hyperparameter tuning, inappropriate missing-data handling, duplicate predictor profiles, and temporal leakage.
OBJECTIVE
We compared 11 classical, neural, and feature-token Transformer classifiers for diphtheria and dengue surveillance data from Yemen using symmetric repeated-seed evaluation, group-aware validation, temporal testing, calibration assessment, decision-curve analysis, and a selection-bias audit.
METHODS
: The diphtheria analytic cohort included 566 records with definitive culture results from a source line list of 2,035 suspected or notified cases. Models were evaluated using identical 5-fold outer and 3-fold inner StratifiedGroupKFold cross-validation and the same 5 random seeds. For dengue, 118 records from 27 predictor-identical groups with conflicting labels were excluded from the primary analysis. Data from 2017-2018 were used for development (n=3,672), and 2019 was reserved as a locked temporal test set. One 2019 record with an exact predictor profile observed during development was excluded from the strict primary test (n=2,314). Hyperparameters for final temporal models were selected using group cross-validation within 2017-2018 only. Macro-F1 was the primary metric. Additional measures assessed classification performance, discrimination, calibration, and decision-curve net benefit. For the primary confidence-interval and paired-comparison analyses, probabilities from the 5 seed-specific fits were averaged at the observation level to form a 5-model probability ensemble. Group-bootstrap 95% CIs and paired group-bootstrap differences were then calculated from these ensemble predictions.
RESULTS
For diphtheria, Deep MLP had the highest 5-seed ensemble Macro-F1 (0.6372, 95% CI 0.5962-0.6757), followed by random forest (0.6263, 95% CI 0.5879-0.6665) and the feature-token Transformer (0.5897, 95% CI 0.5468-0.6322). The Deep MLP-minus-Transformer difference was 0.0475 (95% CI 0.0078-0.0871). Random forest had the highest ROC-AUC (0.6718), the lowest Brier score (0.2278), and a calibration slope closest to 1 (0.8499) among the leading models. The diphtheria selection audit showed substantial differences between retained and excluded records. For dengue, positive prevalence increased from 46.8% during development to 79.8% in the strict 2019 test. Internal performance was high, but all models had lower Macro-F1 on the temporal test. XGBoost performed best in 2019, with a 5-seed ensemble Macro-F1 of 0.5109 (95% CI 0.4886-0.5328), compared with 0.4421 (95% CI 0.4204-0.4639) for the Transformer; the paired difference was 0.0688 (95% CI 0.0523-0.0852). Both models were poorly calibrated after temporal shift.
CONCLUSIONS
The feature-token Transformer did not outperform the strongest comparator on either task. Strong internal dengue performance did not translate to the subsequent calendar period, and the diphtheria selection audit limits inference to cases with classifiable culture results. Prospective recalibration, external validation, and setting-specific utility assessment are required before deployment.
CLINICALTRIAL
na