Abstract
Diabetic retinopathy (DR) is a major cause of preventable vision loss, but timely screening remains difficult in settings with limited retinal specialists and weak diabetes–eye-care pathways. This study developed and evaluated a lesion-guided, explainable, and uncertainty-aware deep learning framework for five-class DR grading, with relevance to screening support in Uganda. The pipeline combined a U-Net–ResNet34 lesion-segmentation module, a five-channel lesion-guided CNN–ViT classifier, quantitative explanation analysis using CAM, Grad-CAM, and Information Bottleneck Attribution against lesion-support regions, and Monte Carlo Dropout for selective referral of uncertain cases. Experiments used a public grading dataset of 3,554 fundus images and IDRiD lesion-annotated subsets for hard-exudate and haemorrhage segmentation. The lesion module achieved Dice scores of 0.5344 for hard exudates and 0.4437 for haemorrhages. On the held-out test set, the hybrid classifier achieved 99.25% accuracy, QWK 0.9884, macro precision 98.86%, macro recall 99.39%, and macro F1-score 99.12%, outperforming lesion-guided VGG16, ResNet50, and ViT baselines. CAM showed the strongest lesion consistency (mean IoU 0.0226 ± 0.0162), while Grad-CAM and IBA remained near 0.0050. Predictive entropy was higher for incorrect than correct predictions (0.3302 vs 0.1047), calibration was strong (ECE 0.0289; Brier score 0.0150), and at a validation-derived operating point targeting ~ 15% referral, 85.58% of cases were automated while the accepted subset achieved 100% accuracy, QWK, and macro F1-score. These findings support the value of combining lesion priors, quantitative explanation checks, and calibrated uncertainty, although prospective local validation is still required before use in Ugandan screening practice.