Background
Cervical cancer remains a leading cause of cancer death among women in sub-Saharan Africa, with Tanzania bearing a disproportionate burden. The critical shortage of trained pathologists, coupled with the unprecedented disease burden in low-resource settings, underscores the urgent need for point-of-care screening and diagnosis to enable timely decision-making. We assessed the diagnostic accuracy of AI-driven cytopathological tools to improve diagnostic efficiency and accessibility.
Methods
This retrospective secondary data analysis evaluated five convolutional neural network (CNN) architectures: EfficientNetB7, MobileNet, ResNet50, ResNet152, and InceptionNetV3, for semi-automated classification of cervical cell abnormalities. A total of 11,955 Pap smear cytological images were used from the publicly available Center for Recognition and Inspection of Cells (CRIC) Searchable Image Database, spanning six cellular classes: Normal, ASC-US, LSIL, ASC-H, HSIL, and carcinoma. Performance was assessed on a hold-out test set (
n
= 961) using macro-averaged metrics.
Results
EfficientNetB7 achieved the highest overall performance, with a macro F1 score of 0.9324 (95% CI: 0.920–0.945), an accuracy of 0.9775 (95% CI: 0.968–0.987), and a sensitivity of 0.9324 (95% CI: 0.920–0.945). ResNet50 ranked second (F1 score: 0.9282 [95% CI: 0.916–0.941]) and ResNet152 third (F1 score: 0.9240 [95% CI: 0.911–0.937]), showing minimal performance gaps. InceptionNetV3 followed closely (F1 score: 0.9220 [95% CI: 0.909–0.935]). MobileNet achieved the lowest F1 score (0.8918 [95% CI: 0.875–0.908]), but its lightweight architecture is suitable for edge deployment. The Carcinoma (CA) class achieved near-perfect recall across all models (> = 0.978). Notable interclass confusion was observed between ASC-US and LSIL, attributable to cytomorphological overlap.
Conclusion
EfficientNetB7 offers promising diagnostic accuracy for automated cervical cancer screening using Pap smear images and shows potential for integration into point-of-care workflows in low-resource settings. Future work should focus on training models on locally representative datasets and exploring whole smear analysis to reduce interclass misclassification.