Logo Lanfrica

Multimodal large language models versus physicians for real-world inpatient diagnosis: Retrospective comparative (Preprint)

Domain:

healthcare

Record type:

paper
Creator:
BruAmySitMic
Publisher:
JMI
Host:
BACKGROUND Large language models (LLMs) are increasingly proposed for diagnostic support. However, validation of their performance in real-world multimodal inpatient care remains limited, particularly in low- and middle-income country (LMIC) hospital settings. OBJECTIVE This study aimed to evaluate the diagnostic accuracy and clinical reasoning of leading multimodal LLMs compared to routine clinical practice within a South African tertiary hospital. METHODS We conducted a retrospective evaluation of 10 leading multimodal LLMs using 539 complete inpatient cases from a South African tertiary public hospital. The dataset included radiology imaging, radiology reports, laboratory results, and clinical narratives. Expert clinician panels adjudicated 300 cases to establish reference diagnoses, differential diagnoses, and clinical reasoning. Model outputs and routine recorded ward diagnoses were scored using a panel-calibrated, three-model LLM Jury. RESULTS Mean LLM scores were tightly clustered despite 50-fold cost differences, and on average all models exceeded routine ward diagnostic performance (>99.9% bootstrap probability). In a senior clinician-adjudicated tie-breaker subset maximised for model–panel disagreement, a predefined Top-3 LLM ensemble exceeded the original expert-panel consensus with 97% bootstrap probability. Residual safety risks and model non-response motivate prospective evaluation with rigorous adjudication and clinical oversight. CONCLUSIONS Multimodal LLMs demonstrated high diagnostic performance in a complex, real-world LMIC inpatient setting, surpassing recorded routine ward diagnoses in 60-80% of cases depending on the model. Our finding that even the cheapest models outperform the routine ward diagnoses and are comparable to the best models suggests that affordable AI clinical decision support in resource-constrained environments is now worth serious consideration.

Similar