Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multimodal large language models versus physicians for real-world inpatient diagnosis: Retrospective comparative (Preprint)

Domaine:

healthcare

Type de record:

paper
Créateur:
BruAmySitMic
Éditeur:
JMI
Hôte:
BACKGROUND Large language models (LLMs) are increasingly proposed for diagnostic support. However, validation of their performance in real-world multimodal inpatient care remains limited, particularly in low- and middle-income country (LMIC) hospital settings. OBJECTIVE This study aimed to evaluate the diagnostic accuracy and clinical reasoning of leading multimodal LLMs compared to routine clinical practice within a South African tertiary hospital. METHODS We conducted a retrospective evaluation of 10 leading multimodal LLMs using 539 complete inpatient cases from a South African tertiary public hospital. The dataset included radiology imaging, radiology reports, laboratory results, and clinical narratives. Expert clinician panels adjudicated 300 cases to establish reference diagnoses, differential diagnoses, and clinical reasoning. Model outputs and routine recorded ward diagnoses were scored using a panel-calibrated, three-model LLM Jury. RESULTS Mean LLM scores were tightly clustered despite 50-fold cost differences, and on average all models exceeded routine ward diagnostic performance (>99.9% bootstrap probability). In a senior clinician-adjudicated tie-breaker subset maximised for model–panel disagreement, a predefined Top-3 LLM ensemble exceeded the original expert-panel consensus with 97% bootstrap probability. Residual safety risks and model non-response motivate prospective evaluation with rigorous adjudication and clinical oversight. CONCLUSIONS Multimodal LLMs demonstrated high diagnostic performance in a complex, real-world LMIC inpatient setting, surpassing recorded routine ward diagnoses in 60-80% of cases depending on the model. Our finding that even the cheapest models outperform the routine ward diagnoses and are comparable to the best models suggests that affordable AI clinical decision support in resource-constrained environments is now worth serious consideration.

Visit

doi.org

Similaires

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier ModelsAI on the Frontline: Evaluating Large Language Models in Real-World Conflict ResolutionMultimodal adaptive cross-mamba for tomato leaf disease diagnosis built on real-world agricultural dataM3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language ModelsLarge Language Models as Clinical Decision Support for Diagnosis: Benefit After AI-Literacy Training Versus Hallucination and Automation BiasTowards Multimodal Cultural Context Modeling for African Languages in Large Language Models

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few e

AI on the Frontline: Evaluating Large Language Models in Real-World Conflict Resolution

This groundbreaking study authored by Nathalie Bussemaker and Mark Freeman and published by the Inst

Multimodal adaptive cross-mamba for tomato leaf disease diagnosis built on real-world agricultural data

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Despite the existence of various benchmarks for evaluating natural language processing models, we ar

Large Language Models as Clinical Decision Support for Diagnosis: Benefit After AI-Literacy Training Versus Hallucination and Automation Bias

This narrative review examines the conditions under which physicians benefit from large language mod

Towards Multimodal Cultural Context Modeling for African Languages in Large Language Models

This preliminary work addresses the critical gap in multimodal Large Language Models (LLMs) for Afri