Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery

Domain:

healthcare

Record type:

paper
Creator:
A RC. A BG M
Publisher:
Oxf
Host:
Abstract Large language models (LLMs) such as OpenAI o1 and DeepSeek-R1 are designed to move beyond factual recall by modelling deliberative thought processes. Their performance in complex spine scenarios remains unclear. This study evaluates whether reasoning-tuned LLMs can generate coherent and clinically relevant management plans when assessed by fellowship-trained spine surgeons. Eleven synthetic case vignettes were presented to OpenAI o1 (full) and DeepSeek R1 with identical prompts requesting diagnostic impression, reasoning, and management plan. Outputs were anonymised and randomised for blind review. Eight fellowship-trained spine surgeons [five consultants, three fellows] from the United Kingdom, Switzerland, Nigeria, and Zambia scored diagnostic accuracy, reasoning, surgical plan appropriateness, and clarity on five-point Likert scales. Eighty-six paired evaluations were analysed using two-tailed paired t-tests with Bonferroni correction, adjusted α=0.0125. OpenAI o1 (full) outperformed DeepSeek R1 across all domains. Means [Standard Deviation] and p values were diagnostic accuracy 4.57 [0.60] vs 4.31 [0.79], p < 0.001, reasoning and thoroughness 4.48 [0.68] vs 4.23 [0.75], p = 0.005, surgical plan appropriateness 4.33 [0.76] vs 4.06 [0.86], p = 0.010, clarity 4.47 [0.68] vs 4.15 [0.85], p < 0.001. All comparisons met the corrected significance threshold, and o1 showed lower standard deviations, which signals more consistent quality across raters and cases. Reasoning-tuned LLMs can emulate elements of expert surgical decision-making. OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs. Transparent validation and reporting remain essential before clinical use.

Visit

doi.org

Licenses

https://academic.oup.com/pages/standard-publication-reuse-rights

Similar

Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity EvaluationCendol: Open Instruction-tuned Generative Large Language Models for Indonesian LanguagesEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningEvaluating Multilingual Long-Context Models for Retrieval and ReasoningAssessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces

Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian Languages

Large language models (LLMs) show remarkable human-like capability in various domains and languages.

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense Reasoning

Commonsense reasoning research has so far been limited to English. We aim to evaluate and improve popular multilingual language models (ML-LMs) to help advance commonsense reasoning (CSR) beyond English. We collect the Mickey Corpus, consisting of 561k sentences in

Evaluating Multilingual Long-Context Models for Retrieval and Reasoning

Recent large language models (LLMs) demonstrate impressive capabilities in handling long contexts, s

Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

Language is not monolithic. While benchmarks, including those designed for multiple languages, are o