Sweet & Bitter

Discordance Between Textual Reasoning and Visual Interpretation in Large Language Models for Low Back Pain: Cross-Sectional Quantitative Evaluation and Exploratory Multimodal Stress Test.

Key Takeaway

BACKGROUND: Large language models (LLMs) are rapidly evolving from text-based agents to multimodal systems capable of interpreting medical images. While their textual reasoning has improved, the safety implications of this shift remain underexplored, specifically regarding the alignment between visu

BACKGROUND: Large language models (LLMs) are rapidly evolving from text-based agents to multimodal systems capable of interpreting medical images. While their textual reasoning has improved, the safety implications of this shift remain underexplored, specifically regarding the alignment between visual interpretation and textual advice in low back pain (LBP) management. OBJECTIVE: This study aims to evaluate the performance of 10 commonly used and highly representative LLMs, including paid models such as ChatGPT 5 Plus (OpenAI) and reasoning-oriented models like Qwen3-Thinking (Alibaba/Qwen team) in the context of LBP. Specifically, it assesses their clinical accuracy, readability, and the "asymmetric evolution" in cross-modal diagnostic consistency. METHODS: We conducted a cross-sectional quantitative evaluation integrated with an exploratory multimodal stress test. First, 10 LLMs, including ChatGPT 5 and ChatGPT 5 Plus, Gemini 3 Pro (Google DeepMind), Claude 4.5 Sonnet (Anthropic), Grok 4 (xAI), DeepSeek, Kimi K2 (Moonshot AI), Doubao (ByteDance Seed [Doubao Team]), Qwen3 (Base), and Qwen3-Thinking, addressed 25 standardized LBP inquiries. A multidisciplinary expert panel extracted 1816 individual recommendations by segmenting the models' point-by-point responses. These were independently coded for clinical accuracy against international guidelines on a 5-point Likert scale. Outcomes also included readability (Flesch Reading Ease [FRE] and Simple Measure of Gobbledygook [SMOG]), understandability (Patient Education Materials Assessment Tool [PEMAT]), and the presence of disclaimers. Second, to probe multimodal capabilities, we conducted an exploratory stress test (N=5) using a curated set of challenging clinical cases, including representative but complex LBP pathologies to assess baseline image recognition, and adversarial cases with deliberate clinical-radiological mismatches to evaluate cross-modal alignment. RESULTS: In textual tasks, models achieved an overall accuracy of 88.4% (1605/1816). ChatGPT 5 Plus demonstrated near-perfect performance, with zero serious errors. Within the only directly matched reasoning comparison, Qwen3-Thinking showed a lower serious-error rate than Qwen3 (Base). However, textual advice remained difficult to read, although understandability scores were robust. In contrast, multimodal performance in the exploratory stress test exhibited a negative divergence. In pure imaging tasks featuring complex but clinically representative pathologies, the average diagnostic score was poor, with models failing to identify core pathologies. In comprehensive diagnostic analysis tasks (Q29-Q30), models exhibited "textual masking," aligning visual findings with text cues rather than imaging evidence. Crucially, safety disclaimer coverage dropped precipitously from 82.8% (207/250) in textual tasks to 28% (14/50) in multimodal interactions. CONCLUSIONS: Under adversarial stress conditions, current LLMs exhibit a capability mismatch: expert-level textual accuracy masks vulnerable visual interpretation. This "high-confidence trap" may cause users to misapply trust in textual logic to visual diagnostics. Although the matched Qwen3 comparison suggested a lower serious-error rate in the reasoning-enhanced setting, poor cross-modal alignment and the systemic absence of medical disclaimers in visual tasks pose substantial risks. Thus, current multimodal LLMs remain unready for direct patient use in LBP management.

Source

Zhang, Ziyu; Chen, Longhao; Lv, Zhizhen; Lv, Hanzhe; Sheng, Wei; Wei, Zicheng; Wang, Binghao; Shen, Ying; Tian, Yu; Hu, Jingwen; Shen, Zhifang; Lv, Lijiang. JMIR medical informatics, 2026. DOI: 10.2196/93522

Share

You may also like