Abstract / Summary
Multimodal large language models (LLMs) show promise in medical imaging by integrating text and image data to help diagnose, improve consistency, and streamline workflows. They can approach specialist-level performance, support clinical decision-making, and assist in metabolic bariatric surgery (MBS) imaging, though applications remain early. Although LLMs perform well on exams, their real-world readiness still demands thorough validation before they can be safely adopted at scale. This study compared ChatGPT 5.2, Grok-4.1, and Gemini-3 Pro in diagnosing postoperative complications after MBS using 12 radiologic scenarios (six unpublished institutional cases and six PubMed-sourced cases). Each LLM received standardized prompts and case imaging. Model responses were compared with the reference diagnoses using top-1 concordance and top-3 inclusion as exploratory performance measures. Across 12 scenarios, ChatGPT 5.2 achieved the highest overall score (18/24), followed by Gemini-3 Pro (17/24) and Grok-4.1 (17/24), with no model achieving full concordance. All models demonstrated better performance in identifying structural and anatomical complications than functional or vascular complications. Published cases showed better observed diagnostic performance than previously unpublished cases; however, this difference should be interpreted cautiously because publicly available cases may have been represented in model training data, and therefore the observed advantage may reflect both recognition of well-documented clinical patterns and potential prior exposure to published case content. In this proof-of-concept evaluation, multimodal LLMs demonstrated potential to support differential diagnosis of postoperative MBS complications but showed variable performance, particularly for less common vascular and bowel-related complications. Higher observed scores for published than unpublished cases should be interpreted cautiously because prior exposure to publicly available case content during model training cannot be excluded and therefore cannot be considered evidence of superior diagnostic accuracy or generalizability. These findings support further evaluation using larger datasets of demonstrably unseen cases before clinical implementation.