Abstract / Summary
Background/Objectives: Treatment allocation for hepatocellular carcinoma (HCC) is a highly complex task that requires multidisciplinary tumor board (MDT) input; however, its accessibility and reliability can vary across the healthcare system. Large language models (LLMs) have emerged as potential clinical decision-making tools to aid MDTs. Our aim was to compare different LLM-generated treatment recommendations for HCC cases with MDT decisions. Methods: We retrospectively analyzed 100 HCC cases discussed during MDT meetings in a tertiary-care hospital in Cluj-Napoca, Romania. Identical prompts and structured clinical information were offered to four different LLMs (ChatGPT-5, a customized Tumor Board ChatGPT, Gemini 2.5 Flash, and Gemini 2.5 Pro), each of which was required to provide treatment recommendations. Concordance with the MDT decisions was assessed across the first, second, and third treatment options. Inter-model differences were evaluated using Cochran’s Q and Holm-adjusted exact McNemar tests. Chance-corrected agreement was assessed using Cohen’s kappa (κ), and generalized estimating equations evaluated associations between concordance and clinical complexity. Results: First-recommendation concordance was 81% (95% CI 72.2–87.5) for GPT-5, 80% (71.1–86.7) for Gemini 2.5 Pro, 78% (68.9–85.0) for TumorBoard ChatGPT, and 63% (53.2–71.8) for Gemini 2.5 Flash; cumulative top-3 concordance reached 95%, 92%, 95%, and 86%, respectively. GPT-5 (κ = 0.754) and Gemini 2.5 Pro (κ = 0.743) showed substantial chance-corrected agreement. Gemini 2.5 Flash showed significantly lower concordance, including after adjustment for BCLC stage, Child–Pugh class, tumor burden, and treatment category (p = 0.006), while these clinical complexity variables were not significantly associated with concordance. Conclusions: LLMs showed moderate-to-substantial agreement with MDT decisions, with meaningful inter-model differences, for a complex disease, such as HCC. Our study highlights that, while the need for human oversight is essential, LLMs could serve as supportive tools for expert guidance in some clinical settings and underscore AI’s potential in enhancing liver cancer care.