Abstract / Summary
Large language models can synthesize clinical records, identify phenotypic features and support patient-specific decisions, but fluent explanations may obscure uncertainty. Surgical practice provides a demanding setting in which to examine this problem because decisions combine heterogeneous evidence, procedural risk and individual preferences. This critical narrative review connects surgical self-assessment research with studies of clinician–LLM interaction, clinical summarization, uncertainty estimation and perioperative implementation. Targeted searches of PubMed records through Europe PMC were supplemented by retrieval of methodological and foundational sources. Selection was purposive, with attention to contradictory findings and the distinction between simulated performance and clinical benefit. Randomized studies using general clinical vignettes and simulated tasks show that LLM assistance does not uniformly improve physician reasoning; their applicability to surgical decisions remains indirect. Direct clinical studies extend from preoperative documentation and triage to shared decision-making and postoperative support, with mixed findings and limited evidence of sustained patient benefit. We distinguish clinician self-appraisal, LLM uncertainty, prognostic calibration, appropriate reliance and patient-specific decision quality, each requiring distinct evaluation criteria. For personalized surgical care, generated interpretations should preserve the provenance, timing and limitations of the underlying patient data and predictive models. We propose a workflow for examining discrepancies between clinician judgment and LLM output, together with an evaluation agenda centered on harmful decision changes, verification burden and patient-relevant outcomes. This framework requires prospective validation.