Abstract / Summary
Background: Large Language Models (LLMs) are increasingly utilized for patient-facing healthcare communication. However, their clinical utility remains constrained by communicative, ethical, and safety challenges. This systematic review provides a cross-specialty evaluation of conversational AI, synthesizing evidence on accuracy, readability, informational quality, and emotional resonance. Methods: Following PRISMA 2020 guidelines and PROSPERO registration (CRD420251021428), PubMed/MEDLINE, Scopus, and Web of Science were searched from 2020 through November 2025. Eligible studies evaluated LLM-generated medical or surgical information against healthcare professionals, validated benchmarks, or alternative AI platforms. Assessed outcomes included factual accuracy, standardized readability indices (e.g., FKGL, FRES), validated quality metrics (e.g., DISCERN, GQS), and emotional or empathetic impact. Results: Eighty-two studies (published 2023–2025) were included. Accuracy was evaluated in 45 studies, demonstrating generally high performance (mean Likert scores > 3.5/5); advanced GPT models frequently showed parity or superiority to comparators, though factual errors persisted in complex clinical reasoning. Readability, assessed in 69 studies, emerged as the primary limitation: baseline outputs consistently required high school or university-level comprehension (FKGL > 10–12), substantially exceeding recommended patient literacy thresholds (≤8th grade). Informational quality, appraised in 35 studies, was rated moderate-to-high across DISCERN and GQS metrics. Six studies evaluating emotional resonance indicated that LLMs can alleviate situational anxiety and deliver structured cognitive empathy comparable or superior to brief clinician text, though lacking genuine emotional depth. Conclusions: LLMs demonstrate strong factual accuracy and structural quality across diverse clinical disciplines. However, autonomous, patient-facing deployment remains premature due to excessive lexical density that threatens health equity. Conversational AI should function as a clinical adjunct to support providers rather than an unsupervised educational tool. Future research must prioritize dynamic literacy alignment and robust clinical validation.