Abstract / Summary
As the population ages and the prevalence of multiple coexisting conditions increases, end-of-life decision-making for older adults is becoming increasingly complex in clinical practice. limited access to relevant medical information and differences in medical knowledge between clinicians and surrogate decision-makers. Large language models (LLMs) are being widely used for accessing health information and facilitating doctor-patient communication, but their value and risks in end-of-life decision-making support for older adults have yet to be systematically evaluated. This study aimed to benchmark the safety, accuracy, observable empathetic communication, information quality, transparency indicators, and readability of responses generated by five publicly accessible consumer-facing generative AI systems to standardized end-of-life questions relevant to surrogate decision-makers for older adults. This study is a cross-sectional benchmark evaluation using a 30-question standardized question set covering six prespecified end-of-life domains. The question set was developed using an evidence-informed approach incorporating clinical guidelines, consensus statements, relevant literature, published qualitative studies involving surrogate decision-makers and family caregivers, and public search terminology. Each question was submitted once, in a new session, to five consumer-facing generative AI systems under their default web-interface conditions on June 14, 2026, yielding 150 paired system–question outputs. Two blinded clinical experts independently evaluated the outputs using predefined criteria for safety, accuracy, textual empathy, DISCERN, EQIP, JAMA transparency benchmarks, GQS, and six readability indices. Inter-rater reliability, paired overall comparisons, effect sizes, and post hoc comparisons with Benjamini–Hochberg correction were assessed. Safe-response rates ranged from 76.7% to 86.7% across systems, with no evidence of an overall between-system difference in safety (Cochran’s Q = 1.368, df = 4, P = 0.850; effect size = 0.011, 95% CI 0.005–0.117). Accuracy differed significantly across systems, although the overall effect was small ( P = 0.003; Kendall’s W = 0.132, 95% CI 0.036–0.332). Differences in expert-rated textual empathy were more pronounced ( P < 0.001; Kendall’s W = 0.373, 95% CI 0.194–0.598). Significant between-system differences were also observed for DISCERN ( P < 0.001; W = 0.332), EQIP ( P = 0.016; W = 0.102), JAMA transparency benchmarks ( P < 0.001; W = 0.751), and GQS ( P = 0.001; W = 0.148). Readability differed significantly across systems across all six indices (all P < 0.001), with substantial variability in grade-level estimates. Inter-rater agreement was good to excellent across the evaluated dimensions. In this time-stamped benchmark, consumer-facing generative AI systems generally produced accurate responses with observable empathetic communication, but clinically relevant limitations remained in safety, transparency, and readability. The sampled systems did not differ significantly in safety, whereas differences in accuracy were small and differences in textual empathy were more pronounced. Because the study evaluated generated text rather than surrogate understanding, trust, preferences, decisions, decisional conflict, behavior, or clinical outcomes, the findings do not establish the effectiveness of these systems as end-of-life decision-support interventions. Consumer-facing generative AI should therefore not be used as a standalone information source for high-stakes end-of-life decisions.