Abstract / Summary
Generative AI chatbots are increasingly used for health information, but their performance in safety-sensitive sudden cardiac death (SCD) consultation remains insufficiently characterized. This study evaluated six publicly accessible chatbots on SCD-related public questions involving symptom triage, emergency response, CPR/AED use, inherited risk, myocarditis-related concerns, return to exercise, screening, and ICD decision-making. This cross-sectional, comparative, text-level study evaluated ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, Doubao, and Perplexity. A standardized set of 56 SCD-related public consultation questions was developed from Google Trends, online forums, public question-answer platforms, clinical guidelines, expert consensus statements, systematic reviews, and qualitative interviews. Each question was submitted once to each chatbot between May 10 and May 16, 2026, yielding 336 model-response items. Five blinded senior cardiology raters assessed safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, and GQS. Readability was calculated using six formula-based indices. Between-model comparisons were performed using Cochran’s Q test and Friedman tests, with Kendall’s W reported as the effect size. Responses containing potential safety concerns were additionally examined qualitatively. Inter-rater agreement was high across all manually assessed metrics, with Fleiss’ kappa of 0.856 for Safety and ICC values ranging from 0.801 to 0.887 for other domains. Safe responses predominated across all models, although responses containing potential safety concerns occurred in every chatbot. ChatGPT Plus 5.5 thinking and Perplexity had the highest observed safety proportions during this query window, whereas Doubao had the lowest. Qualitative analysis identified recurrent concerns involving delayed emergency activation for chest pain, syncope, or post-viral symptoms; over-reassurance about prevention; fixed return-to-exercise timelines after infection or COVID-19; pulse-check instructions that could delay CPR; and oversimplified screening, electrolyte, or ICD-related advice. Most flagged responses contained a single potential safety concern. The qualitative severity distribution was concentrated in low and low-to-moderate concerns, whereas higher-severity concerns were less frequent. Accuracy differed significantly across models ( P < 0.001; Kendall’s W = 0.864), with ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, and Perplexity achieving similar median scores and Doubao performing lower. Empathy also differed significantly ( P < 0.001; Kendall’s W = 0.531), with DeepSeek-V4-pro having the highest observed score. Significant differences were observed for DISCERN, EQIP, JAMA, GQS, and all readability indices (all P < 0.001). Within this dataset, DeepSeek-V4-pro had the highest DISCERN and EQIP scores and the most favorable readability profile, Gemini 3.5 Flash showed comparable EQIP performance, and Microsoft Copilot smart had the highest JAMA score. However, transparency remained limited overall, and all models produced responses with relatively high reading-grade requirements. In this multidimensional evaluation of six generative AI chatbots for SCD-related public consultation, most responses were safe and clinically relevant, but potential safety concerns occurred across all models. Important limitations involved emergency escalation, CPR/AED guidance, return-to-exercise advice, individualized risk interpretation, transparency, and readability. These tools may support general SCD education, but they should not replace emergency medical services or clinician assessment. Because each question was sampled once during a defined week and all prompts were submitted in English through public interfaces accessed in China, the findings should be interpreted as a time-, language-, and access-specific snapshot rather than a stable ranking of the underlying models. Future development should prioritize evidence-linked content, clear red-flag escalation, plain-language communication, and ongoing expert review.