Abstract / Summary
Abstract Background Pediatric burn injuries are common household injuries, particularly among young children, and hot-liquid scalds are a major cause of pediatric burn morbidity. The quality of burn first aid provided by caregivers can affect pain, burn wound depth, wound healing, the need for surgical intervention, and scarring. As large language model (LLM)-driven chatbots become widely available sources of health information, caregivers may consult these tools in time-critical situations requiring burn first aid; however, their performance in pediatric burn scenarios remains insufficiently characterized. Objective To compare the safety, accuracy, empathy, information quality, and readability of four widely accessible LLM chatbots in answering caregiver-facing questions about pediatric burn first aid and inpatient care. Methods In this cross-sectional comparative evaluation, a 36-item question set was developed from an initial pool of 66 caregiver-facing questions derived from clinical guidelines, consensus statements, official patient-education materials, systematic reviews, and common pediatric burn care concerns. The same questions were entered independently into ChatGPT, Gemini, DeepSeek, and Doubao, yielding 144 chatbot responses. Five blinded raters evaluated outputs using predefined safety and accuracy criteria, an empathy scale, DISCERN, EQIP, JAMA benchmark criteria, and the Global Quality Score. Readability was assessed using the Automated Readability Index, Coleman-Liau Index, Flesch-Kincaid Grade Level, Gunning Fog Index, Simple Measure of Gobbledygook, and Flesch Reading Ease Score. Inter-rater agreement and paired question-level comparisons were examined. Results Inter-rater agreement was high for safety assessment (Fleiss κ = 0.860; 95% CI, 0.784–0.924; P < 0.001) and for the other rater-scored outcomes ( ICC range, 0.831–0.899; all P < 0.001). Responses were classified as safe in 33 of 36 cases for ChatGPT (91.7%), 31 of 36 for Gemini (86.1%), 30 of 36 for DeepSeek (83.3%), and 29 of 36 for Doubao (80.6%). Accuracy scores differed across models ( P = 0.002), although the overall effect size was small (Kendall's W = 0.137). Empathy scores showed a larger between-model difference ( P < 0.001; Kendall's W = 0.564), with Doubao receiving lower ratings than the other models. DISCERN, EQIP, and Global Quality Scores differed across models (all P < 0.001), whereas JAMA benchmark scores did not ( P = 0.330). All readability indices differed across models (all P < 0.001). Gemini and DeepSeek generally produced responses with lower estimated grade levels than ChatGPT and Doubao. Conclusions The evaluated LLM chatbots generally received favorable safety and accuracy ratings when answering caregiver questions about pediatric burns. However, potentially unsafe responses, limited source transparency, variable empathy, and substantial readability barriers remained. Future pediatric burn information systems should prioritize concise, action-oriented guidance, plain-language communication, transparent evidence attribution, and clear indications for urgent medical assessment. Clinical trial number: Not applicable.