Abstract / Summary
Objective To evaluate the information quality, reliability, readability, and content deficits of ChatGPT, DeepSeek, and Doubao in answering frequently asked questions about bee stings. Methods Twenty-five high-frequency patient questions were selected based on Google Trends, domestic health platforms, and allergy guidelines, and submitted to the three LLMs, yielding 75 responses. Two emergency physicians independently assessed quality using GQS, mDISCERN, QAMAI, EQIP, and PEMAT-U/A. Readability was evaluated using six indices (ARI, FRES, GFI, FKGL, CLI, SMOG). Content deficits were classified qualitatively. Twenty emergency department patients evaluated the comprehensibility and credibility of AI responses. Results Significant differences were found among the three models on all quality metrics (all P < 0.01). DeepSeek achieved the highest mDISCERN score (median 21.00); ChatGPT led on QAMAI (25.00) and EQIP (75.00%); Doubao ranked lowest on all quality indicators but highest on PEMAT-U (73.00%). No model met the sixth-grade reading level standard. Eleven content deficits were identified: outdated medical advice (45.45%), omissions (36.36%), and hallucinations (18.18%). Doubao received the highest patient preference (45.0%). Conclusion The three LLMs exhibited distinct strengths: DeepSeek in source reliability, ChatGPT in overall quality, and Doubao in understandability. All models had excessively high readability levels, and outdated medical advice was the predominant content deficit. Objective To evaluate the information quality, reliability, readability, and content deficits of ChatGPT, DeepSeek, and Doubao in answering frequently asked questions about bee stings. Methods Twenty-five high-frequency patient questions were selected based on Google Trends, domestic health platforms, and allergy guidelines, and submitted to the three LLMs, yielding 75 responses. Two emergency physicians independently assessed quality using GQS, mDISCERN, QAMAI, EQIP, and PEMAT-U/A. Readability was evaluated using six indices (ARI, FRES, GFI, FKGL, CLI, SMOG). Content deficits were classified qualitatively. Twenty emergency department patients evaluated the comprehensibility and credibility of AI responses. Results Significant differences were found among the three models on all quality metrics (all P < 0.01). DeepSeek achieved the highest mDISCERN score (median 21.00); ChatGPT led on QAMAI (25.00) and EQIP (75.00%); Doubao ranked lowest on all quality indicators but highest on PEMAT-U (73.00%). No model met the sixth-grade reading level standard. Eleven content deficits were identified: outdated medical advice (45.45%), omissions (36.36%), and hallucinations (18.18%). Doubao received the highest patient preference (45.0%). Conclusion The three LLMs exhibited distinct strengths: DeepSeek in source reliability, ChatGPT in overall quality, and Doubao in understandability. All models had excessively high readability levels, and outdated medical advice was the predominant content deficit.