Abstract / Summary
Large language models (LLMs) are increasingly being introduced into clinical environments, but their influence on physician decision-making, particularly when responses are incorrect, remains insufficiently understood. The aim of the study was to assess whether pediatricians are influenced by LLM-generated answers in their clinical decision-making and whether they can distinguish correct from incorrect AI-generated responses. Ninety-eight pediatrician participants were recruited through convenience sampling from university hospitals and private clinics in Greece between February and July 2025. Mixed-effects logistic regression models were used to compare differences in participant behavior when the LLM provided correct versus incorrect answers. The main outcomes were the ability of pediatricians to distinguish correct from incorrect LLM-generated answers and the extent to which AI-generated responses influenced changes in their clinical decisions. Subgroup analyses evaluated differences by training level and practice setting. Among 98 pediatricians, most were able to distinguish between correct and incorrect LLM-generated answers using their clinical experience and training ( p < 0.001). Overall, most participants were able to use LLM output constructively without being misled by incorrect responses. Most pediatricians retained the ability to recognize incorrect LLM-generated information and were able to use LLMs as supportive diagnostic tools. However, pediatric residents ( p = 0.02) and private-practice pediatricians ( p = 0.005) were more likely to change their answers based on AI-generated responses, even when the LLM-provided answer was incorrect. The greater susceptibility of residents and private-practice pediatricians to incorrect LLM output suggests a risk of authority or priming effects. These findings support the need for targeted training and implementation safeguards before widespread clinical deployment of LLMs in pediatric practice.