Abstract / Summary
Abstract Background Generative artificial intelligence (AI) is increasingly used in medical education, but most current tools remain reactive: learners must initiate questions before receiving assistance. Whether an AI tutor that proactively delivers contextualized, individualized learning prompts improves immediate learning and delayed retention remains uncertain. Thyroid cancer guidelines, which require frequent updating and multidisciplinary integration, provide a suitable test case. Methods We conducted a nationwide online, parallel-group, assessor-blinded, active-controlled randomized trial from 1 March to 1 May 2026. Resident physicians in standardized residency training were recruited online and randomized 1:1 to a proactive push AI-assisted learning group or a reactive response chatbot control group. Both groups used the same Feishu-OpenClaw platform, knowledge base, question bank, prompts, and ChatGPT-5.4 model. The proactive group received two individualized micro-quizzes daily with immediate feedback; the control group accessed the same content only when initiating interaction. The primary outcome was the 40-item total knowledge score at Week 3 (T3w). Key secondary outcomes included delayed retention at Week 7 (T7w), module scores, engagement, guideline-application self-efficacy, and user experience. Results A total of 120 residents were randomized; 105 (87.5%) completed the T3w primary assessment, and 90 (75.0%) completed all scheduled assessments through T7w. In intention-to-treat analyses, the adjusted T3w total score favored the proactive group but did not reach conventional statistical significance (adjusted difference, 0.65 points; 95% CI, -0.04 to 1.34; p = 0.065). The proactive group scored higher on the T3w adverse drug reaction module (0.89 points; 95% CI, 0.13 to 1.65; p = 0.022) and T7w delayed retention (1.49 points; 95% CI, 0.36 to 2.62; p = 0.009). It also showed more login days, more micro-quiz completions, more learner-initiated questions, and greater self-efficacy gains (2.83 points; 95% CI, 1.51 to 4.16; p < 0.001). Conclusions When content, knowledge base, and model were held constant, proactive push AI tutoring did not clearly improve the prespecified immediate total score, but it produced more consistent advantages in delayed retention, adverse drug reaction knowledge, engagement, self-efficacy, and perceived learning experience. AI proactivity may act primarily by increasing distributed retrieval practice rather than by immediately increasing short-term total scores. Trial registration: Chinese Clinical Trial Registry, ChiCTR2600121101. Retrospectively registered on 25 March 2026.