Abstract / Summary
Abstract Purpose To develop and evaluate the accuracy and safety of a customized, domain-restricted chatbot providing postoperative patient education after vitreoretinal surgery. Methods We evaluated a domain-restricted GPT-4 chatbot accessed through the OpenAI API and implemented with retrieval-augmented generation, a curated vitreoretinal postoperative knowledge base, a persistent system prompt, and rule-based safeguards. Nineteen frequently asked postoperative questions were each submitted in three independent fresh sessions, generating 57 responses. Each response was rated by one of three retinal specialists as accurate and sufficient, partially accurate and sufficient, or inaccurate, and screened for hallucination. Results Across 57 reviewer ratings, 53 (93.0%; 95% CI 83.0–98.1) were rated accurate and sufficient and 4 (7.0%; 1.9–17.0) partially accurate and sufficient; none was inaccurate. Sixteen of 19 questions (84.2%; 60.4–96.6) received unanimous accurate-and-sufficient ratings. Disagreement was confined to three activity-related questions: lifting, exercise, and driving. Hallucinations occurred in 3 of 57 ratings (5.3%; 1.1–14.6), all in activity advice and each flagged by a single reviewer; none arose in red-flag symptom responses. Inter-rater agreement was high (pairwise agreement 89.5%; Gwet’s AC1 0.88). Conclusion A customized, domain-restricted chatbot is a feasible and safe approach to postoperative vitreoretinal patient education, with the only residual inconsistencies arising in areas where clinical practice itself varies between surgeons.