Abstract / Summary
Objective: To evaluate whether a context-specific large language model, OpenEvidence, can generate gastroenterology/hepatology eConsult responses for metabolic dysfunction-associated steatotic liver disease (MASLD), while characterizing the frequency and potential severity of errors. Patients and Methods: In this retrospective cohort study at a tertiary academic medical center, 125 specialist electronic consultations (eConsults) for MASLD (from January 2020 to December 2024) were paired with OpenEvidence-generated eConsults produced using OpenEvidence and standardized prompts. Five blinded physician reviewers evaluated all eConsults using a previously validated instrument. A board-certified gastroenterologist independently reviewed OpenEvidence-generated eConsults to identify errors and assess potential for harm. Results: OpenEvidence-generated eConsults received higher overall eConsult Specialist Quality of Response scores than specialist eConsults (mean difference, 6.4; 95% CI, 6.0-6.9; P <.001) and higher global quality ratings (odds ratio [OR], 16.1; 95% CI, 12.0-21.5; P <.001). OpenEvidence-generated eConsults were more likely to be rated as providing an appropriate amount of information (70.9% vs 33.9%; OR, 6.4; 95% CI, 4.8-8.5; P <.001). Of 125 OpenEvidence-generated eConsults, 12 contained minor errors and 1 contained a potentially major error. Conclusion: In this blind comparison of paired eConsults for MASLD, OpenEvidence-generated responses were rated higher in perceived quality and informational adequacy than specialist-authored responses, but occasionally clinically relevant errors occurred. These findings support further evaluation of OpenEvidence-assisted eConsult workflows for protocol-driven conditions.