Abstract / Summary
ABSTRACT Introduction Large language models (LLMs) can enhance learning engagement and accessibility, though their performance in dentistry is inconsistent. Bitewing radiographs are critical for caries detection but prone to diagnostic errors from operator positioning. This study compared the performance of four LLMs with reasoning capabilities, GPT o1, GPT o3‐mini, Gemini 2.0 Flash and Grok 3, in generating scenario‐specific educational feedback from a structured textual format of bitewing radiographs. Materials and Methods Ten positioning error scenarios were created by DMFR specialists. LLMs were evaluated using an original prompt (OP) and an engineered prompt (EP), resulting in a 4 × 2 LLM‐by‐prompt design. Each LLM was tested via its model‐specific access mode: API or web interface. Eighty outputs were assessed by two DMFR specialists for detail, content, clarity and relevance using a 5‐point Likert scale. Ordinal category‐level scores were analysed using a cumulative link mixed‐effects model, with LLM, prompt type, scenario group, assessment category and LLM × prompt as fixed effects, and scenario as a random intercept. Results The LLM–prompt interaction was significant, indicating prompt engineering affected models differently. GPT o3‐mini had significantly lower odds of higher scores than Gemini 2.0 Flash, while EP reduced Grok 3 performance. Positioning error type influenced performance, with significant model‐dependent differences across scenario groups. No significant LLM–assessment category interaction was observed, suggesting model differences did not vary across assessment categories. Conclusion LLM‐generated feedback varied by model, prompt design and positioning error type, but not consistently across assessment categories. Outputs require validation, model‐specific prompt testing and specialist review before implementation.