Abstract / Summary
Patients with osteoporosis have information needs that routine care inconsistently addresses. In this multicenter real-world evaluation, the LLM-based, guideline-grounded chatbot OsteoCHAT showed high participant acceptability and achieved expert-rated answer quality comparable to general-purpose LLMs. Despite challenges in complex treatment and safety-related topics, guideline-grounded chatbots may represent useful educational adjuncts. Patients with osteoporosis have information needs that routine care inconsistently addresses. Guideline-grounded chatbots may offer scalable educational support We developed OsteoCHAT, a retrieval-augmented conversational system based on the current German osteoporosis guideline, and conducted a multicenter evaluation at seven centers in Germany. Participants interacted with OsteoCHAT, provided real-time feedback on individual responses, and completed a standardized questionnaire assessing usability, comprehensibility, usefulness, trust, perceived quality, and preference over conventional internet search. In parallel, OsteoCHAT was benchmarked against three general-purpose large language models (LLMs) using ten osteoporosis frequently asked questions derived from the BfO patient guideline as the gold standard. Five clinical experts independently rated these responses across five domains, each scored 0–3. Overall, 1,075 question-answer interactions were recorded. Of 389 individually rated responses, 379 (97.4%) received positive feedback. Among 217 complete questionnaire responses, 84.3% rated answers as easily understandable, 79.8% found OsteoCHAT easy to use, 77.2% perceived time savings, and 78.7% considered it a useful addition to existing educational materials. Trust was comparatively lower (64.7% agreement). Blinded expert benchmarking showed acceptable-to-high-quality responses across all LLMs (median total scores 14/15 for Gemini, ChatGPT, and OsteoCHAT and 13/15 for Meta AI), with expert-identified weaknesses mainly concerning pharmacological treatment indication and communication of adverse and rare safety-critical events. OsteoCHAT demonstrated high user acceptability and usability in a real-world multicenter evaluation, was rated as a valuable addition to existing educational materials, and achieved expert-rated response quality comparable to general-purpose LLMs, although limitations in safety communication were identified across all examined systems.