Abstract / Summary
Abstract Background Large language models (LLMs) are increasingly used to obtain information on knee osteoarthritis (OA) and total knee arthroplasty (TKA), but evidence-based answer quality and consistency across repeated outputs require separate evaluation. Methods We included 24 standardized questions derived from American Academy of Orthopaedic Surgeons (AAOS) guidelines: eight on non-surgical knee OA and 16 on TKA surgical management. Four LLMs were tested in three independent sessions per question, yielding 288 answers anonymously rated by eight experts. Relevance, accuracy, understandability, and safety were scored from 1 to 5; their equally weighted mean was the primary outcome. Three-response consistency was rated separately. Quality outcomes were analyzed with maximum-likelihood mixed-effects models, whereas consistency used paired Friedman and Wilcoxon tests. Pairwise comparisons were Holm-adjusted, with paired t-test sensitivity analyses. Inter-rater reliability was evaluated using absolute-agreement intraclass correlation coefficients (ICCs). Results All ratings were complete. The overall model effect for overall quality was significant (Wald χ²=78.30, df = 3, P < 0.001). Adjusted means were 4.527 for GPT, 4.311 for DeepSeek, 3.866 for ERNIE, and 3.675 for Kimi. GPT exceeded DeepSeek by 0.216 points (95% CI 0.040–0.392; Holm-adjusted P = 0.032 in the mixed model), whereas paired t-test and Wilcoxon sensitivity analyses yielded P = 0.055 and P = 0.069, respectively. GPT scored highest for relevance, accuracy, and safety, whereas DeepSeek scored highest for understandability. Consistency means were 4.750, 4.490, 4.198, and 4.047, respectively (Friedman χ²=45.471, P < 0.001). The eight-rater average-measure ICCs were 0.966 for overall quality and 0.852 for consistency. The model-by-question-category interaction was not significant (P = 0.805). Conclusions The four LLMs differed in answer quality and repeated-output consistency in this benchmark, with advantages varying by evaluation domain. The overall-quality difference between GPT and DeepSeek was sensitive to the inferential method and should not be interpreted as a fixed absolute ranking. These findings reflect expert-rated textual performance and do not establish benefits for patient understanding or clinical outcomes.