Abstract / Summary
Abstract Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify the optimal approach. Twenty referral scenarios spanning the urgency spectrum were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded to produce a consensus reference standard; inter-rater reliability was quantified. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per referral). The experiment was run with a simple prompt and repeated with an advanced prompt supplying explicit triage expectations and worked examples. In line with published literature, agreement among the four rheumatologists was moderate (Fleiss kappa 0.60; mean pairwise linear-weighted kappa 0.79), with all four assigning categories within one level of each other in 90% of cases. All 2760 model calls returned an interpretable result. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho = 0.42; P =.047) and accuracy tracked cost. Advanced prompting reduced between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P =.01), abolished the size-accuracy association (rho=-0.05; P =.83) and removed the accuracy-cost relationship. The best models matched expert consensus on 18 of 20 cases, comparable to the reproducibility observed among the rheumatologists themselves. Under-triage errors persisted with some LLMs. Contemporary LLMs categorised rheumatology referral urgency with an accuracy and reproducibility comparable to that reported for human triage systems, although human triage was not tested in this study. Advanced prompting appeared to substitute for the reasoning capability of larger models, suggesting that adequate performance on this task may not require the most expensive models. These findings indicate that automation of this administrative task is technically feasible. Candidate models that warrant prospective evaluation are identified.