Abstract / Summary
Background: Evaluations of large language models (LLMs) in medicine rest largely on diagnostic accuracy. We report what we learned evaluating four frontier LLMs on psychiatric diagnosis, where we asked whether clinician assessment of model reasoning adds information that accuracy does not. Methods: Four LLMs produced ranked five-diagnosis differentials for 196 psychiatric vignettes (135 screened published case reports; 61 unpublished, clinician-authored). An automated adjudicator scored accuracy. On a random 30-vignette subset, five psychiatrists blinded to model identity judged correctness and rated reasoning on a 0-4 rubric adapted from the ACGME Psychiatry Milestones. We validated the adjudicator against the clinicians, measured the reproducibility of accuracy across generation runs and repeated adjudication, and tested whether computable reasoning trace features could substitute for clinician ratings. Results: On all 196 vignettes, automated top-5 accuracy ranged from 0.759 to 0.837, and the majority verdict of three adjudication passes separated one of six model pairs. On the 30 rated vignettes, clinician-determined accuracy separated no pair, while reasoning separated five. Two systems were equally accurate (0.827 each; difference bounded within +/-0.059, or +/-0.10 after Holm adjustment) yet differed by 0.70 reasoning points (95% CI, 0.50 to 0.90). The adjudicator agreed with held-out clinician majorities less closely than individual clinicians did (0.83 to 0.85 vs 0.89 to 0.96), regeneration changed top-1 correctness for 18 of 119 pairs, and reference strings that combined several diagnoses were scored as misses until split. Trace features that tracked ratings across model averages did not track them across individual traces. Conclusions: Accuracy did not distinguish models that clinicians rated differently, and automated accuracy depended on how it was scored. Where LLMs are meant to support clinical reasoning, structured clinician assessment of reasoning belongs in pre-deployment evaluation.