Abstract / Summary
Large language models (LLMs) are being explored for medical risk assessment, yet performance on standardized benchmarks does not necessarily establish clinical competency. The present study evaluated whether LLM-generated Alcohol Use Disorder (AUD) risk judgments aligned with epidemiological evidence and remained consistent as clinical and demographic context changed. Three LLMs—GPT-5.5, LLaMA-2-70B, and MEDITRON-70B—completed four pairwise-comparison experiments involving DSM-5 AUD criteria, demographic characteristics, criteria paired with demographic characteristics, and narrative clinical vignettes. LLM-implied rankings were compared with epidemiological benchmarks derived from the National Epidemiological Survey on Alcohol and Related Conditions, Wave Three (NESARC-III). Our results show that none of the LLMs reproduced the implied severity ordering of the AUD criteria. Demographic cues consistently influenced AUD risk judgments, although the direction and magnitude of these effects varied across LLMs, demographic attributes, and clinical contexts. Demographic rankings also showed only partial stability as contextual information changed. LLMs differed substantially in decisiveness, with GPT-5.5 assigning substantially higher probability to selected responses than LLaMA-2-70B or MEDITRON-70B. These findings demonstrate a gap in current LLM risk-assessment capabilities and highlight the importance of evaluating the direction, magnitude, stability, and certainty of LLM judgments alongside accuracy.