Abstract / Summary
Aim: In large language model (LLM) examination studies, models answer the same items repeatedly, so observations are clustered and conventional intervals overstate precision. We aimed to measure the accuracy of seven LLMs on a subspecialty-level anaesthesiology question bank and to determine which between-model and subgroup differences remain distinguishable under a cluster-respecting analysis.Materials and Methods: A 120-item, five-option multiple-choice bank in the style of the Anaesthesiology and Reanimation Subspecialty Examination, written by the authors and reviewed against current guidelines and textbooks, was administered to seven LLMs in five runs at temperature 0, with option order re-randomised per run (4,200 responses from 120 unique items). Accuracy is reported with cluster bootstrap 95% confidence intervals (CIs) resampling items (4,000 replicates) and naive Wilson intervals for comparison; subgroup analyses are exploratory.Results: Pooled accuracy was 80.1% (cluster-robust 95% CI 75.8-83.9; naive Wilson 78.8-81.3). Between-model differences were large and robust, spanning 60.8% (54.0-67.5) to 94.0% (90.3-97.0) with non-overlapping extremes. Cluster-robust intervals were wider in 25 of 26 estimates (median factor 2.1, range 0.8-4.0). Most subgroup comparisons were not distinguishable, intervals overlapping substantially: vignette (75.4%, 62.5-86.5) versus non-vignette (80.9%, 76.6-84.9) and guideline-dependent (73.0%, 64.6-80.7) versus other items (80.9%, 76.6-84.9). Only contrasts between weakest and strongest domains persisted. Between-run standard deviation was 1.26-2.80 points.Conclusion: Between-model differences and run-to-run instability are robust; commonly emphasised domain and item-characteristic differences are mostly not distinguishable once clustering is respected. Benchmarks of 120 items can separate models whose accuracies differ widely but cannot reliably localise their weaknesses.