Abstract / Summary
Abstract We report a real (executed, not projected) model-size scaling comparison for a psychiatric Multi-Agent Debate (MAD) system, run entirely on local open-weight models, with an ablation isolating debate from Judge synthesis and repeated independent trials to establish statistical reliability. On 24 synthetic DSM5-TR-style vignettes, zero-shot accuracy rises with scale (75.0% at 1B → 83.3% at 3B), and PsychMAD scores numerically lower than zero-shot at both scales. A single-run paired comparison is not significant at N = 24 (p = 0.18 and p = 0.63), but this is a power problem, not an absence of effect: across 5 independent trials per condition, PsychMAD’s mean accuracy is significantly lower than zero-shot’s at both 1B (73.3% vs. 57.0%, exact permutation p = 0.024) and 3B (88.3% vs. 69.2%, p = 0.008). An ablation isolating debate from Judge synthesis complicates the natural follow-up hypothesis that Judge synthesis is responsible: debate alone (specialists + vote, no Judge) already underperforms zero-shot, and the Judge’s marginal effect on category-level accuracy is inconsistent across scales (it further lowers accuracy at 1B but partially recovers it at 3B). What the ablation establishes cleanly is narrower: on the two ground-truth high-risk cases, debate itself (not the Judge) recovers a risk signal zero-shot misses at 1B, while the Judge specifically (not debate) dilutes a correct risk signal at 3B — two different mechanisms, at two different scales, on two different outcome axes (category accuracy vs. risk-tier correctness), neither large-N-confirmed given only two risk-positive cases. As secondary, exploratory context, several of our conditions are numerically within or above the range of published human-psychiatrist vignette accuracy (66%, 1,038 psychiatrists across 19 countries, Boberg et al., 2026; and a small literature of matched same-instrument LLM-vs-clinician studies that is itself mixed-to-favorable for LLMs, Lenz et al., 2026; Kim et al., 2024b; Patel et al., 2025; Chen et al., 2026; Shan et al., 2025). We did not run a matched study ourselves and our vignettes are author-labeled; we present this context to avoid both an overclaim (“LLMs beat psychiatrists”) and an uninformed dismissal (assuming human accuracy is a fixed high ceiling it is not).