Abstract / Summary
Abstract Trained on vast text corpora, large language models (LLMs) encode statistical expectations about language. We tested whether deviations from these expectations, quantified as LLM surprisal, capture clinically relevant speech alterations. Surprisal was defined as median sentence-level cross-entropy of Qwen2.5-32B-Instruct conditioned on a healthy-speaker persona. In this cross-sectional German sample with schizophrenia-spectrum (SSD, n = 48), bipolar (BD, n = 26) or depressive disorders (MDD, n = 119) and healthy controls (n = 178), surprisal was elevated in SSD (β = 0.49**) and BD (β = 0.39*), but not MDD, after adjustment for demographic confounders and average sentence length. Surprisal correlated with eight of 18 symptom scales, including positive formal thought disorder (ρ = 0.23**), largely reflecting between-group rather than within-group differences. It improved model fit beyond sentence length, confounders and semantic coherence, although discrimination remained modest (AUC 0.74 SSD, 0.62 BD), and was robust across prompts and models. LLM surprisal may provide a scalable marker of language alterations across psychiatric disorders, requiring further validation.