Abstract / Summary
Background Speech carries cues to variation in mental state in schizophrenia, typically indexed with clinician-rated scales such as Positive and Negative Syndrome Scale (PANSS). In this study, we aggregate multi-centre recordings to assemble a large corpus and assess symptom-prediction models at scale, to enable rater-independent, efficient, and reproducible assessments for early detection of relapse-related signals from speech. Methods In this computational secondary analysis, we compiled data from 453 patients with schizophrenia spectrum disorders, recruited between 2005 and 2026 from ten global sites in the Czech Republic, United States, Spain, Chile, France, Switzerland, the Netherlands, Turkiye, Germany, and China, and clipped their speech recordings into 6,664 segments. Across three featuresets, acoustic-prosodic profile, pretrained multilingual embeddings, and their concatenation, we compared 16 algorithms in separate item-specific models to predict eight relapse-related PANSS items, including three positive (P1, P2, P3), three negative (N1, N4, N6), and two general (G5, G9) items, on speaker-disjoint splits. Performance was assessed by root-mean-squared-error (RMSE) at both segment and participant (median-aggregation) levels. Best model per item underwent bias checks for age, sex, education, and symptom severity. Findings Best-performing models predicted symptoms with prediction errors of 1.5 or lower: P1 1.544/1.527, P2 1.318/1.107, P3 1.487/1.542, N1 1.071/1.030, N4 1.521/1.430, N6 0.869/0.855, G5 0.951/0.882, G9 1.298/1.282 (segment/participant). Performance of pretrained embeddings surpassed acoustic-prosodic features and concatenation. Results were comparable in languages or language variants with fewer available speech and language computational resources. We found no bias by age, sex, or education, aside from reduced N4 accuracy in males; but performance degraded with higher symptom severity. Interpretation Speech can support automatic assessment of schizophrenia symptoms using pretrained embeddings, even without transcripts. Such models show promise as clinically meaningful, efficient, and low-burden tools for real-time monitoring of symptom trajectories. Future work should validate these models prospectively in higher-severity samples and improve clinical interpretability. Funding European Union Horizon.