Abstract / Summary
Objective: To identify speech features distinguishing adults with depression from healthy controls, evaluate whether these features track with depression severity during naturalistic treatment, and assess whether baseline speech features predict treatment response. Design: Cross-sectional group comparison (Objective A) followed by longitudinal evaluation in an independent sample (Objectives B and C). Candidate features surviving false discovery rate (FDR) correction in the cross-sectional analysis were carried forward to longitudinal and predictive analyses. Setting: A community mental health center in New York. Participants: Dataset I (cross-sectional): 82 adults with chart-diagnosed depressive disorders and 80 healthy controls (N=162), recruited via convenience sampling. Dataset II (longitudinal): 85 adults with depressive disorders followed for up to four timepoints over a mean of 10.4 weeks during naturalistic treatment; 63 with baseline PHQ-9 >=10 and at least one follow-up were included in prediction analyses. Recruitment occurred under protocols approved between 2020 and 2021. Interventions: None; participants received naturalistic clinical care. Outcome Measures: Cross-sectional: diagnostic group membership (depressive disorder vs. healthy control). Longitudinal: Patient Health Questionnaire-9 (PHQ-9) total score. Predictive: treatment response, defined as >=50% PHQ-9 reduction from baseline to last follow-up. Results: Across three speech tasks (journaling, picture description, phonemic fluency), 11 voice quality features, six shimmer variants and five harmonics-to-noise ratio (HNR) variants, survived FDR correction (all q<.05) in all tasks, with higher shimmer and lower HNR in the depressed group (largest adjusted effect beta=1.89, q<.001). In longitudinal analysis, one phonemic fluency voice quality feature was associated with PHQ-9 at FDR<.05 (maximum HNR beta=-1.43, SE=0.41, FDR=.006), with the remaining ten consistent in direction at FDR<.10; no speech x time interactions were significant. Sex moderated voice quality-depression associations in picture description and phonemic fluency. For treatment response prediction, phonemic fluency voice quality features achieved cross-validated AUC=0.727 [95% CI: 0.662, 0.792], comparable to clinical features alone (AUC=0.700-0.715); the combined model across all tasks and clinical features achieved AUC=0.730 [0.639, 0.820]. No pairwise AUC comparisons between predictor configurations reached significance after correction (all corrected p>.54). Conclusions: Voice quality features, particularly shimmer and HNR, were the most consistent speech correlates of depression across cross-sectional, longitudinal, and predictive analyses and across multiple speech tasks. These features are computationally simple to extract from brief recordings and showed preliminary predictive utility comparable to standard clinical variables. However, the modest sample sizes, chart-based diagnoses and heterogeneous treatment exposures preclude clinical application at this stage. Larger, prospectively designed studies with standardized recording protocols, structured diagnostic assessment, and recruitment at treatment initiation are needed before voice quality features can be recommended as clinical biomarkers.