Abstract / Summary
Depression is a prevalent mental health disorder that requires timely and accurate identification to support early intervention and improve patient outcomes. Although recent advances in deep learning (DL) have improved automated depression detection, many existing approaches rely on a single modeling architecture, such as Transformers or Recurrent Neural Networks (RNNs), limiting their ability to capture complementary temporal and multimodal information. This study proposes MHTCI-Net, a Multi-branch Hybrid Temporal Cross-Interaction (HTCI) framework for multimodal depression detection. The proposed architecture jointly learns temporal dependencies and cross-modal interactions from heterogeneous data sources, enabling richer representations than conventional single-model approaches. To further enhance temporal feature learning, the framework incorporates Adaptive Temporal Representation Refinement (ATRR) and a Temporal Attention Pooling (TAP) mechanism to improve feature discrimination and robustness. The proposed model was evaluated on the Large-scale Multimodal Vlog Dataset (LMVD) using a subject-independent five-fold cross-validation protocol and Compared with recent multimodal depression detection methods. Comprehensive experiments, including ablation studies, threshold optimization, and model training analyses, demonstrate that MHTCI-Net effectively learns complementary multimodal temporal representations and achieves competitive performance on the LMVD benchmark, with particularly favorable Accuracy and AUC among the compared methods. The proposed framework provides a research-oriented approach for multimodal depression-risk classification on the LMVD benchmark. External validation on independent cohorts and prospective clinical studies are required before considering its use for clinical screening, diagnosis, or decision-support applications.