Abstract / Summary
Predicting thyroid cancer recurrence from clinical and multi-omics data remains a challenging information-processing task because available datasets are often small, class-imbalanced, heterogeneous, and highly redundant. To address these issues, we propose CARE, a Causal-Augmented Representation and Evidence framework designed as a unified three-stage information-processing pipeline. First, CARE learns a stochastic-gated variational representation and imposes a sparse acyclicity-constrained directed dependency structure on selected latent variables, enabling minority-class augmentation in structured latent space rather than unconstrained interpolation in the observed feature space. Second, a Causal Specificity Index (CSI) identifies the most specific latent representation, which is integrated with observed variables and purified through ElasticNet regularization to derive a compact predictive feature subset. Third, downstream classifiers are trained on the CARE-purified information space, and SHapley Additive exPlanations (SHAP) are used for post hoc interpretation. CARE was evaluated on a 383-patient clinical cohort and a 489-patient TCGA-THCA clinical-plus-RNA-seq cohort. On internal held-out test data, the CARE-AdaBoost configuration achieved a macro-averaged F1 score of 0.9843 and an AUC of 0.9996 in the clinical cohort, and an AUC of 0.7587 in the multi-omics cohort. Ablation analyses supported the complementary roles of the DAG constraint, latent-space augmentation, and specificity-driven feature purification. The CSI-selected latent variable Z1 showed interpretable associations with clinical risk, treatment response, and recurrence status in the clinical cohort. These findings suggest that a unified structure-aware workflow can improve robustness, compactness, and interpretability in small-sample imbalanced biomedical prediction.