Abstract / Summary
Objective: This study aimed to screen and identify hub genes correlated with septic shock through integrative analyses of bulk transcriptomics, machine learning, weighted gene co-expression network analysis (WGCNA), single-cell transcriptomics, and in silico gene knockout, and further dissect their underlying immunoregulatory mechanisms.
Methods: Multiple bulk transcriptomic datasets of septic shock were integrated and processed via batch correction and differential expression analysis, followed by WGCNA. Candidate genes were screened by intersecting disease-related differentially expressed genes, module genes derived from WGCNA, and predicted drug/compound targets. A diagnostic classifier was established by integrating 113 machine learning algorithms, with its predictive performance assessed using a training cohort and an independent external validation cohort. SHAP analysis was implemented to quantify and interpret the feature importance of model biomarkers. Single-cell transcriptomic data from the dataset GSE175453 were subsequently used to characterize the cell-type localization of hub genes, and scTenifoldKnk was adopted to conduct in silico gene knockout experiments on pivotal genes. Finally, gene set enrichment analysis (GSEA) based on Z-score ranking was performed to dissect the underlying regulatory signaling pathways of these hub genes.
Results: Batch correction markedly optimized the distribution of samples from distinct datasets. Combined differential expression analysis and WGCNA identified prominent transcriptional perturbations associated with immunity and inflammation in septic shock. Multidimensional intersection screening yielded candidate hub genes, from which core genes including SORT1, Turquoise, blue, and brown modules IL18RAP, PLEKHA1, SLC38A1, and LDHA were prioritized. Machine learning evaluations demonstrated that diagnostic models constructed from these candidate genes exhibited strong discriminatory capacity in both training and independent external validation cohorts. Notably, the RF + SVM hybrid model yielded AUC values of 0.998, 0.986 and 0.972 for the training set, GSE137340 and GSE26440, respectively. SHAP analysis verified that SORT1 exerted the largest predictive contribution to the classification model. Single-cell transcriptomic profiling illustrated that SORT1 was predominantly enriched in monocytes, especially the CD16+ monocyte subset, with significantly elevated expression in patients with septic shock. In silico gene knockout of SORT1 led to pronounced transcriptional dysregulation of B2M, CD74 and multiple HLA class II family genes. Gene set enrichment analysis (GSEA) performed on Z-score ranked genome-wide transcripts further uncovered robust enrichment of biological pathways governing antigen processing and presentation, MHC class II antigen presentation, as well as interferon-mediated immune responses.