Abstract / Summary
Background: Prognostic models for metastatic spinal tumors have evolved from conventional point-based scores to models incorporating treatment, laboratory, and machine-learning variables. However, comparisons in contemporary Japanese surgical cohorts remain limited. Methods: This retrospective study included 171 of 208 consecutive patients who underwent surgery for metastatic spinal tumors between January 2017 and August 2024 and had ascertainable 1-year survival status. Seven prognostic systems were compared using receiver operating characteristic area under the curve (AUC) analysis for 90-day and 1-year survival. Pairwise AUC comparisons were adjusted using the Holm method. Calibration and Brier scores were also assessed for the Skeletal Oncology Research Group (SORG) machine learning algorithm (MLA). Results: For 90-day survival, AUCs were 0.77 (95% confidence interval [CI], 0.58–0.95) for the SORG MLA and 0.76 (95% CI, 0.60–0.91) for the New Katagiri score. For 1-year survival, AUCs were 0.79 (95% CI, 0.72–0.86) for the New Katagiri score, 0.76 (95% CI, 0.68–0.83) for the SORG classic scoring algorithm, and 0.74 (95% CI, 0.66–0.82) for the SORG MLA. After Holm adjustment, no significant pairwise differences were found among these three models at either prediction horizon. SORG MLA calibration deviated from perfect calibration. Conclusions: In this single-center surgical cohort, the New Katagiri score and SORG models showed higher discrimination than several conventional scoring systems. However, these findings do not establish model equivalence or the clinical benefit of model-guided decisions. Broader external validation and assessment of clinical utility are needed.