Abstract / Summary
Abstract Background Refractory Mycoplasma pneumoniae pneumonia (RMPP) has a protracted course and frequent complications, and its diagnosis is often delayed; early identification of children at high risk is therefore crucial. This study aimed to develop and internally validate a combined machine-learning model integrating CT radiomics features and early clinical indicators to predict progression from Mycoplasma pneumoniae pneumonia (MPP) to RMPP in children. Methods This single-center retrospective study enrolled 527 children with MPP (204 RMPP, 323 non-refractory MPP [NRMPP]) between January 2022 and December 2024, randomly split into training (n = 368) and internal test (n = 159) cohorts at a ratio of approximately 7:3. From 1834 CT radiomics features, 25 selected by stability (intraclass correlation coefficient ≥ 0.85), Mann–Whitney U screening, correlation-based redundancy removal and LASSO dimensionality reduction formed a radiomics score (Rad-score); seven machine-learning classifiers were compared. A clinical model was built by multivariate logistic regression with backward stepwise selection of variables with P < 0.1 on univariate analysis, and a combined model integrating the clinical prediction and Rad-score was visualized as a nomogram. Performance was evaluated using the area under the receiver operating characteristic curve (AUC), calibration curves, the Hosmer–Lemeshow test and decision curve analysis (DCA). Results The combined model achieved AUCs of 0.876 (95% CI: 0.839–0.913) and 0.869 (95% CI: 0.815–0.923) in the training and test cohorts, respectively, significantly outperforming the clinical model (0.704 and 0.638; P = 1.42 × 10⁻⁸ and P = 1.59 × 10⁻⁶); in the test cohort, its AUC was also higher than that of the radiomics model (0.851), although not significantly (P = 0.060). The final clinical model included LDH, peak body temperature and ALT. DCA demonstrated a higher net benefit of the combined model across a wide range of threshold probabilities. The Hosmer–Lemeshow test indicated some degree of miscalibration of the combined model in the test cohort (P = 0.032). Conclusions The combined machine-learning model showed good discrimination for predicting progression from MPP to RMPP (internal test AUC 0.869) and significantly outperformed the clinical model; however, its calibration was suboptimal in the test cohort (Hosmer–Lemeshow P = 0.032), and external validation and recalibration are required before clinical deployment.