Abstract / Summary
Abstract Background Whether fixed pretreatment MRI signal summaries improve pathological complete response probability prediction beyond age and receptor status was examined across two trial cohorts. External probability accuracy and the stability of individual predictions were assessed to inform the requirements for treatment-response communication. Methods Models were developed retrospectively using public, deidentified I-SPY2 data and evaluated externally using I-SPY1 data. Twelve baseline dynamic contrast-enhanced MRI summaries were combined with age, hormone receptor status and HER2 status. Clinical, MRI and fusion models were fitted by penalized logistic regression with nested five-fold cross-validation. All pipelines were frozen before external prediction. The paired external Brier score difference, clinical minus fusion, was estimated with a 95% interval from 2000 patient bootstrap resamples conditional on the fitted models. Discrimination, calibration and conventional MRI comparisons were assessed. Development and external patients were independently resampled in a further 1000 replicates, with preprocessing, tuning and fitting repeated. Results Models were developed in 982 patients with 316 complete responses. External evaluation was performed in 166 patients with 48 responses after five missing original outcomes and one prespecified phase-annotation discrepancy were excluded. Internal Brier scores of 0.20143 and 0.19975 were obtained for clinical and fusion models, respectively. Corresponding external scores of 0.17543 and 0.18454 were obtained, with a difference of − 0.00911 (95% interval − 0.02087 to 0.00218). External areas under the receiver operating characteristic curve were 0.773 for clinical, 0.528 for MRI and 0.707 for fusion models. Mean predicted response probabilities exceeded the observed 28.9% frequency. With repeated development, the median external Brier difference was − 0.01598 (95% interval − 0.07367 to 0.00646). Mean absolute prediction changes were 3.1, 8.4 and 7.3 percentage points for clinical, MRI and fusion models. Conclusions Improved external response probability prediction was not established after these fixed pretreatment MRI summaries were added. The MRI increment should be assessed against clinical and conventional imaging comparators, with calibration and uncertainty reported before treatment-response communication is considered. Neither equivalence nor clinical utility was established.