Abstract / Summary
Purpose: Videourodynamics (VUDS) are used to assess bladder dysfunction and risk of upper urinary tract deterioration in patients with spina bifida (SB), but interpretation is subject to interrater variability. Machine learning models have previously classified bladder dysfunction severity and predicted incident hydronephrosis using VUDS data. This study evaluated the performance of adapted versions of these models in an independent external cohort. Materials and Methods: VUDS data were collected from SB patients who underwent VUDS at a single institution between 2016 and 2025. Previously developed machine learning frameworks were adapted for implementation in the independent dataset. Three models were utilized: a deep learning convolutional neural network model using pressure-volume data, a deep learning imaging model using fluoroscopic imaging data, and an ensemble model that averaged the risk data from the pressure-volume and fluoroscopic data. The models were evaluated for prediction of incident hydronephrosis and classification of bladder dysfunction severity compared to expert pediatric urologist reviewers. Results: 70 patients were included in the hydronephrosis cohort, of whom 13 (18%) developed incident hydronephrosis. The ensemble model outperformed individual modalities, with a concordance index of 0.77 (pressure-volume only: 0.70, imaging only: 0.72) and an overall AUROC of 0.75 (pressure-volume: 0.67, imaging: 0.74). For bladder dysfunction classification in 95 VUDS studies, the ensemble model achieved 71% accuracy compared to 60% for pressure-volume data alone and 66% for imaging alone, with no substantial disagreements with expert reviewers. Conclusions: The machine learning models that were developed to predict incident hydronephrosis and classify the severity of bladder dysfunction using VUDS data were able to be adapted to an external, independent dataset and perform with similar discrimination and accuracy as previously reported.