Abstract / Summary
Background: Emergency triage of anterior circulation large vessel occlusion (LVO) is constrained by delays in vascular imaging, specialist interpretation, and transfer decision-making. Noncontrast computed tomography (NCCT) is often obtained first in suspected stroke, but visual recognition of LVO in NCCT images is difficult outside specialist settings. NCCT-based AI may provide an early human-in-the-loop escalation signal.
Objective: This study aimed to compare validated NCCT-based AI paradigms with human reader paradigms for detection of anterior circulation LVO and assess whether current evidence supports prospective evaluation of AI-assisted escalation pathways.
Methods: We searched the PubMed, Embase, Web of Science Core Collection, and Cochrane Library databases from inception to May 24, 2026. Eligible studies evaluated NCCT-based LVO detection in validation datasets independent of model development, included within-study head-to-head comparisons, and provided reconstructible 2 × 2 data. Four nodes were compared: expert readers, nonexpert readers, unimodal imaging AI, and clinically informed multimodal AI. A Bayesian bivariate hierarchical diagnostic test accuracy network meta-analysis estimated sensitivity, specificity, diagnostic odds ratio, and absolute differences. Risk of bias and applicability were assessed using the Prediction Model Risk of Bias Assessment Tool for AI and the complementary Quality Assessment of Diagnostic Accuracy Studies-3, and certainty was rated using the Grading of Recommendations Assessment, Development, and Evaluation (GRADE).
Results: All 10 included studies were retrospective validation studies, comprising 11 validation datasets, 28 study node arms, and 3632 patients. Unimodal imaging AI had a sensitivity of 0.78 (95% credible interval 0.68-0.86), specificity of 0.88 (95% credible interval 0.81-0.94), and diagnostic odds ratio of 30.89 (95% credible interval 14.19-60.27). Expert readers and nonexpert readers had lower sensitivity estimates of 0.62 (95% credible interval 0.48-0.75) and 0.60 (95% credible interval 0.45-0.75), respectively, with similar specificity estimates of 0.86. In league table comparisons, unimodal imaging AI showed higher sensitivity than expert and nonexpert readers by 0.16 (95% credible interval 0.03-0.29) and 0.18 (95% credible interval 0.03-0.32), respectively, with no clear specificity separation. Clinically informed multimodal AI had a sensitivity of 0.81 (95% credible interval 0.64-0.92) and specificity of 0.92 (95% credible interval 0.80-0.98), but this sparse node was connected to human readers only through indirect evidence.