Abstract / Summary
Abstract Purpose Generative artificial intelligence (AI) models are increasingly evaluated for diagnostic tasks in radiology, yet accuracy, study designs, endpoints, and comparators vary widely. The purpose was to synthesize diagnostic accuracy of generative AI for radiology and compare performance with physicians. Materials and methods A systematic review and meta-analysis was prospectively registered in PROSPERO (CRD420251040000) and conducted in accordance with PRISMA-DTA guidance. Searches of Medline, Scopus, Web of Science, Cochrane Central, and medRxiv (June 2018–March 2025) identified studies validating generative AI on diagnostic tasks in radiology. Two reviewers independently screened, extracted data, and assessed risk of bias with PROBAST+AI, with disagreements resolved by a third reviewer. Multilevel random-effects meta-regression was performed to compare AI performance with physician performance and identify sources of heterogeneity, with study-clustered robust inference to account for multiple estimates per study. Results In total, 48 studies met inclusion criteria. Pooled diagnostic accuracy of generative AI was 42.9% (95% CI, 35.8–50.1%) for free-text tasks and 58.1% (95% CI, 50.0–66.2%) for choice tasks. Generative AI overall showed significantly lower accuracy than expert physicians (difference in accuracy [physicians minus AI], +13.0 percentage points [95% CI, 0.9–25.2]; P = .038). Text-only input was associated with higher accuracy than image-only input (difference in accuracy [image-only minus text-only], −25.9 percentage points [95% CI, −39.1 to −12.8]; P = .001) and text-and-image input (− 10.6 percentage points [95% CI, −20.8 to −0.3]; P = .046). Conclusion Generative AI remained less accurate than expert physicians on diagnostic tasks in radiology. Text-oriented assistive use warrants further evaluation, but the observed accuracy difference may reflect task difficulty and information content rather than the input modality itself. Standardized, transparently reported, adequately powered evaluations are warranted before clinical deployment.