Abstract / Summary
Abstract Background Liver MRI is central to non-invasive hepatocellular carcinoma (HCC) diagnosis. Deep learning (DL) is increasingly studied for detection and interpretation, but the maturity, transportability, and patient-level diagnostic evidence for these systems remain uncertain. Main body We searched MEDLINE, Embase, and Web of Science from 2010 through July 2025 (the exact original execution day was not retained) and updated the same strategies through 20 July 2026. Eligible studies evaluated stand-alone DL or radiologist-interpreted DL-assist MRI and provided reconstructable patient-level 2 × 2 data. Two reviewers independently selected studies, extracted data, and assessed QUADAS-2 and AI-reporting items. Four retrospective Chinese tertiary-centre cohorts ( N = 698; three stand-alone systems and one radiologist-interpreted DL-assist system) were included; the update identified no additional eligible study. Model-derived exploratory pooled estimates were sensitivity 0.862 (95% CI 0.819–0.896) and specificity 0.916 (0.881–0.941); MCMC median LR + was 10.19 (95% interval 7.22–14.55), LR- was 0.151 (0.114–0.199), and diagnostic odds ratio was 67.59 (41.10–111.10). Between-study variance estimates lay at the boundary and the confidence and prediction regions were nearly identical; with four studies, this does not establish statistical or clinical homogeneity. Patient selection and flow/timing commonly had high or unclear risk of bias, and validation, threshold, blinding, explainability, workflow, and regulatory reporting were incomplete. Conclusions DL applied to liver MRI yielded promising model-derived patient-level summaries, but four retrospective, geographically concentrated cohorts do not establish a transportable accuracy benchmark, radiologist augmentation, or clinical benefit. The principal finding is an evidence-maturity gap requiring prospective multi-centre reader studies, locked testing, transparent reporting, and evaluation of workflow and safety.