Abstract / Summary
Abstract Background Digital pathology and deep learning have produced rapid advances in automated prostate cancer detection and Gleason grading. However, headline performance from internal benchmark datasets does not necessarily translate to reliable clinical performance across institutions, scanners, staining protocols, specimen types, or patient populations. This review synthesizes evidence on automated Gleason grading and prostate cancer classification from histopathology images, with particular emphasis on external validation, generalizability, failure modes, human–AI interaction, and deployment in low-resource settings. Methods The review was structured according to PRISMA 2020 principles. PubMed-indexed studies, publisher records, systematic reviews, and methodological guidance were examined, with emphasis on original studies of digital histopathology AI for prostate cancer and Gleason grading. Fifty key publications were selected for the evidence synthesis, supplemented conceptually by broader literature identified in prior systematic reviews. Because the evidence is heterogeneous in specimen type, endpoint, annotation strategy, architecture, and validation design, a narrative synthesis was used rather than an unverified pooled estimate. Results The evidence demonstrates that deep learning can approach or exceed the agreement of general pathologists and can achieve expert-level performance in selected external cohorts. The PANDA challenge demonstrated strong cross-continental performance using 10,616 digitized prostate biopsies, while subsequent real-world evaluations showed wider performance variation among algorithms. External validation on prostatectomy specimens has also demonstrated strong agreement for a biopsy-trained commercial model. Prospective multicenter evidence published in 2026 indicates that AI assistance can alter Gleason grading in routine practice and increase diagnostic confidence. Persistent weaknesses include domain shift, tissue and scanning artifacts, class imbalance, difficulty distinguishing adjacent Gleason patterns, limited prospective validation, dataset leakage, opaque reference standards, and underrepresentation of low-resource settings. Conclusion AI-assisted prostate histopathology has moved beyond proof-of-concept research, but clinical readiness should not be inferred from accuracy on a single dataset. Future studies should prioritize independent multicenter validation, prespecified clinical endpoints, transparent reference standards, calibration and uncertainty reporting, robustness testing, prospective workflow studies, and implementation strategies compatible with low-resource pathology laboratories.