Abstract / Summary
Diagnostic performance varies across glomerular disease (GD) subtypes, and whether artificial intelligence (AI) outperforms pathologists with different experience levels remains uncertain.
To evaluate pathology-based AI models for GD classification and compare their performance with pathologists.
PubMed, Embase, Web of Science, and Cochrane Library were searched through 15 July 2026. Studies using pathology images and pathology diagnosis as the reference standard were included. Random-effects models pooled sensitivity, precision, accuracy, F1 score and area under the curve (AUC).
Fifteen studies comprising 39,536 validation sample units, not necessarily unique patients, were included. For subtypes with at least 10 validation datasets, AI achieved high performance for membranous nephropathy (MN; sensitivity 0.96, precision 0.94, accuracy 0.96, F1 score 0.95, AUC 0.98), IgA nephropathy (IgAN; sensitivity 0.92, precision 0.91, accuracy 0.94, F1 score 0.90, AUC 0.96), and minimal change disease (MCD; sensitivity 0.92, precision 0.87, accuracy 0.96, F1 score 0.89, AUC 1.00). AI also showed higher accuracy than senior pathologists for IgAN, MN, and MCD; however, comparator evidence was sparse and should be interpreted cautiously. Most included studies were retrospective, and substantial heterogeneity was observed across datasets, imaging modalities, model architectures, and validation strategies.
Pathology-based AI shows strong potential for GD classification, but current head-to-head evidence is insufficient to establish superiority over pathologists, particularly senior pathologists. Prospective multicenter studies integrating multimodal clinical data and standardized external validation are needed.
Artificial intelligence based on renal pathology images is being developed to support classification of glomerular disease.
This systematic review synthesizes pathology-image AI models and evaluates performance in relation to validation context and heterogeneity.Potential impact: These models may support reproducible screening and quantitative assessment under pathologist oversight, but prospective multicenter validation remains necessary.