Abstract / Summary
While clinical guidelines remain essential for standardizing medical care, their utility decays as manual updating fails to keep pace with the exponential growth in medical evidence. To address this challenge, we developed a Multi-Agent Retrieval-Augmented Generation (MARG) system dedicated to clinical guideline development and maintenance. We employed a multi-phase evaluation framework to assess MARG’s performance across agent-agent interaction and human-AI collaboration tasks. The system was evaluated against seven 2025 liver disease guidelines, encompassing 137 clinical scenarios. Here we show MARG achieve 77.4% overall concordance with the guidelines, with consistency rates rising from 45.8% under weak consensus to 87.8% at full consensus. Domain-specific performance is 65.5% for diagnosis, 75.0% for prevention, and 81.0% for treatment. MARG also assists in evidence integration, outperforming manual review in recency (6.3% vs. 0.9%), reproducibility (25.5% vs. 7.5%), and quality (24.2% vs. 8.0% Level I evidence). Through four rounds of structured debate, the system resolves 95.8% of recommendation conflicts and achieves strong consensus (>85%) in 74.5% of final outputs. In a small-scale simulation involving 60 scenarios assessed by physicians with 5, 15, and 30 years of experience, MARG is linked to correction of 40-60% of guideline-inconsistent decisions, with the net gain (50%) most apparent in the early-career tier. These proof-of-concept findings suggest that MARG may support clinical guideline development and maintenance. Effective deployment will depend on a tiered human-AI collaboration model built upon curated evidence repositories, machine-interpretable grading protocols, guideline-specific benchmarks, and independent ethical governance. Clinical practice guidelines help doctors make consistent, evidence-based decisions. However, manual updating often falls behind new research, risking outdated recommendations. To address this, we developed an AI system called MARG to review and synthesize new evidence for guideline updates. Tested against seven liver disease guidelines, MARG accurately reproduced most recommendations, identified higher-quality evidence, and resolved disagreements through structured AI debate. In a small simulation with clinicians of varying experience, MARG resolve guideline inconsistencies, especially for early-career doctors. While promising, the system is not meant to replace human expertise. Future work should focus on real-world testing to ensure safety, transparency, and ethical governance before any clinical use. Here the authors develop multi-agent AI that acts like a digital committee to review evidence and flag potential guideline updates. In liver disease benchmarks, it replicated expert consensus and cut early-career clinicians’ guideline errors by half.