Abstract / Summary
Abstract Integrating Large Language Models (LLMs) into radiology is often hindered by the inability of monolithic architectures to balance clinical reasoning with pedagogical safety. This study evaluates MedGemma-27B-IT across two specialized agentic frameworks: a clinical "Radiology Sparring Partner" and a "Socratic Teacher". For clinical decision support, we compared three architectures: Raw (base model output, no prompt), Minimal Prompted (single-shot prompt), and Multi-Agentic, using a 4-stage pipeline for clinical synthesis. Separately, the educational track employs a dual-agent architecture for progressive educational scaffolding. A pooled evaluation of 120 interaction transcripts (10 per architecture with three independent LLM-as-a-judge runs) was finalized using an expert-anchored combined mean approach, anchored via Bühlmann credibility calibration against radiologist-authored references (5 per architecture). Results demonstrate that agentic orchestration outperforms monolithic baselines in high-level reasoning. Most notably, Pathophysiological Logic scores rose from 3.20 (Raw) to 4.14 (Agentic) on a 5-point Likert scale, with a simultaneous reduction in standard deviation from 1.23 to 0.70. While this Proof of Concept (PoC) highlights the advantages of modularity, results remain subject to evaluator formulation bias and a measurable "latency tax" inherent in sequential agent hand-offs. Future work will focus on longitudinal clinical validation and optimization of cached memory architectures to refine reasoning depth while mitigating latency.