Abstract / Summary
Vision-language foundation models (VLMs) have demonstrated broad competence across medical imaging tasks, raising the question of whether they can match purpose-built, specialised AI systems for high-stakes 3D CT screening tasks. This talk presents a retrospective benchmark comparing eight VLM configurations — MedGemma, M3FM, CT-CHAT, MedSigLIP, and COLIPRI (each evaluated zero-shot and/or fine-tuned) — against seven specialized, CT-native cancer detection AI models on the NLST held-out test set (n=2,065; 54 cancers). Specialized models consistently exceed AUC 0.94, led by the Eyonis LCS ensemble at 0.98, while VLMs range from near-chance zero-shot performance to AUC 0.84-0.91 depending on fine-tuning and prompt design — notably, COLIPRI's malignancy-grounded zero-shot prompting reaches AUC ~0.91 at a fraction of the compute cost (~15 TFLOPs/scan) of larger VLMs (up to 6,700 TFLOPs/scan). The key driver of this gap is CT-native domain pretraining, not model scale or general vision-language capability: COLIPRI, the only VLM in our benchmark pretrained natively on CT volumes, is also the only one to approach specialised-level AUC, while Pillar-0's CT-native vision-only encoder matches top CNN specialists using nothing more than a linear probe on frozen embeddings. To ground these results in clinical relevance, I will also draw on two companion studies benchmarking against radiologists directly. In a 12-reader study, native MedGemma (AUC 0.70) fell short of clinically relevant performance, while fine-tuned MedGemma (AUC 0.83) reached only the level of less-experienced readers (mean radiologist AUC 0.90; range 0.80-0.94). By contrast, in an independent multi-reader evaluation, the specialised Eyonis LCS AI outperformed all radiologists. Combined with the markedly higher computational burden of generic vision-language models, this gap currently favours CT-native specialised systems for safer near-term clinical use, with fewer diagnostic errors and unnecessary procedures. For the SAFER community, this underscores that faithful evaluation of foundation models must be paired with genuine domain adaptation — or, better still, dedicated 3D CT pretraining from the outset — not just scale, before deployment in screening pipelines.