Abstract / Summary
Mammography interpretation requires tying each finding to supporting evidence and withholding judgment when that evidence is unavailable. Vision-language models are increasingly used for image-based medical question answering, yet existing benchmarks suggest that they struggle to jointly support accurate answers, evidence grounding, and appropriate abstention. We propose CARE-MVLM, built on Qwen2.5-VL-7B-Instruct, to address this gap through separate modules for answer prediction, answer-conditioned evidence generation, and answerability selection. On 7,358 triplets curated from CBIS-DDSM, MIAS, and VinDr-Mammo, CARE-MVLM substantially improves lesion-sensitive selective behavior over same VLM backbone baselines and exhibits markedly stronger lesion-sensitive abstention than GPT-5.6 and Gemini-3.5-flash while attaining the highest joint answer, grounding, and abstention reliability in the comparison.