Abstract / Summary
Purpose Open-source vision-language models (VLMs) can be locally deployed without external internet access. This study aimed to evaluate the slice-level lesion-detection performance of multiple general-purpose and medical-purpose open-source VLMs for brain metastases on contrast-enhanced (CE) MRI, and to assess their accuracy in characterizing detailed imaging features on lesion-positive images. Materials and Methods One hundred lesion-positive axial CE T1-weighted images and 100 matched lesion-negative images from 100 patients were analyzed using eight evaluable VLMs: five general-purpose VLMs-InternVL3-8B, Qwen2.5-VL-7B-Instruct, MiniCPM-V-4.5, DeepSeek-VL2-tiny, and Phi-3.5-vision-instruct-and three medical-purpose VLMs-MedGemma-4B-it, LLaVA-Med v1.5, and HuatuoGPT-Vision-7B. Lesion-detection performance was assessed using sensitivity, specificity, and balanced accuracy. On lesion-positive images, accuracy was evaluated for lesion count, laterality, anatomic location, enhancement pattern, necrosis, and mass effect. Balanced accuracy was reported with patient-level cluster bootstrap 95% confidence intervals (CIs). Models were compared using the Friedman test and paired patient-level permutation tests with Benjamini-Hochberg correction. Secondary outcomes were compared with majority-class baselines using exact McNemar tests. A revised prompt was tested post hoc. Results The median age of the study patients was 65 years (IQR, 59.0-70.0), and 58 patients were male (58.0%). MiniCPM-V-4.5 showed significantly higher balanced accuracy (80.5%; 95% CI, 75.0-85.5) than all other models (FDR-adjusted P < 0.001). Seven of the eight models showed predicted-positive rates of 81-100%, and 70.8% (17/24) of responses to synthetic control images without anatomic content were classified as lesion-positive. With the revised prompt, balanced accuracy increased from 59.5% to 77.0% in HuatuoGPT-Vision-7B and from 51.5% to 68.5% in InternVL3-8B. No model significantly outperformed the majority-class baseline for enhancement pattern or for mass effect. Conclusion This slice-level technical benchmark demonstrated variable lesion-detection performance of open-source VLMs for brain metastasis in an image-only setting, with most models showing a strong positive-response tendency.