Abstract / Summary
Introduction: Vision-language models (VLMs) are increasingly explored for clinical interpretation tasks. Electrocardiographic (ECG) identification of acute coronary occlusion (ACO) is a high-risk application requiring accurate pattern recognition, and medical VLM performance on this task remains poorly characterised outside benchmark datasets.,Methods: A medical-domain VLM (MedGemma 1.5) was tested between 30 January and 22 February 2026 using 15 de-identified 12‑lead ECGs: 14 source-labelled ACO or high-risk occlusion phenotypes and one normal control. Each ECG was submitted with a standardised prompt requesting a structured interpretation and overall concern level (low, intermediate, or high), without iterative prompting. The same ECGs were submitted to a general-purpose VLM comparator (Gemini 3.1 Pro).,Results: The model assigned low concern in 8/15 cases (53%) and intermediate concern in 7 (47%); no case reached high concern, including among the 14 high-risk ECGs. It under-recognised ST-segment elevation, failed to integrate reciprocal and territorial patterns, and missed hyperacute T-wave morphology. Conduction abnormalities and quantitative parameters were assessed inconsistently. The comparator assigned high concern to all 14 high-risk cases but also to the normal control, citing an incorrect ECG pattern.,Conclusion: A medical VLM showed clinically important interpretation failures and did not escalate concern in any high-risk ECG. A general-purpose comparator was more consistent across high-risk cases but misclassified the normal control, suggesting some failure modes may not be modelspecific. These findings support task-specific validation and governance before multimodal AI is deployed in time-critical cardiovascular settings.