Abstract / Summary
Abstract Cardiovascular diseases, particularly myocardial infarction (MI), remain a leading cause of global mortality, necessitating accurate and early diagnostic solutions beyond conventional operator-dependent echocardiographic analysis. The proposed hybrid deep learning framework in this study combines the region-specific and global representations for automatic cardiac disease classification from echocardiographic videos. It uses transformer-based UltraSAM to accurately segment the left-ventricular (LV) myocardium, and then extracts various features such as shape, texture (GLCM), motion (optical flow), and keypoint-based descriptors to characterize the myocardial structure and dynamics. Concurrently, a Swin Transformer is used to obtain the global spatial representations from complete echocardiographic frames. The heterogeneous attributes are fused and optimized by employing Ant Colony Optimization (ACO), and temporal dependency is modeled by applying LSTM for robust classification. Experimental evaluation on the HMC-QU dataset achieves an accuracy of 94.7% with significant sensitivity (93%) and specificity (97%) for MI detection. In addition, the model generalizes well on the CardiacNet dataset with an accuracy of 90.3% in multi-class classification (PAH, ASD, Healthy), especially in the healthy subjects (95%). The results here confirm the efficiency of the proposed hybrid framework in capturing ventricular abnormalities at the local level and capturing the general structure of the heart for the reliable detection of cardiovascular disease.