Abstract / Summary
Early identification of respiratory and cardiovascular problems depends on the accurate characterization of cardiopulmonary sounds. Nevertheless, the majority of current research concentrates on the independent analysis of either heart or lung sounds and frequently employs a single feature representation. In response to these shortcomings, this paper presents XCardioPulmoNet, an image-based framework that develops an approach to classifying heart and lung sounds. Through the utilization of Mel spectrograms and Wavelet transforms, the proposed method converts audio inputs into representations that are based on time and frequency. The subsequent step involves the reconstruction of a hybrid image by merging the information that is retrieved from both representations, which are complementary to one another. To improve multiclass classification performance and discriminative feature learning, the proposed EAGR_Net, a residual attention-guided convolutional neural network, is used to categorize the produced hybrid images. Furthermore, Gradient-Weighted Class Activation Mapping (Grad-CAM) is used to depict the discriminative regions that influence the model’s predictions and improve interpretability. Five-fold cross-validation is used to test the baseline CNN and Attention CNN models. With a ROC-AUC of 0.9851 and PR-AUC of 0.9558 on the heart sound dataset and a ROC-AUC of 0.9978 and PR-AUC of 0.9917 on the lung sound dataset, respectively, the experimental findings show that EAGR_Net outperforms the baseline models. These findings demonstrate the framework’s robustness, interpretability, and diagnostic potential for automated multiclass cardiopulmonary sound classification and disease detection.