Abstract / Summary
Heart sounds provide a valuable non-invasive source of information for the early detection of cardiac abnormalities. However, reliably distinguishing normal from abnormal heart sounds remains challenging in routine clinical practice because interpretation depends strongly on expert auscultation and is often affected by signal variability, recording artifacts, and background noise. Therefore, this study is presented as a transparent comparative evaluation of classical machine-learning methods for binary heart sound classification rather than as a novel deep-learning approach. The proposed framework consists of preprocessing, segmentation, handcrafted feature extraction, and comparative classification using Support Vector Machines (SVM), K-Nearest Neighbors (KNN), and ensemble classifiers. Three experiments were designed to assess the contribution of different handcrafted feature sets: (i) statistical, entropy-based, and frequency-domain descriptors (14 features), (ii) Mel-frequency cepstral coefficients (MFCCs) only (13 features), and (iii) a fused 27-feature representation combining all handcrafted descriptors. The results show that statistical, entropy-based, and frequency-domain descriptors alone provide a useful baseline but are clearly weaker than MFCC-centered representations. MFCC-only features produced a substantial improvement in classification performance, confirming their strong discriminative value for heart sound analysis. The best results were achieved with the fused 27-feature representation, showing that complementary handcrafted descriptors further improve robustness when combined with MFCCs. In this setting, KNN Fine achieved the highest overall cross-validation accuracy and F1-score, reaching $$0.944 \pm 0.008$$ accuracy, $$0.963 \pm 0.005$$ F1-score, and $$0.955$$ held-out test accuracy. Overall, the findings highlight the dominant contribution of MFCCs, the added benefit of feature fusion, and the value of evaluating classifiers with multiple complementary metrics.