Abstract / Summary
Abstract Epileptic seizures affect approximately 65 million people worldwide, and standard antiseizure medications fail to achieve seizure control in roughly one-third of patients [4], [6]. Automated, EEG-based seizure detection offers a promising avenue for improving diagnosis and enabling timely clinical intervention. In this study, we evaluate and compare logistic regression, random forest, support vector machines (SVM), and XGBoost for binary classification of seizure versus non-seizure EEG activity using the UCI Epileptic Seizure Recognition dataset (11,500 EEG segments, 500 subjects) [16], [17]. We evaluate frequency-domain (FFT) feature augmentation, hyperparameter optimization, and two class-imbalance mitigation strategies (class-weight balancing and SMOTE). XGBoost trained on combined time- and frequency-domain features achieved 98.19% test accuracy with a seizure-class F1-score of 0.95. A SMOTE-resampled voting ensemble achieved a comparable 98.05% accuracy while further improving seizure recall to 0.98, a trade-off that would likely be favorable in clinical seizure-detection contexts, where missed seizures typically carry a higher cost than false alarms, though this comparison is benchmark-level rather than clinically validated. Five-fold cross-validation (97.63% ± 0.35% for random forest) confirms these results are stable. We further report and correct a data-provenance issue encountered during the study, in which an initial experimental run inadvertently used a synthetic placeholder dataset due to a network connectivity fallback, underscoring the importance of validating dataset integrity in reproducible machine learning research.