Abstract / Summary
EEG-based machine-learning studies of attention-deficit/hyperactivity disorder (ADHD) are vulnerable to optimistic performance estimates when feature selection or model choice is informed by data outside the training fold. We compared expert-reconstructed, large language model (LLM)-guided and full-space feature strategies under leakage-controlled subject-level validation in a frozen cohort of 77 children (45 ADHD, 32 controls) represented by 1,185 gamma-free EEG features. Four classifiers (SVM, GMM, random forest [RF] and XGBoost) were evaluated with nested leave-one-subject-out cross-validation (LOOCV), with all data-dependent operations restricted to outer-training subjects. Across 100 repetitions of stratified 10-fold cross-validation, RF/expert achieved mean accuracy 62.8% (SD 3.5%; empirical 95% interval 55.8% -- 69.5%), mean balanced accuracy 62.3% and mean AUC-ROC 0.618. The corresponding LOOCV point estimate was 71.4% accuracy and 71.0% balanced accuracy (AUC-ROC 0.673). GMM/expert achieved the highest AUC-ROC (0.699), indicating metric-dependent model ranking. The expert-reconstructed strategy achieved the highest accuracy and balanced accuracy for every classifier, whereas the LLM-guided strategy showed no consistent advantage over the full feature space. Pairwise classifier differences were not significant after Holm correction. A pre-specified 200 V artifact gate reduced RF/expert accuracy to approximately 55.8%, below the 58.4% no-information rate; this drop may reflect removal of artifact-related signal and/or reduced epoch counts leading to noisier feature estimates. These findings support cautious benchmarking and external validation rather than claims of algorithmic or LLM superiority. Keywords: ADHD; electroencephalography; machine learning; benchmarking; feature engineering; subject-level validation; nested cross-validation; random forest; clinical prediction