Abstract / Summary
Preprint (author's original manuscript) submitted to Biomedical Engineering / Biomedizinische Technik on 26 September 2026; not yet peer reviewed. Code, the leakscope Python package and all replication materials: https://doi.org/10.5281/zenodo.22974070 Objectives: Machine-learning studies on physiological signals often let segments from the same person appear in both training and test data. We quantified when and how much such subject leakage inflates reported performance, and how often it occurs. Methods: Random-segment, within-subject block, night-wise and subject-wise evaluation were compared on five public datasets (MIT-BIH, PTB-XL, UCI HAR, WESAD, Sleep-EDF) with four model families. We derived a closed-form ceiling for identity-driven inflation and a pre-training leakage risk index, tested them in 720 simulated datasets, and audited 100 randomly sampled studies (2023–2026) that used the same datasets. Results: Inflation of macro-F1 ranged from ≤0.01 (PTB-XL) to 0.49 (MIT-BIH) and followed exposure and subject identifiability. In simulation, the risk index explained 85–88 % of identity inflation for flexible models, and leakage produced 91 % accuracy without any physiological signal. Overlapping windows added a separable adjacency component; per-subject standardisation reduced inflation. In the audit, 12 % of studies explicitly evaluated on already-seen subjects, 40 % reported too little to tell, and official subject-disjoint splits did not prevent leakage. Conclusions: Subject leakage is predictable from data properties before training. We provide a reporting checklist and an open-source toolkit, leakscope.