Abstract / Summary
Biomedical Machine Learning (ML) is increasingly evaluated through Independent and Identically Distributed (IID) benchmarks, even when deployment involves irregular observation, missingness, distribution shift, and high failure costs. This position paper argues for a stability-oriented reporting standard for safety-relevant biomedical ML. The central requirement is not a specific model class or optimizer, but auditable evidence: performance should be reported under clinically plausible perturbations such as thinning, timestamp jitter, bursty missingness, channel dropout, and Signal-to-Noise Ratio (SNR) shift. When authors invoke training-stability diagnostics, those diagnostics should be hazard-linked: their assumptions, operational meaning, and relationship to stress-test degradation should be stated explicitly. We use controlled mechanism probes to illustrate two hazards that IID scores can hide: observation-process dependence under irregular sampling, and perturbation sensitivity associated with high-curvature training regimes such as Edge of Stability (EoS). We do not propose Neural ODEs, continuous-time models, Sharpness-Aware Minimization (SAM), or EoS metrics as default prescriptions. Instead, we propose a minimum reporting standard that makes robustness claims testable and supports escalation to more complex modelling or optimization only when stress tests justify the cost.