Abstract / Summary
Voice-based cognitive load estimation is a long-standing target for unobtrusive workload monitoring, but existing data-driven classifiers are rarely grounded in the physiological mechanisms that generate cognitive load related acoustic change. This paper presents the Analytical Physiology Acoustics (APA) framework, a closed form model derived from established vocal fold biomechanical and autonomic theory, in which every constant is sourced from peer reviewed publications and validated against five independent published benchmarks before any training data are generated. The validated APA model is used to synthesize a 30 speaker, 1080 utterance dataset, on which a Physiology Informed Deep Neural Network (PI-DNN) is trained; its source filter heads, vocal-fold physics residual, and privileged physiology auxiliary pathway each correspond to a named APA derivation stage. The evaluation compares seven models under leave one subject out cross validation, including two independent gradient-boosting classifiers as modern baselines for tabular acoustic features. All pairwise comparisons against PI-DNN are corrected for multiple comparisons using the Holm–Bonferroni procedure with bootstrap confidence intervals on effect size, and a Friedman omnibus test with Nemenyi post-hoc analysis is reported across all models. The omnibus test indicates that the models are not all equivalent, but the pattern of differences does not favour architectural complexity: PI-DNN significantly outperforms only the support vector machine and one gradient boosting baseline after correction, and by mean rank across folds it places second behind ordinary logistic regression, a gap well inside the Nemenyi critical difference and therefore not statistically distinguishable. PI-DNN and the plain multilayer perceptron agree on the large majority of predictions. We report this outcome without embellishment: on the present synthetic dataset, the principal defensible contribution of this work is the validated generative framework and its validation-first methodology, not a demonstrated classification advantage over strong conventional baselines. All experiments are conducted exclusively on data generated by the authors’ own validated model; this is a closed-loop evaluation and is explicitly not a claim of real-world generalization. We discuss this limitation prominently, identify four publicly documented speech-bearing candidate datasets for future cross-dataset validation with verified access terms, clarify which widely cited affective datasets are unsuitable for speech-based cognitive-load validation, and release the complete data-generation, modelling, and statistical-analysis code.