Abstract / Summary
To understand spoken language, listeners must extract meaning from acoustic input that is variable and noisy. The noisy-channel framework posits that listeners overcome this uncertainty by rationally inferring the speaker's intent, combining prior expectations with a model of how the message may have been corrupted during transmission. However, the exact nature of this noise model and its sensitivity to context (e.g., speaker or acoustic environment) remain unknown. Here, I use eye-tracking to measure human speech processing as it unfolds on the millisecond timescale, and I use the representational geometry of a self-supervised speech model (HuBERT) to estimate the noise model. I show that spoken word comprehension is well captured as real-time Bayesian belief updating. Further, listeners flexibly tune their expectation of noise to the specific acoustic environment. These results extend the noisy-channel framework to the moment-by-moment comprehension of speech. More broadly, these findings demonstrate that the human language system continuously adapts its internal model to the statistical structure of the communicative environment. By linking noisy-channel inference with real-time perceptual dynamics, this work highlights the remarkable flexibility of human cognition in navigating uncertainty and opens the door to investigating internal representations in human and artificial speech recognition systems through their fine-grained temporal dynamics.