Abstract / Summary
Infectious disease forecasting supports public health preparedness and response by informing resource allocation, intervention planning, and situational awareness. Infectious disease forecasting capabilities in the United States have grown over the past decade yet whether forecast quality has improved over this time remains unclear. Current evaluation metrics describe relative performance for a common set of forecast targets but cannot evaluate performance across different outcomes, seasons, and pathogens. We addressed this challenge by developing a novel skill score metric to assess forecasts based on performance relative to two benchmark models: a hindcast as an idealized upper bound of performance and a baseline model as a minimal performance standard. To study the impact of the baseline choice, we considered three types of baseline models with varying levels of complexity. Using this skill score framework, we retrospectively evaluated short-term influenza (influenza-like illness and hospitalization) and COVID-19 (case, death, and hospitalization) forecasts, focusing on how the performance of team and multi-model ensemble forecasts changed over time. The ensemble outperformed 99% of forecasts from a simplistic, naive baseline, and 84-91% of forecasts from more complex baseline models that incorporate recent data. We found consistent improvements in ensemble and team forecast skill over time for both influenza forecast targets, with 4-week horizon forecasts in the most recent year having higher skill scores than 1-week horizon forecasts in the first year of each challenge. Results were mixed for COVID-19 forecasts. This framework provides a comprehensive picture of progress in infectious disease forecasting over the past decade enabled by improved standardization, data, models, and collaboration.