Abstract / Summary
Background. Public-health use of influenza surveillance - situational awareness, planning for clinical surge, and issuing advisories or warnings - needs more than an accurate point estimate of next week's count. What matters operationally is whether a rise is caught early, whether the direction of change is right, and whether an alert threshold is about to be crossed, all with honest uncertainty. For such use a forecast must also respect reporting delay, later data revision, and the calibration of its prediction intervals. Using Japan's sentinel surveillance, we evaluated short-term probabilistic forecasts that respect the data release history, and we measured what adding other surveillance series contributes, focusing on the properties that matter for public-health decisions. Methods. The target was the number of influenza reports per sentinel site in each of the 47 prefectures. For each forecast we predicted, as quantiles, the value 1-4 weeks after the most recent observed week. We compared three models: a ridge regression trained jointly across prefectures; a gradient-boosting quantile regression (GBQR) that learns from many surveillance series; and their equal-weight combination on the log scale. Six seasons were used to develop the models. The 2025/26 season was held out as a temporal evaluation period after the specification was fixed; because its epidemic profile had already been seen, it is not a fully independent prospective test. We scored forecasts with the weighted interval score (WIS) and checked prediction-interval coverage. For the Japanese target series we used the value first published for each week; some auxiliary series were taken as retrieved and therefore include later revisions. Results. In 2025/26, the combination's WIS relative to a persistence forecast was 0.841, 0.811, 0.794 and 0.818 at horizons of 1-4 observed weeks - that is, 16-21% lower than persistence. Coverage of the 80% prediction interval was 79.7%, 77.1%, 74.5% and 75.1%, and of the 95% interval 94.3%, 93.4%, 92.6% and 91.8%. Adding auxiliary series improved the GBQR over the development period. However, on the original scale the combination's edge over ridge regression was not statistically clear, and in some low-level rising phases the combination did worse than ridge. Conclusions. Over the period studied, models that share information across regions scored better than a persistence forecast. For public-health use, however, the operative questions are the real lead time once reporting delay is accounted for, whether rises from low levels and threshold exceedances are caught in time, and whether the upper prediction intervals are wide enough; on these, the present evidence is not yet sufficient. The 2025 change to the sentinel system also means that pre-change advisory and warning thresholds cannot be read mechanically. A prospective evaluation with strictly controlled data availability is needed before such forecasts are used to support advisories or planning.