
Image: METAL
Summary
- In the CDC's FluSight 2025-2026 season evaluation released on September 30, Google's Google_SAI-FluEns ranked first among individual team submissions out of the 39 models analyzed.
- Google's model scored a relative WIS of 0.56, ahead of three models at 0.58 and the CDC ensemble (0.62, 7th), though other models did better in Texas, Massachusetts and Connecticut.
- The forecasting model was built with ERA, Google's tool that uses Gemini to write and optimize scientific code, which was published in Nature in May.
A flu hospitalization forecasting model built by Google Research ranked first in the U.S. Centers for Disease Control and Prevention's (CDC) evaluation of 2025-2026 season flu forecasts. In its FluSight season evaluation report released on September 30 (local time), the CDC said Google's Google_SAI-FluEns delivered the best results among individual team submissions out of the 39 models analyzed. Google built the forecasting model with its own AI research tool, Empirical Research Assistance (ERA). The original Google announcement that METAL reviewed runs four paragraphs and does not say what input data the model used.
FluSight is a CDC program that collects weekly flu hospital admission forecasts from academic, industry and government forecasting teams and combines them into one. Submissions for this season ran from November 22, 2025, to May 20, 2026, and forecasts were published on the FluSight webpage starting December 12, 2025. Each team had to forecast hospital admissions for the current week and up to three weeks ahead for all 50 states, Washington, D.C., and the nation as a whole. The CDC uses the combined forecast each week to communicate anticipated state-level demand for medical services. The CDC has hosted the forecasting challenge every year since the 2013-2014 season, except for 2020-2021, when flu activity was limited.
This season, 34 teams submitted 53 models, and 39 models that submitted at least 75% of forecast targets were included in the evaluation. The primary scoring metric was the relative weighted interval score (WIS). It measures how well a forecast's prediction intervals matched observed values, with a baseline model that carries the prior week's admissions forward set at 1; a score below 1 means a forecast beat the baseline. Of the 39 models, 33 beat the baseline.
In Table 1 of the CDC report, Google's model had the lowest relative WIS at 0.56. OHT_JHU-nbxd, CMU-TimeSeries and UGA_flucast-INFLAenza followed at 0.58, and UMass-flusion scored 0.59. Google's model hit 51.2% coverage on its 50% prediction intervals and 93.72% on its 95% intervals, close to the 50% and 95% targets, and it submitted 98% of all forecast targets. The FluSight ensemble, which the CDC uses when issuing forecast messaging, scored 0.62 and ranked 7th.
By state, the rankings diverge. In the same report's jurisdiction-level table, Google's model scored 0.37 in California, 0.42 in Texas and 0.45 in North Carolina. In Texas, NEU_ISI-AdaptiveEnsemble (0.35), NAU-vulPES (0.40) and UMass-flusion (0.41) scored lower than Google, and in Massachusetts and Connecticut, CMU-TimeSeries beat Google's 0.51 and 0.59 with 0.37 and 0.40. The CDC ensemble was one of 12 models that consistently outperformed the baseline in all jurisdictions.

The season itself was hard to forecast. The CDC classified it as moderate, and weekly hospital admissions peaked above 40,000. Admissions began rising in mid-November 2025 and peaked nationally in the week ending December 27. The CDC said in the report that ensemble forecasts "may not always reliably predict rapid changes in disease trends." Indeed, the ensemble's 50% and 95% prediction intervals failed to anticipate the surge in late December and the drop in mid-January, and during the peak week fewer than 25% of two-week-ahead prediction intervals across jurisdictions contained the observed values. Ensemble coverage stabilized near 95% from February 2026.
ERA, the tool behind Google's model, is a research tool that uses Gemini to write and optimize scientific code. Google formally introduced ERA in a paper published in Nature on May 19, 2026, titled "AI system designed to help scientists write expert-level empirical software." Given a scientific problem and a measure of success, ERA searches the literature, writes code, explores solutions, combines techniques and evaluates the results. It uses a tree search approach to weigh thousands of options and refine code toward its goal. The paper reported that ERA achieved expert-level performance across benchmarks in genomics, public health, satellite imagery analysis, neuroscience prediction, time-series forecasting and mathematics. Lizzie Dorfman, a product manager at Google Research, and research scientist Michael Brenner said "one of AI's greatest potential benefits to humanity is increasing the speed and scope of scientific discovery."
Infectious disease forecasting is a flagship application of ERA. According to Google Research, models built with ERA forecast weekly hospital admissions for flu, COVID-19 and respiratory syncytial virus (RSV) at the state level up to four weeks ahead, and ranked at or near the top of public CDC leaderboards for all three viruses. According to reports, the researchers described in a May 15 preprint a system in which a large language model guides a tree search that writes, tests and refines forecasting software, and reported that in a real-time prospective evaluation an ensemble of machine-generated models performed as well as or better than the CDC's human-curated hub ensembles. The researchers say it also produced workable results for RSV forecasting starting from almost no data. The epidemiological forecasting work was led by Zahra Shamsi, Sarah Martinson, Nicholas Reich, Martyna Plomecka and Brian Williams.
ERA is also being used beyond infectious disease. Google said a California snowmelt runoff forecasting model built with ERA produced more accurate early predictions than Bulletin 120, the state's official water supply outlook. It also produced a model that estimates carbon dioxide concentrations every 10 minutes from GOES-East geostationary weather satellite data, and a retail forecasting model that rivals the Chicago Fed's retail sales forecast. ERA's underlying technology is offered to trusted testers as a pilot science tool, and Google is opening Computational Discovery, built with ERA and AlphaEvolve, as a pilot tool within Gemini for Science. METAL previously reported that Google launched Gemini for Science, an AI tool for research.

Zahra Shamsi of Google Research, who wrote the post on the forecasting model, said "this performance validates our confidence that AI combined with human ingenuity will improve our ability to forecast diseases worldwide." The gap between first and second place was 0.02 in relative WIS, and state-level rankings were mixed. Still, in a competition where people have long refined statistical and mechanistic models, forecasting code written and revised by AI finished at the top of a season's scorecard. The CDC will soon release its evaluation of forecasts for the share of emergency department visits due to flu in the same season.
Sources
- Google — Google's science AI ranks #1 in CDC evaluation →
- CDC — FluSight 2025-2026 Evaluation →
- Google Research — Empirical Research Assistance (ERA): From Nature publication to catalyzing Computational Discovery →
- Unite.AI — Google AI Model Tops CDC Evaluation of Flu Hospitalization Forecasts →
- Google Research (X) — Today the @CDCgov announced that of 39 eligible models, Google's science AI model ranked highest →





Comments