The Second Opinion
This site publishes two independent forecasts of the same seasons — the analog engine and SEAS5 — and they disagree about the sign of the temperature anomaly over roughly a quarter of the world. Publishing both without saying which to believe was a gap we made ourselves. This page closes it, and the answer is not the one we were hoping for.
Matched on
SEAS5
Analog wins
Blending gains
The verdict
matched over 24 years · ACC, global then landBoth engines scored on identical footing: the same 24 years, the same three-month target, detrended over those years, on one grid. Each is still used as it would be operationally — the analog draws on its full 1940–2025 library to forecast each year, SEAS5 gets only what a reforecast gives it — but the scoring is the same.
| Season | Analog | SEAS5 | 50/50 blend | Where the analog wins |
|---|---|---|---|---|
| ASO 2026 | 0.376 0.172 | 0.655 0.490 | 0.597 0.412 | 6% 11% land |
| SON 2026 | 0.346 0.204 | 0.516 0.340 | 0.500 0.336 | 18% 27% land |
| OND 2026 | 0.300 0.169 | 0.468 0.316 | 0.449 0.296 | 19% 27% land |
| NDJ 2027 | 0.288 0.156 | 0.442 0.313 | 0.425 0.277 | 22% 25% land |
The blend is a flat 50/50 of the two standardised forecasts, untuned. Searching for the best weight finds 0.1–0.2 and buys about +0.01 ACC — fitted on the very data it is scored against, over 24 years. That is not a gain, it is a measurement of how little the analog adds once SEAS5 has spoken.
Why the published numbers mislead
four mismatches at onceRoute 09 publishes an analog ACC of 0.475 and route 10 a SEAS5 ACC of 0.481. Set side by side those look like a tie, and an earlier version of this comparison concluded the analog was ahead at the longer leads. It was wrong, and it was wrong for four reasons at once — different target (seasonal against monthly), different period (86 years against 24), different detrending window, different grid.
| Comparison | Analog | SEAS5 | What is being measured |
|---|---|---|---|
| As published | 0.475 | 0.481 | Route 09's 86-year seasonal ACC against route 10's 24-year monthly ACC. Looks like a dead heat. It is not a comparison. |
| Matched | 0.376 | 0.655 | Same 24 years, same three-month target, detrended over those years, on one grid. SEAS5 wins by a distance. |
Seasonal averaging suppresses noise, so a three-month ACC sits well above the mean of its three monthly ones — which is why SEAS5 scores 0.66 here and 0.48 on route 10. And skill is not stationary: the analog engine's 0.475 is earned over 86 years, but over these particular 24 it manages 0.38.
Where each one wins
4 layers · 7 views · 4 seasons
The difference in cross-validated skill between the two engines, scored on identical footing. Red is where SEAS5 has verified better over 1993–2016, blue is where the analog engine has. Red covers most of the world — but not all of it, and the blue is not noise.
How to read this
Four things it is notVerdict
Not a reason to delete the analog
SEAS5 wins, and it should: it is initialised with a full three-dimensional ocean and atmosphere, while the analog engine sees a single monthly SST field. What the analog offers is independence — it shares none of the model's errors — and a forecast a reader can audit by looking up four years. Where the map is blue, it is the better of the two on this record.
Geography
The blue is not scattered
The analog engine wins over the high-latitude North Atlantic, Scandinavia, the Arctic and parts of the Southern Ocean — mid- and high-latitude regions where the signal is internal atmospheric variability rather than a tropical teleconnection, and where coupled models have long been weakest. SEAS5 wins almost everywhere the tropics drive the season. That is a physically coherent split, not noise.
Sample
Twenty-four years decides this
The reforecast period is the binding constraint on the whole comparison, and 24 is a small number. A pointwise ACC difference of 0.1 between two engines is well inside what this sample can produce by chance. Read the large-scale pattern of the verdict map; do not read one gridpoint.
Blending
Two forecasts are not better than one
Combining independent forecasts usually helps, and here it does not — because the two are not of comparable quality. A 50/50 blend lands below SEAS5 at every season, and the best weight found by searching is barely above zero. When one source is clearly better, averaging it with a weaker one is a way of throwing skill away.