FORECASTSEASON.COMRoute 03 · openSEAS5 ExplorerProbabilityAnalogsSecond OpinionCitiesModelsDriversAbout
SEASON90Forecast SeasonSeasonal Outlooks · U.S. · Europe · Global
Forecast Season/The Copernicus Multi-System
Route 03 · Multi-system & consensus

The Copernicus Multi-System

Every outlook you read comes from one centre, and almost none of them tell you whether the others agree. This puts ten systems from nine centres on one grid and one calendar, each calibrated against its own reforecast, and shows both what they say together and how far apart they are. Then it does the part that matters: it checks whether the consensus is actually better than the best single model, and publishes the answer either way.

SEASON90

Systems

10
from 9 centres — ECCC contributes two different models, not two versions of one

Members

704
and every system gets one vote regardless, for the reason below

Consensus vs ECMWF

+0.038 RPSS
September 2026 temperature over land — what the other nine centres add

Agreement

7.2 of 10
September 2026: average systems concurring over land

The plate

9 layers · 7 views · 6 months
Layer
View
Month

Each centre's probability computed inside its own ensemble against its own hindcast, then averaged. Climatology is 33 %. Tercile boundaries are moved to today's climate, for the reason route 10 sets out at length.

* Lead 1 is the initialisation's own month, so part of it has already happened. It is kept because the models produced it, not because it is a forecast for days that are over.
Every panel except the skill layers is stippled where the multi-system RPSS is zero or below — where issuing these probabilities would be no better than saying nothing. That includes the agreement layer, deliberately: ten models concurring in a place none of them can forecast is exactly the picture this site should not publish unmarked.

Agreement is not skill

the warning this route exists to give

A map showing ten models in agreement is the most persuasive thing this site could publish, and one of the easiest ways it could mislead you. These systems are not independent witnesses. They share ancestry, they share resolution limits, several share components outright, and they share the same well-known biases in how the tropical ocean drives the atmosphere. When they agree, they may simply be agreeing on an error — and a consensus of ten wrong answers looks exactly like a consensus of ten right ones.

That is why the agreement layer is stippled on skill rather than left clean, why the scoreboard below is on the same page rather than a link away, and why the panel worth reading first is not the consensus but the gain — what the other nine centres actually add to ECMWF alone, measured, and drawn whichever way it comes out.

Temperature

MonthChance above normalSystems agreeingLand unanimousSpread (pts)
August 202654% 56%7.0 /1017%22
September 202654% 53%7.2 /1020%16
October 202654% 54%7.4 /1023%15
November 202655% 54%7.5 /1026%14
December 202656% 55%7.6 /1028%13
January 202757% 58%8.1 /1036%12

Rainfall

MonthChance above normalSystems agreeingLand unanimousSpread (pts)
August 202636% 33%6.5 /107%16
September 202636% 33%6.6 /106%12
October 202637% 38%6.9 /109%12
November 202637% 38%7.2 /1012%12
December 202636% 37%7.1 /1010%12
January 202736% 36%6.9 /109%11

"Systems agreeing" counts how many put their highest probability on the same tercile as the published consensus, averaged over land. The floor is not zero but roughly 3 — with three categories and no signal at all, a third of the systems land on the same one by chance.

The scoreboard

1993–2016 reforecasts vs ERA5 · n=24

Every system scored inside one leave-one-out, over land, on identical footing: for each held-out year each system's tercile boundaries are rebuilt from its own other 23 reforecast years before the ten are averaged. Calibrating once on all 24 and holding out only at the averaging step would score each model partly against its own answer, and the consensus would inherit ten times that flattery rather than escape it. The best single system in each row is marked; the last two columns say by how much the consensus beat it and how many systems beat the consensus.

Temperature

Monthbom-2cmcc-4dwd-22eccc-4eccc-5ecmwf-51jma-4meteo_france-9ncep-2ukmo-610Consensusvs bestBeat it
August 2026+0.161+0.116-0.127+0.084+0.197+0.240+0.126+0.150+0.064+0.192+0.263+0.0230
September 2026+0.036-0.002-0.073+0.025+0.044+0.067-0.002+0.059-0.009+0.069+0.105+0.0370
October 2026+0.047+0.034-0.022+0.046+0.035+0.073+0.002+0.062+0.013+0.063+0.107+0.0350
November 2026+0.024+0.004+0.013+0.031+0.025+0.057-0.029+0.061-0.010+0.046+0.087+0.0260
December 2026+0.005+0.026-0.014+0.004+0.023+0.031-0.016+0.030+0.008+0.032+0.073+0.0420
January 2027+0.013+0.018-0.016+0.015-0.003+0.041-0.026+0.040+0.000+0.031+0.071+0.0290

Rainfall

Monthbom-2cmcc-4dwd-22eccc-4eccc-5ecmwf-51jma-4meteo_france-9ncep-2ukmo-610Consensusvs bestBeat it
August 2026+0.057+0.029-0.130-0.007+0.053+0.128+0.011+0.026-0.002+0.081+0.139+0.0100
September 2026-0.035-0.031-0.061-0.043-0.025-0.011-0.074-0.007-0.039-0.011+0.028+0.0360
October 2026-0.026-0.003-0.028-0.037-0.025+0.002-0.070-0.008-0.039-0.011+0.032+0.0290
November 2026-0.011-0.010-0.001-0.023-0.009+0.008-0.065+0.003-0.024+0.005+0.040+0.0320
December 2026-0.003+0.007-0.018-0.020-0.011+0.003-0.052-0.005-0.025-0.009+0.038+0.0320
January 2027-0.027-0.013-0.031-0.035-0.043-0.014-0.088-0.018-0.024-0.015+0.021+0.0340

The rainfall result is the strongest case this route makes. Taken one at a time, seasonal rainfall terciles are close to worthless: at every full lead between 7 and 10 of the 10 systems score negative RPSS over land — issuing their probabilities is worse than admitting you do not know — and in September 2026 that is 10 of 10, with ecmwf-51 itself at -0.0109. Averaged, the same ten systems are positive at 5 of 5 leads. Route 10 concluded that one model's rainfall terciles were mostly not worth issuing, and that stands. This is the amendment: as a consensus they cross into being worth issuing, barely — which is a smaller claim than it sounds and a larger one than nothing.

Twenty-four years is not many, and a per-system RPSS from 24 cases is noisy — the ranking between two close systems is not meaningful, though the gap between the consensus and the field usually is. The benchmark column uses ecmwf-51, the system route 10 publishes on its own.

Who is in

10 systems · 704 members

Membership is not a fixed club. CDS is asked which systems carry each initialisation month and the answer changes through the year — the Met Office moved from system 604 to 605 to 610 across 2026, and JMA has no system at all for some months. What follows is who was actually in this run.

CentreSystemCDS idMembersReforecast members
Bureau of MeteorologyACCESS-S2bom-212127
CMCCSPSv4cmcc-45030
DWDGCFS2.1dwd-225030
Environment and Climate Change CanadaCanESM5.1 / GEM5.2-NEMOeccc-42020
Environment and Climate Change CanadaCanESM5.1 / GEM5.2-NEMOeccc-52020
ECMWFSEAS5.1ecmwf-515125
JMACPS3/CPS4jma-415510 regridded from 145×288
Météo-FranceSystem 9meteo_france-95131
NCEPCFSv2ncep-212424
Met OfficeGloSea6ukmo-6106228

Member counts range from 20 to 155, which is why each system gets one vote rather than one vote per member: pooling members would let the largest ensembles outvote the rest and produce a map mostly of who runs the most members. Note also that a system's reforecast is usually much smaller than its live ensemble — the probabilities are as fine-grained as the live members allow, but the skill estimate is only as good as the reforecast behind it.

How to read this

Four things it is not

Consensus

Not a better model

Averaging systems cancels the errors they do not share and keeps the ones they do. That is usually worth a little skill, and occasionally worth none at all. Route 11 found the same thing blending two engines: when one source is clearly better, averaging it with a weaker one throws skill away. The scoreboard is on this page so you can check rather than assume.

Independence

Ten opinions, fewer minds

Several of these systems share an ocean model, a land surface, or a convection scheme. Two of them come from the same centre. Treating the count of agreeing systems as a sample size — as though nine of ten were nine independent confirmations — overstates the evidence, sometimes by a lot.

One grid

Nearly true

Nine of the ten arrive on the common one-degree mesh. JMA arrives on its native 1.25° grid and is interpolated onto the others. That interpolation adds no information — JMA's forecast still holds only what its own grid holds — and it is named in the roster rather than hidden in a method note.

Calibration

Each against itself

Every system's probabilities come from its own reforecast, never from a shared climatology. A coupled model's drift is its own, and the boundaries that make its members mean anything are the boundaries of its own past runs. This is also what makes the ten comparable at all: they are each answering the same question about themselves.