Why long-range forecasts need calibration before anyone can price on them
Weather models are skilful well beyond two weeks, but the raw output is biased and overconfident at exactly the horizons where energy, insurance, and agriculture decisions are made. This note explains what calibration does, why it has to be local, and how to tell a calibrated forecast from a raw one.
A utility buying power for January, a desk pricing a temperature contract, an insurer setting a frost premium: each of these decisions is made weeks or months before the weather arrives. The question is whether a forecast at that range is worth anything, and if so, how to use it.
What a raw ensemble gives you
Global weather centres run their models many times from slightly different starting points. The spread of those runs is meant to describe the uncertainty of the forecast. At short range this works well. Beyond about two weeks the picture changes in two ways.
First, the forecast drifts. Every model has systematic errors that depend on place and season, and at long range these biases are often as large as the signal itself.
Second, the spread is too narrow. The runs agree with each other more than they agree with what eventually happens. A user who reads the spread as a probability will be surprised far more often than the forecast implies.
Neither problem means the forecast is useless. It means the forecast is an input, and the output people need has to be built from it.
What calibration does
Calibration compares decades of past forecasts with what actually happened, location by location, season by season, lead time by lead time. From that record it learns two corrections: where to move the centre of the distribution, and how much to widen it. The result is a forecast whose stated probabilities match observed frequencies. When it says 30 percent, the event happens about 30 percent of the time.
Calibration has to be local because the errors are local. A model's warm bias on the coast in spring says little about its behaviour inland in winter. It also has to be lead-specific, because the corrections for a seven-day forecast differ from those for a seven-month forecast.
How to tell the difference
Three questions separate a calibrated forecast from a raw one.
- Is it scored? A calibrated product comes with a verification record, ideally on an out-of-sample period, using a proper scoring rule such as the continuous ranked probability score.
- Is it compared with the right baseline? At long range the relevant baseline is climatology, the 30-year average and its spread, because that is what most users fall back to. A forecast should show where and when it improves on that baseline, and say plainly where it does not.
- Are the intervals honest? An 80 percent interval should contain the outcome about 80 percent of the time. A forecast that reports its own coverage, including the cases where it falls short, is one you can plan on.
Why this matters now
The raw model output that used to be expensive is increasingly open. What remains scarce is the calibration and verification layer that turns it into a number a desk, an underwriter, or a grower can act on. That is the layer we build.
