Verification

Every forecast scored. Every number traceable.

This page sets out how we score ourselves, what an engagement looks like from your side, and what the verification report you receive contains. The same discipline applies to every offering.

How we score

The rules, fixed before the results.

CRPS, a proper scoring rule

The continuous ranked probability score is the primary metric. It rewards calibration and sharpness together and penalises overconfidence. MAE is reported alongside.

Operational climatology as the benchmark

Skill is stated relative to the baseline most users otherwise rely on. A number without a benchmark is not a result.

Out-of-sample splits

Backtests use forward and rolling splits over 20 years, so no forecast is scored on data its model saw in training.

Interval coverage

If the 80 percent interval is stated, outcomes should land inside it about 80 percent of the time. Coverage is reported unvarnished.

Significance and sample size

Results carry their sample sizes. Where a difference is within noise, it is described as within noise.

Null results disclosed

Where a station, month, or lead shows no improvement, the result is reported and the product returns climatology there.

What to expect

An engagement, from your side of the table.

01

Name the locations

You name the stations, contracts, fields, or regions, and the quantities that matter to you.

02

Protocol in writing

Locations, quantities, splits, benchmark, and score are fixed in writing before any result is seen.

03

Backtest, full record

We run the backtest and hand over every number, including the misses and the stations where we add nothing.

04

Live scoring

If you proceed, every issued forecast is archived as issued and scored as outcomes settle, for the life of the engagement.

The deliverable

What the verification report contains.

  • CRPS skill against climatology, by location, month, and lead time
  • Interval coverage at the stated probability levels
  • Trigger and exceedance frequencies where an index is involved
  • Null results, stated as plainly as the positive ones
  • A point-in-time archive reference, so any number can be reproduced later
A public scoreboard is in preparation.

Updated as outcomes settle. Until then, the full records behind this page are shared with prospective clients under a short confidentiality note.

In use today

Where the work is being used.

The discipline above is not a proposal. It runs in production and has been exercised under protocol.

Scored against market prices

Our own daily practice: where an exchange lists a temperature futures contract, we compare our calibrated settlement forecast with the market price on every trading day and score it at settlement. The market price is a demanding benchmark, the consensus of professionals with money at risk. The daily record is archived.

Pre-registered backtests

Benchmarks run under protocols fixed in writing before any result is seen: locations, quantities, out-of-sample splits, benchmark, and score. Results are handed over in full, including the locations and months that show no improvement over climatology.

NJ AI Challenge dashboards

Probabilistic weather and demand-risk dashboards for utilities, residents, and trading desks, built with Plug and Play and supported by an NJEDA grant. Open on this site.

Freeze-risk index

A freeze-risk index for fruit growers in the Northeast United States, in development with regional partners, built to the same scoring discipline.

Start with a benchmark on your own locations.

We score our calibrated forecasts against climatology on the stations, contracts, or fields that matter to you, and share the full record. The benchmark runs before any commercial commitment.