
Every forecast scored. Every number traceable.
This page sets out how we score ourselves, what an engagement looks like from your side, and what the verification report you receive contains. The same discipline applies to every offering.
The rules, fixed before the results.
The continuous ranked probability score is the primary metric. It rewards calibration and sharpness together and penalises overconfidence. MAE is reported alongside.
Skill is stated relative to the baseline most users otherwise rely on. A number without a benchmark is not a result.
Backtests use forward and rolling splits over 20 years, so no forecast is scored on data its model saw in training.
If the 80 percent interval is stated, outcomes should land inside it about 80 percent of the time. Coverage is reported unvarnished.
Results carry their sample sizes. Where a difference is within noise, it is described as within noise.
Where a station, month, or lead shows no improvement, the result is reported and the product returns climatology there.
An engagement, from your side of the table.
Name the locations
You name the stations, contracts, fields, or regions, and the quantities that matter to you.
Protocol in writing
Locations, quantities, splits, benchmark, and score are fixed in writing before any result is seen.
Backtest, full record
We run the backtest and hand over every number, including the misses and the stations where we add nothing.
Live scoring
If you proceed, every issued forecast is archived as issued and scored as outcomes settle, for the life of the engagement.
What the verification report contains.
- CRPS skill against climatology, by location, month, and lead time
- Interval coverage at the stated probability levels
- Trigger and exceedance frequencies where an index is involved
- Null results, stated as plainly as the positive ones
- A point-in-time archive reference, so any number can be reproduced later
Updated as outcomes settle. Until then, the full records behind this page are shared with prospective clients under a short confidentiality note.
Where the work is being used.
The discipline above is not a proposal. It runs in production and has been exercised under protocol.
Our own daily practice: where an exchange lists a temperature futures contract, we compare our calibrated settlement forecast with the market price on every trading day and score it at settlement. The market price is a demanding benchmark, the consensus of professionals with money at risk. The daily record is archived.
Benchmarks run under protocols fixed in writing before any result is seen: locations, quantities, out-of-sample splits, benchmark, and score. Results are handed over in full, including the locations and months that show no improvement over climatology.
Probabilistic weather and demand-risk dashboards for utilities, residents, and trading desks, built with Plug and Play and supported by an NJEDA grant. Open on this site.
A freeze-risk index for fruit growers in the Northeast United States, in development with regional partners, built to the same scoring discipline.
Read the method.
Start with a benchmark on your own locations.
We score our calibrated forecasts against climatology on the stations, contracts, or fields that matter to you, and share the full record. The benchmark runs before any commercial commitment.
Or write to info@ilikallc.com
