The Ledger
Our own forecasts, dated when made and scored when they resolve. This is the only page here that can be wrong, which is the reason it exists.
Every other page on this site describes how forecasting should be judged. This one submits to it.
The rules
The question must resolve without argument
A date, a source of truth, and a binary outcome fixed in advance. "Will the incumbent win" resolves. "Will the economy improve" does not, and does not get entered.
The probability is recorded before the outcome is known
Timestamped at entry. An entry is never edited after the fact. If reasoning changes, a second dated entry is added and both stand.
Every entry made is an entry scored
An entry stands whatever the outcome, because a ledger that can be quietly pruned carries no information. The whole value of this page rests on it being impossible to edit backwards.
The comparison is the base rate, not zero
Scores are reported against what simply predicting the historical frequency would have achieved. Beating nothing is not skill.
Where the programme stands
Fifty-nine methods are in use across the seventeen fields, each one having cleared all five gates. Sixteen remain open, thirty are queued, and six were rejected or superseded on the evidence. Three results stand as constraints on everything else.
The forecast ledger is the newest layer of that work. Its rules were published before the first entry was made, which is the order that matters, because a standard set in advance is the only kind that cannot be adjusted to fit the result. Every forecast from here is dated when made, scored when it resolves, and stands either way. The record compounds from the first line.
What an entry looks like
The entry format. Every live entry carries the same four fields, and none of them can be edited afterwards: the date it was made, the question in resolvable form, the probability stated, and the outcome once it lands.
Calibration, and why it is drawn the way it is
Below is the instrument, running on adjustable inputs rather than on real entries. Drag the slider to see what different kinds of wrong look like. A perfectly calibrated forecaster sits on the diagonal. Everyone starts off it.
Reliability measures the gap between what was said and what happened. Resolution measures whether the forecasts discriminated between cases at all. A forecaster who says fifty per cent to everything scores perfectly on the first and zero on the second, and is useless.