Anteproof · Calibration

Calibration record

How much to trust a number before you use it. Every forecast we have ever issued is scored against what actually happened; this page is that scorecard, by question family, with nothing left out.

The launch gate

Four statistical criteria, all required, evaluated on point-in-time backtests ("pastcasts") of the forecaster version that is live. Passing the gate is what lets us market a family; failing it is published just the same.

Loading…

Reliability

When we say 30%, does it happen about 30% of the time? Each point is one probability bin: what we said on the horizontal axis, what happened on the vertical. Perfect calibration sits on the diagonal. Point size shows how many forecasts are in the bin.

Horizon

By question family

Brier score is the mean squared error of a probability: 0 is perfect. Two baselines: skill is against always saying 50% (Brier 0.25); vs base rate is against always saying the family's own observed frequency, a harder bar and the one that matters when a family's outcomes are lopsided. Families with fewer than 30 resolved forecasts are shown but not judged.

Pastcast pool

Loading…

Live track forecasts issued in real time, all versions

Loading…

Every gate evaluation

The complete history, verbatim. Sample size restarts whenever the question scheme or the forecaster version changes, because pooling regimes was measured to flatter the numbers.

How to read this

Pastcast. A forecast issued as if on a past date, with retrieval, thresholds and the model's own knowledge clamped to what was available then, then scored on the real outcome. It is how a young forecaster earns a sample size before it has lived long enough to accumulate one. Every pastcast is stored with the same append-only discipline as a live forecast.

Live track. Forecasts issued in real time and timestamped before the outcome was knowable. Small today; it grows every week and will eventually replace the pastcast pool as the record that matters.

The criteria. At least 150 resolved forecasts; a Brier skill score of at least 0.10 over the always-50% baseline; that skill statistically significant under a block bootstrap (blocks by metric and date, so clustered questions do not count as independent evidence); and two bias tests, calibration-in-the-large and Spiegelhalter's conditional test, both inside |z| < 2.

Two baselines. Always-50% is the classical Brier baseline and the one the gate uses. Always-base-rate ("climatology") is stricter: it credits nothing for knowing that a family resolves YES 46% of the time, only for telling questions apart. We compute it on the pool's own observed base rate, which flatters the baseline and understates our skill, the conservative direction. A family that beats 50% but not its base rate has learned the base rate and little else.

Horizon. Days from a forecast's as-of date (its issue date, for a live forecast) to resolution, bucketed. Skill at seven days says nothing about skill at ninety, so the selector above scopes every table on this page; buckets with no resolved forecasts are greyed. The page opens on the horizon the launch gate judges, because the archive also holds pastcasts at other horizons and pooling them was measured to flatter the headline; "all" pools them on purpose.

What this does not claim. Calibration on one family does not transfer to another, and a passing gate on backtests is not a passing gate on the live track. That is why both are shown, side by side, and why the live numbers carry a flag until they are large enough to mean something.