What we said · What happened

Calibration: what we said, and what happened

A hit rate tells you how often we were right. It cannot tell you whether a 70% was really a 70%. This page can: every probability we have ever published, grouped by what we said, against how often it came true — with the sample size on every row, because a bucket of two calls proves nothing in either direction.

Every call, in one chart

0%0%20%20%40%40%60%60%80%80%100%100%perfect calibration We said 38% · it happened 35% · 159 calls We said 44% · it happened 47% · 455 calls We said 54% · it happened 54% · 1821 calls We said 64% · it happened 64% · 629 calls We said 75% · it happened 76% · 733 calls We said 84% · it happened 87% · 268 calls We said 93% · it happened 97% · 64 calls What we saidWhat happened
Stated probability against observed frequency per bucket, with sample sizes
We saidAverageHappenedGapCalls
30–40%38%35%−3159
40–50%44%47%+3455
50–60%54%54%01821
60–70%64%64%0629
70–80%75%76%+2733
80–90%84%87%+2268
90–100%93%97%+364

Across 4129 graded calls, the average gap between what we said and what happened is 1.1 points. The widest readable bucket is 30–40%: we said 38% and it happened 35% on 159 calls.

By market

Stated probability, observed frequency and sample size per bucket and market
We said1X2Double chanceOver/Under 2.5BTTS
30–40%38% → 35% · 159———
40–50%44% → 47% · 455———
50–60%55% → 61% · 241—55% → 53% · 74254% → 53% · 838
60–70%64% → 73% · 11869% → 74% · 3164% → 61% · 27463% → 61% · 206
70–80%75% → 83% · 5275% → 76% · 63873% → 79% · 3472% → 89% · 9
80–90%84% → 85% · 2784% → 87% · 23881% → 100% · 3—
90–100%94% → 100% · 193% → 97% · 63——

What we said → what happened · calls. Buckets under 20 calls are readings of nothing; they are printed so you can see they are thin.

How to read it

The diagonal is a model that means exactly what it says. A point above it is a bucket where things happened more often than we said — we were too modest. A point below is one where they happened less often — we were too confident. The distance from the line is the error in percentage points, and the table beside the chart prints it with its sign.

Why a 90% call must lose one time in ten

A model that said 90% and was right every time would not be a better model. It would be a worse one, hiding a 97% behind a 90%. Honest probabilities lose exactly as often as they promise to, which is why a screenshot of one 90% pick that failed proves nothing about the model that made it, ours or anyone else's. The only fair test is the one on this page: many calls at the same number, and how many came in.

Why the sample size is on every row

Because a bucket of two calls can read as 100% or 50% and mean nothing either way. The chart refuses to draw a point below 20 calls; the table keeps the small buckets and says how small they are. As the ledger grows the thin buckets fill in and the points appear on their own — the page is rebuilt from the settled rows several times a day.

What is here, and what is not

Four markets: match result, double chance, over/under 2.5 goals and both teams to score. Those are the markets whose stated probability the public ledger carries per call; over/under 1.5 and 3.5 and correct score are graded on the track record but not plotted here. Every model that has published on this site is included, retired ones too. The rows are the same ones you can download from the open dataset, so anyone can rebuild this chart and check it.

Frequently asked questions

What does it mean for a prediction to be calibrated?

That the number means what it says. If every call we made at 70% is collected together, about seven in ten of them should have come true. More is not better: a set of 70% calls that lands nine times in ten was underselling itself, and a set that lands half the time was overselling. Calibration is the property a hit rate cannot measure, because a wide market wins often whatever the model knows.

How is this chart built?

Every settled call on the public ledger contributes one point: the probability we published for the side we called, and whether that side came in. Calls are grouped into ten-point buckets by what we said; each bucket plots its average stated probability against the share that happened. The chart is recomputed from the append-only ledger at every build and never edited by hand.

Why does every bucket print its sample size?

Because a bucket of two calls can read as 50% or 100% and mean nothing either way. A point is only drawn once a bucket holds at least 20 calls; smaller buckets stay in the table with their count so you can see exactly how thin they are.

Which markets are included?

Match result (1X2), double chance, over/under 2.5 goals and both teams to score — the markets whose stated probability the public ledger carries. Over/under 1.5 and 3.5 and correct score are graded on the track record but not plotted here. Every model that has published on this site is included.