Transparency

The Brier score: how to judge a football prediction model

By · Published · Updated 25 Sep 2026 · 3 min read

Here is an uncomfortable fact about football prediction: a model that picks the home team in every match is “right” about 44% of the time — 44.1% across the 17,495 finished matches in our dataset. So when a tipster advertises fifty-something percent accuracy, the honest question is never “is that high?” It is “compared to what, and how confident were the calls?”

The Brier score answers both, and it is the number we hold our own model to on the public track record.

What it measures

A prediction is not a pick. It is a set of probabilities — say 55% home, 25% draw, 20% away. The Brier score compares those probabilities with what actually happened, treated as 100% / 0% / 0%. Take each difference, square it, sum them:

  • Home win happens: (0.55 − 1)² + (0.25 − 0)² + (0.20 − 0)² = 0.305
  • Away win happens instead: (0.55 − 0)² + (0.25 − 0)² + (0.20 − 1)² = 1.005

Lower is better. Confident and right scores near zero; confident and wrong is punished hard; timid predictions land in the middle. Averaged over hundreds of matches, luck washes out and calibration remains.

Why it cannot be gamed the way accuracy can

This is the property that makes it worth the arithmetic. Brier is a proper scoring rule: it is minimised, in expectation, only by reporting the probabilities you actually believe. Shading a number to look bolder costs you, and so does hedging.

Accuracy has no such protection. A forecaster who only ever calls heavy favourites will post a fine accuracy figure and tell you nothing, because the picks carried no information you did not already have. Run the two measures against the same record and they can rank two forecasters in opposite orders — which is precisely why a site quoting only accuracy is showing you the flattering angle.

What good looks like in football

Numbers mean nothing without benchmarks. For 1X2 in top European leagues:

Predictor Typical Brier (3-outcome)
Always 33/33/33 ~0.667
Historical league frequencies (≈44/25/31) ~0.65
A competent statistical model 0.57–0.60
De-vigged closing betting odds ~0.55–0.56

Two lessons hide in that table. The gap between naive and excellent is small in absolute terms — football is genuinely hard to predict, and anyone claiming huge margins is selling something. And the closing market sits at the bottom: it is the strongest public predictor in existence, which is why most value bets are not.

Our own number, since we are asking

We do not quote a backtest here, because a backtest is a claim about the past that a reader cannot check. Our Brier score is the live one: every prediction we have published, scored on matches the model had never seen, updates continuously on the track record with its sample size beside it. What counts as a good value there is worked through in what is a good Brier score.

The two things a Brier score cannot tell you

It does not separate calibration from discrimination. A score can improve because probabilities got better sorted, or because they got better tuned, and the headline number will not say which. Splitting them apart requires decomposing the score or plotting a calibration curve — which is the check worth running before trusting any single figure.

It says nothing on its own about sample size. A Brier score over forty matches is close to noise; the same number over four thousand is evidence. Any published score without an n beside it is an incomplete claim, and that is true of ours as much as anyone’s.

So when you evaluate a prediction service — ours included — ask for the Brier score, the accuracy, the sample size and the calibration plot, and be suspicious of anyone offering only the one that flatters them. A fuller version of that argument, aimed at the industry rather than the metric, is in how accurate are football predictions.