Transparency

How accurate are football predictions? The honest ceiling

Published

How accurate are football predictions? It is the first question anyone asks of a prediction site, and almost every answer you will find is either useless or dishonest. Useless, because “accuracy” is never defined. Dishonest, because the advertised figure was produced by choosing which matches to count. The interesting answer is not a number scraped off someone’s homepage — it is the ceiling: the accuracy a perfect football prediction model would reach, given that football is a genuinely uncertain game. That ceiling is lower than most people expect, and knowing roughly where it sits turns a vague question into a working lie detector.

What “accuracy” means before it means anything

Accuracy is a hit rate: the share of matches where the outcome the model called as most likely actually happened. That single sentence hides three levers, and every one of them moves the number without any change in forecasting skill.

The market changes the ceiling. A 1X2 call has three possible outcomes; an Over/Under 2.5 call has two. Guessing at random gets you 33% on the first and 50% on the second. A 62% hit rate on Over/Under is unremarkable. The same 62% on 1X2 would be extraordinary. Headline accuracy figures rarely say which market they refer to — and quoting a double-chance hit rate next to a 1X2 competitor’s is not a comparison, it is a category error.

The fixture list changes the ceiling. A model that only ever calls matches where a dominant side hosts a struggling one will post a high hit rate with no skill at all. Anyone can predict a top club at home against the bottom side. Skill lives in the close matches, and close matches are exactly the ones a selective publisher quietly omits.

The reporting window changes the number. Over a few dozen matches, accuracy swings wildly for reasons that have nothing to do with the model. A run of 20 correct calls happens to competent and incompetent forecasters alike.

So before comparing two accuracy claims, three things must match: the market, the set of matches, and the sample size. They almost never do.

The ceiling: what a flawless model would score

Here is the part that reframes the whole question. Imagine a model with perfect knowledge of the true probabilities of every match — not a good model, the theoretical maximum. It still cannot exceed a hard limit, because football outcomes are random draws from those probabilities. If the true probability of a home win is 70%, the home side loses or draws 30% of the time, and no forecaster who ever lived can do anything about it.

For a perfectly calibrated model that always calls its most likely outcome, expected accuracy equals the average of the highest probability across the fixture list. Work it through with an illustrative spread of match types from a typical top division — these numbers are hypothetical, chosen to show the shape of the calculation:

Match type Home / Draw / Away Top probability Share of fixtures
Heavy favourite 70% / 19% / 11% 70% 15%
Clear favourite 55% / 25% / 20% 55% 30%
Slight favourite 44% / 28% / 28% 44% 30%
Coin flip 36% / 29% / 35% 36% 25%

Expected accuracy is the weighted average of that third column:

0.15×0.70 + 0.30×0.55 + 0.30×0.44 + 0.25×0.36 = 0.49

Roughly 49%. A model that knows the truth exactly, calling every match in a full league season, gets about half of them wrong. Shift the fixture mix toward stronger favourites and the ceiling rises into the mid-50s; shift it toward tight matches and it falls toward 40%. That range — call it 50% to 55% on 1X2 across a complete fixture list — is the neighbourhood in which honest full-coverage accuracy figures live.

This is not pessimism about models. It is a property of the sport. A single football match delivers one bit of noisy evidence, and the favourite loses often enough to keep the league interesting.

Why draws hold the ceiling down

Draws are the structural reason football accuracy caps lower than other sports. In most top divisions, roughly a quarter of matches end level — but a draw is almost never the single most likely outcome. Even in the most draw-prone leagues, draw probability tops out around 30% to 35%, which means it rarely wins the three-way comparison against a home or away price.

A model calling the most likely outcome therefore predicts a draw almost never, and concedes a quarter of the fixture list before kick-off. This alone puts a lid of about 75% on 1X2 accuracy, and the ordinary uncertainty between home and away wins takes it down from there to the 50s. It also explains why accuracy varies by competition: a defensive, draw-heavy league mechanically produces lower hit rates than a high-variance, favourite-friendly one, with the same model doing the same work.

What a 70% accuracy claim actually tells you

Apply the ceiling as a filter. When a site advertises 70%, 77% or 85% accuracy on match outcomes, only a few explanations exist, and none of them is “better model”:

  1. It is a different market. Over/Under, BTTS or double chance calls clear 60% routinely, because they have two outcomes and the model can favour the common one.
  2. The fixtures were selected. Only heavy favourites were counted, or only the picks that landed.
  3. The sample is small. Twenty or thirty matches produce spectacular hit rates in both directions.
  4. The record is not complete. Losing calls were removed, or the history starts at a flattering moment.
  5. The number is invented. In a niche where nobody publishes their raw predictions, this costs nothing.

None of these require the site to be sophisticated. They require only that the accuracy figure be unauditable — which is why the useful question is never “what is your accuracy?” but “where is your complete, timestamped, unedited history?”

Measure the probabilities, not the hit rate

The deeper problem is that accuracy throws away most of the information in a forecast. A model saying 51% and a model saying 95% both get scored identically when the favourite wins, and identically when it loses — yet one of them is telling you far more, and taking far more risk in doing so.

Proper scoring rules fix this by grading the stated probability against the outcome, which is what a Brier score does. Under that scoring, confident predictions that come off are rewarded, confident predictions that fail are punished hard, and a forecaster who hedges everything toward the base rate cannot hide. Two consequences follow: a model can be judged in hundreds of matches rather than thousands of bets, and inflating the number by picking easy fixtures stops working, because easy fixtures earn small scores for everyone.

That is why our published record leads with scored probabilities rather than a hit-rate banner. Every prediction we make is stored with its model version and settled automatically, wins and misses alike, on the track record page — including the stretches where the model reads a league badly. The methodology behind the numbers is public for the same reason: an accuracy figure you cannot reconstruct is a marketing asset, not evidence.

So how accurate are football predictions?

On full-coverage 1X2 calls, a genuinely excellent model lands somewhere in the low-to-mid 50s, and a perfect one would not do dramatically better. On two-outcome markets, higher. On a hand-picked list of one-sided fixtures, higher still, and meaninglessly so.

The number itself is nearly useless without the fixture list, the market and the sample size attached — and once those are attached, the honest ceiling is close enough to a coin flip that any prediction service quoting a figure far above it is telling you something about its marketing rather than its model.