Why football prediction models disagree, and how to read the gap
By FootInsights · Published · 6 min read
Open two prediction sites on the same fixture and you will often find two different numbers: one gives the home side a 52% chance, the other 38%. The instinct is to ask which one is right. That question has no answer available before kick-off, and it is not the useful one. Understanding why football prediction models disagree — and how far apart two numbers have to be before the disagreement means anything — turns a confusing pair of figures into a genuine piece of information about the match.
The short version: a small gap is noise, a large gap is a statement about the fixture rather than about the modellers, and the only thing that can arbitrate between two models over time is a public record that can be scored.
Why football prediction models disagree in the first place
A prediction model is a chain of decisions, and almost none of them are forced. A model has to pick a way of rating teams — a rating system that updates after each result, or separate attack and defence strengths fitted to a goals model. It has to choose how goals are distributed, and whether to apply a correction for the low-scoring scorelines that a plain Poisson gets wrong. It has to decide how much history counts and how fast old matches fade. It has to decide whether draws are modelled directly or fall out of the goal distribution. And it has to decide what it will refuse to touch at all: team news, motivation, travel, weather.
Each of those choices moves a probability by a couple of points. Chained together, they move it by ten or more. Two models built by competent people, on the same data, with no errors in either, will produce visibly different numbers for the same match as a matter of routine. Disagreement is the normal state, not a symptom that one of them is broken.
There is a less flattering source too, and it is worth naming: many sites do not run a model at all. They resell a feed, or convert bookmaker odds into a “prediction” and present it as independent analysis. Two sites agreeing tells you nothing if both are downstream of the same price.
How big a disagreement has to be before it matters
Percentage points are hard to feel, so convert them into prices and compare them against something with a known scale — the bookmaker’s margin.
Take a three-way market quoted at 2.05 home, 3.40 draw, 3.75 away. The raw implied probabilities are 48.8%, 29.4% and 26.7%, and they sum to about 105%. That surplus is the margin, and removing it properly leaves the home side at roughly 46.5%. The book is charging a bit over two percentage points on that outcome.
Now put the two models against that scale:
- A gap of two or three points — say 46% against 49% — is smaller than the margin the bookmaker charges to quote the price at all. It is also smaller than the honest uncertainty around any single model’s output. Treat it as agreement.
- A gap of five to eight points is worth a look. Usually one input explains all of it: one model weights recent matches much more aggressively, or one has a materially different view of a promoted or newly rebuilt side.
- A gap of ten points or more, as in the 52%-versus-38% example, is structural. The two models are not disagreeing about a detail; they disagree about what kind of match this is.
That last case is the interesting one, and the usual reading of it is backwards. A large gap is not an invitation to pick the number you prefer. It is evidence that the fixture is genuinely hard to price — which is a reason to commit less to it, not more.
Disagreement measures the match, not the modellers
Line up several independently built models and the spread between them behaves like an uncertainty estimate for the fixture itself. Ordinary league matches between two settled sides with plenty of history tend to produce tight clusters. The scatter widens on exactly the fixtures you would expect: cup ties with rotated squads, early rounds after heavy squad turnover, sides with barely any comparable matches on record, matches where one team’s motivation is in doubt.
So when the models spread out, the finding is “this match is hard”, not “these modellers are bad”. That reframing is worth the price of admission on its own, because it stops the search for the one oracle who has it figured out.
Averaging usually beats picking a side
If two models are roughly comparable in quality and genuinely built differently, the average of their probabilities will usually score better on a proper scoring rule than either model alone. Their errors are partly independent, so averaging cancels some of each.
Notice what happens in the running example: 52% and 38% average to 45%, which lands close to the margin-free market price of 46.5%. That is not a coincidence — it is the same effect that makes the market itself hard to beat, working on a sample of two.
Three conditions attach to this, and they matter:
- Average the probabilities, not the picks. Two models that both “predict a home win” while disagreeing 52% to 38% are not agreeing about anything useful, and a majority vote on picks throws away the only information they gave you.
- Only average sources that are actually independent. Blending three sites that all derive from the same odds feed does not reduce error; it just triples the confidence you feel while making the same mistake.
- Averaging costs you something if one model is clearly better. Blending a strong model with a weak one drags the good number toward the bad one. That is only worth doing when you have no reliable way to tell which is which — which brings us to the part that does settle it.
The tiebreak is the record, not the match
No single result can adjudicate between 52% and 38%. If the home side wins, both models said it was possible; the one that said 38% was not thereby refuted. Probabilities are only judged in bulk, over hundreds of matches, with a scoring rule that rewards being right and being honest at the same time — which is what the Brier score is for.
That gives you a tiebreak that works before you ever look at a specific fixture. Prefer, consistently, the model that:
- publishes probabilities rather than picks, so it can be scored at all;
- publishes them for every fixture it covers, not a selection assembled afterwards;
- timestamps them before kick-off, so nothing can be revised once results exist;
- and attaches a proper scoring rule to the whole set, in public.
A model that is confident and unscoreable should lose to a model that is hedged and scoreable, every time. That is the standard we hold ourselves to: our probabilities are dated, complete, and scored on the public track record, where the numbers are the site’s own results rather than a claim about them.
What to do with a real disagreement
When the gap survives the noise test, work through it in this order. Find the input that explains it — usually one assumption is doing all the work. Ask whether the divergent model knows something the other does not, or merely assumes something the other declines to assume. Check both against the margin-free market price: a model that disagrees with a rival and with the market is making a much bigger claim than one that disagrees with a rival while sitting where the market sits. Then size your interest to your confidence, which in a genuine disagreement should be lower than usual.
Do this for a few weeks, logging both models’ numbers and the outcome, and you stop needing anyone’s opinion about which site is better. You will have built the small version of the record that should have been published in the first place.