Model

Can you beat the bookmaker with maths?

Published

The internet’s answer to this question is remarkably consistent, and it comes almost entirely from people selling something. The recipe is always the same: build a model, compare its probabilities to the bookmaker’s, bet whenever yours is higher, profit arrives. The academic literature is more careful and occasionally reports genuine returns. The gap between those two bodies of writing is where most people lose money.

The honest answer is yes, it is possible, and it is much harder than the recipe implies — for a reason that has nothing to do with effort or cleverness.

What you are actually betting against

The price you see is not one bookmaker’s opinion. By kick-off it is the settled output of a market: professional traders, syndicates with better data than yours, and money from everyone who disagreed, all of it pushed into a single number that moves whenever anyone with conviction disagrees strongly enough to stake on it.

That process is good. In our own database of 16,628 matches carrying closing prices across nine leagues, the market’s shortest price — its own favourite — came in 53.1% of the time, and the win rate falls in an orderly line as prices lengthen, from 83.6% for favourites under 1.30 down to 36.7% for those at 2.50 and above. Whatever else you think of bookmakers, their ranking of outcomes is close to correctly ordered.

That is the bar. Not “predict football well” — predict football better than the aggregated judgement of everyone who has already bet on it, by enough to clear the margin they charge for the privilege.

The margin is the first hurdle, and it never sleeps

A bookmaker’s odds imply probabilities that sum to more than 100%. The excess is the overround, and it is charged on every bet whether you win or lose. Beating the market’s estimate is not enough; you have to beat it by more than the fee, every time, forever.

Worse, the fee is not applied evenly. It leans disproportionately on longer prices — the favourite-longshot bias — which is precisely where a model most often thinks it has found something. The outcomes that look most attractive to a naive edge detector are the ones carrying the heaviest charge.

If you have not worked through what a price actually claims once the margin is stripped, start there; the arithmetic changes how big an “edge” looks.

Why apparent edges are usually your error

Here is the trap that catches nearly everyone, and it is a statistical trap rather than a football one.

Your model produces a probability. The market produces a probability. They differ. There are two possible explanations: the market is wrong, or your model is. Betting the difference assumes the first without testing it — and for any model that isn’t calibrated, the second explanation is far more likely.

An uncalibrated model doesn’t fail evenly. It fails hardest exactly where it is most confident, which means its largest apparent edges are its largest errors. Selecting bets by “biggest disagreement with the market” is therefore a machine for finding your own worst estimates and staking money on them. This is why calibration is not an academic nicety: an uncalibrated model with a good hit rate is still a losing bet, because the hit rate never told you whether 70% meant 70%.

We know this concretely rather than theoretically. When we backtested our own model against closing odds and simulated betting every apparent edge at flat stakes, the result was a clear loss — the details are on our methodology page, published alongside the model’s other numbers. Our probabilities were good enough to beat naive guessing comfortably and not good enough to beat the closing market. Both halves of that sentence are the point.

What would actually constitute an edge

The honest checklist is short and the bar is high:

  1. Calibration before selection. Probabilities must be demonstrably honest across the full range before any of them are used to pick bets. Score them with a proper scoring rule, not a win rate.
  2. Measured against closing odds, not opening. Beating a Tuesday price that drifts by Saturday means you were early, not right. Closing line value is the benchmark that survives scrutiny.
  3. A sample that could tell skill from luck. Hundreds of bets minimum, and the arithmetic of small samples is unforgiving about anything less.
  4. Information the market lacks. This is the real one. Every public model runs on public data, and public data is already in the price. Genuine edges historically come from being faster, from data nobody else has, or from markets too small for professionals to bother with — not from a better fit to the same numbers everyone else can download.

Point four is why the famous success stories are famous. They are notable because they are rare, and most of the people capable of it end up working for a bookmaker or a fund rather than publishing tips.

Where that leaves us

We do not sell value bets, and we will not until calibration demonstrably clears the bar in simulation. Publishing probabilities is a claim we can support and grade in public; publishing “bet this, it’s value” would be a claim our own backtest currently contradicts.

The track record is where that gets settled — every prediction frozen before kick-off, graded afterwards, misses included. If the gap to the market ever closes, it will be visible there before it is announced anywhere.

So: can you beat the bookmaker with maths? In principle yes. In practice, the market is a hard, well-calibrated opponent, and the confident answer to “I found an edge” is almost always “you found a bug”. Treat any model’s output as a probability rather than a plan, and stake only what you can afford to lose.