Model

How to backtest a football prediction model without fooling yourself

By · Published · Updated 25 Sep 2026 · 5 min read

Every football model looks brilliant in a backtest. That is not a compliment to the models — it is an indictment of most backtests. To backtest a football prediction model is to replay history and ask how the model would have done, and the exercise is only as honest as its rules: the model may use nothing it would not have known before kick-off, it must be judged on data it never trained on, and any betting simulation must pay real-world prices. Break any of those rules and the backtest stops measuring the model and starts measuring the leak. This post walks through the three ways backtests silently go wrong, how to structure one that can actually fail, and the reason even a clean backtest is the beginning of the evidence rather than the end of it.

What it means to backtest a football prediction model

A backtest is a simulation with a strict information diet. For every historical match, the model receives only what existed before that match was played — earlier results, earlier ratings, pre-match odds — and produces a probability. The probabilities are then scored against what actually happened, ideally with a proper scoring rule such as the Brier score rather than a simple hit rate, because a probability model’s job is to be calibrated, not merely to pick winners.

Done honestly, this is the cheapest way to learn whether a modelling idea holds up across thousands of matches before a single real prediction is published. Done carelessly, it is a machine for manufacturing confidence. The difference sits in three failure modes.

Failure mode 1: look-ahead leakage

Leakage is any route by which information from a match’s future sneaks into its prediction, and it is far easier to introduce than to spot.

The blatant version is training a model on seasons 1–5 and then “testing” it on matches inside those same seasons: the model has already seen the answers. The subtle versions are worse because they pass a casual code review. A team-strength rating computed over a full season and then used to predict matches from that season’s opening months has leaked — those early predictions are using information from games not yet played. A league-average home advantage estimated across the whole dataset and applied to every match in it leaks the same way. Even data hygiene can leak: if you drop teams that were later relegated because their data is patchy, the surviving sample is quietly biased toward stability.

The discipline that prevents all of these is the walk-forward backtest: process matches in strict chronological order, and before each one, rebuild every input — ratings, averages, adjustments — from only the matches that preceded it. It is slower to run and more tedious to code, and it is the only version whose results deserve trust.

Failure mode 2: overfitting to a finished past

Give a model enough parameters, or a strategy enough filters, and either will fit historical noise perfectly. The classic tell is the rule stack that grows one condition at a time — only home favourites, only after a loss, only in months with an “r” — with each addition improving the historical result and eroding the logical one. Football data is finite; with enough searching, patterns that mean nothing are guaranteed to appear.

Two habits keep this contained. First, insist that every input earns its place with a footballing reason before you test it, not a results-based justification after. Second, hold data back. Fit the model on one span of seasons, freeze it, and evaluate on a later span it has never touched — the out-of-sample test. A model that shines in-sample and collapses out-of-sample was memorising, not learning, and the out-of-sample number is the only one worth quoting. Held-out data also needs to be large: as with any betting record, a few hundred matches of sample size leaves a noise band wide enough to hide both real edges and real flaws.

Failure mode 3: fantasy odds

The first two failure modes corrupt a model’s probabilities; this one corrupts the profit simulation stacked on top of them. Backtested betting returns are routinely inflated by prices no real bettor could have taken: best-of-market odds when only one book offered them, closing odds for bets the strategy would have placed hours earlier, and — most commonly — ignoring the bookmaker’s margin entirely. A margin of a few percent per bet, compounded over a season of wagers, is the difference between a strategy that backtests as profitable and one that actually is.

The honest simulation bets at a single realistic price per match, subtracts the margin, and then asks a harder question: does the edge survive against closing odds, the market’s most informed price? Run that experiment on almost any apparent edge and a few seasons of closing odds turn a promising signal into a losing strategy — which is precisely the kind of thing a backtest is for. A backtest that cannot possibly say “no” is not a test.

Why a clean backtest still isn’t a track record

Suppose the backtest is immaculate: walk-forward, out-of-sample, real prices. It remains a claim about the past made by the person selling you the future, and it carries two irreducible weaknesses.

The first is selection. You see the backtest that survived, never the dozen quiet variants that didn’t — and the survivor of a dozen searches is measurably lucky even if every individual test was honest. The second is drift: football and its markets keep moving, and a relationship that held for five historical seasons is under no obligation to hold next season, when the market has adapted and the model’s assumptions have aged.

The only evidence immune to both problems is a prediction published before the match, timestamped, scored after it, and never edited — accumulated in public, misses included. That is why a live, append-only track record outranks any backtest, ours or anyone’s: it cannot be re-run until it flatters, and it is scored on matches nobody had seen when the prediction was made. A backtest earns a model the right to make live predictions; only the live predictions themselves earn it any trust.

A checklist before you believe any backtest

Whether the backtest is yours or a stranger’s sales page, the same five questions decide its worth:

  1. Chronology — was every input rebuilt walk-forward, using only pre-match information?
  2. Held-out data — are the quoted results from seasons the model never trained on?
  3. Real prices — one realistic price per bet, margin included, no best-of-market hindsight?
  4. Sample size — enough matches that the result sits outside the luck band?
  5. Survivorship — how many variants were tried before this one, and where are their results?

Most published backtests fail at least three. The ones that pass all five have earned exactly one thing: the chance to be tested for real, one publicly logged prediction at a time.