Head-to-head records in football predictions: signal or noise?
By FootInsights · Published · Updated 26 Sep 2026 · 6 min read
Open almost any preview of a football match and you will find the same paragraph: one side has won four of the last five meetings, so they are the pick. Head-to-head records in football predictions are the most quoted evidence in the entire genre and, measured honestly, one of the weakest. The problem is not that past meetings contain no information. It is that they contain so little of it, so noisily, that a rating model built on thousands of matches has usually priced the fixture better before a single head-to-head result is consulted.
This post puts arithmetic on that claim: how often a lopsided H2H record appears by pure chance, why the sample can never grow fast enough to settle the question, and the narrow set of cases where pairing history genuinely tells you something.
Why head-to-head records feel more informative than they are
Two forces make H2H persuasive out of proportion to its content.
The first is that it comes with a story. “They just don’t like playing at that ground” is a causal explanation, and human reasoning accepts a causal explanation far more readily than it accepts variance. The second is selection: nobody writes a preview about the pairing that has split its last six meetings evenly. You only ever see the streaks, because the streaks are what got picked out and pointed at.
That combination — a vivid pattern plus an available explanation — is the standard recipe for reading meaning into noise. The way out is to ask a boring question first: how often would this pattern show up if nothing special were going on at all?
The arithmetic of five meetings
Take a fixture where one side is genuinely the better team but not overwhelmingly so — say their true probability of winning any single meeting is 40%, with the rest split between draws and defeats. Now look at what a run of five meetings does.
| Record over five meetings | Chance of it happening at a true 40% win rate |
|---|---|
| Wins all five | 1.0% |
| Wins four or more | 8.7% |
| Wins three or more | 31.7% |
Read the bottom row slowly. A team with a modest edge wins the majority of a five-meeting sequence about a third of the time, and that is exactly the record a preview would describe as dominance. In a twenty-team division there are 190 possible pairings; if each had met five times, dozens of “dominant” head-to-head records would exist purely as a by-product of shuffling.
The row above it matters too. Four wins in five — the record that reliably gets called a hoodoo — arrives roughly once in eleven attempts without any pairing-specific effect whatsoever.
And the sample cannot be fixed by waiting. League opponents meet twice a season. Five meetings is two and a half seasons of history; a hundred meetings, which is roughly what you would need to separate a true 55% from a true 45% with any confidence, is fifty years. Long before you accumulate that, the teams involved have stopped existing in any meaningful sense.
Two teams are never the same two teams twice
The statistical problem is compounded by a football one. A meeting from three seasons ago was played by a squad that has since turned over substantially, very possibly under a different manager, with different tactical instructions and a different set of players available on the day.
Head-to-head history treats all of those as one continuous entity called “the fixture”. A rating model does the opposite: it applies time decay, so recent evidence counts heavily and old evidence fades toward irrelevance. Those two approaches cannot both be right, and only one of them survives contact with how squads actually change.
There is a subtler version of the same trap. Because opponents meet home and away, a six-meeting record is really two three-meeting records, one at each ground. Any H2H figure quoted without splitting by venue is mixing two different questions — and three matches at a venue is a sample that supports essentially no conclusion at all.
The ratings already contain those meetings
This is the argument that does most of the work. When a Dixon-Coles style model prices a fixture, each team’s attack and defence ratings have been estimated from every match that team has played in the fitted window — hundreds of matches per side, tens of thousands across the dataset. The effective evidence behind the model’s view of a fixture is three orders of magnitude larger than the head-to-head sample.
So head-to-head data can only add value if it contains something the ratings do not: a genuine interaction between these two specific teams, over and above their individual strengths. That is a real statistical concept, and it is also the hardest kind of effect to estimate. Interaction terms need far more data than main effects, and the data available for any single pairing is the handful of meetings we just showed to be dominated by noise.
Adding H2H to a rating model, then, usually does one of two things. Either it re-states information the ratings already carry — both teams’ recent results are in there — which double counts. Or it fits pairing-level noise, which makes predictions worse in exactly the confident, specific way that is most expensive.
When head-to-head data does carry information
Not never. The honest cases have a shape in common: the effect is real, but it is better modelled as a general variable estimated on the full dataset than as a memory attached to one pairing.
- Stylistic mismatches. It is plausible that a deep-defending side systematically frustrates a possession-heavy one. If true, that is a property of playing styles, measurable across every match involving those styles — not a property of these two clubs.
- Venue effects. Unusual pitch dimensions, altitude, or a genuinely hostile ground affect every visitor, so they belong in a home-advantage term fitted on all matches at that venue, not in one fixture’s history.
- Competition context. Cup ties, derbies and matches with qualification stakes may well behave differently. That is a context variable, and it applies league-wide.
The pattern is consistent: whenever a head-to-head story turns out to be true, it generalises — and once it generalises, you can estimate it on thousands of matches instead of five, which is the entire ballgame.
How to test a head-to-head claim before trusting it
A short checklist, usable on anyone’s preview including ours:
- Count the meetings and split them by venue. If either half has fewer than about ten matches, stop; there is nothing to interpret.
- Ask what the baseline said. Would a plain ratings model have predicted those results anyway? An H2H record that merely agrees with team strength adds nothing — it has to beat the baseline to count.
- Score it probabilistically, across many pairings. One striking fixture proves nothing. Convert the claim into probabilities, apply it to every pairing in a league, and measure it with a Brier score against the baseline. If it does not lower the score, it is decoration.
- Remember why you are looking at this fixture. You noticed it because it looked striking, which means it is pre-selected for extremity — the one condition under which regression to the mean is guaranteed to be waiting.
How we handle it
The FootInsights model reads years of results, fixtures and prices — how that becomes a probability is described on the how our predictions work page. Head-to-head history is not among those inputs. Past meetings appear on match pages as context for readers who want it, not as a lever that moves the probabilities.
That is a modelling judgement, not a certainty, and the useful thing about publishing it is that it stays checkable. Every prediction is graded after full time and kept, so the consequences of choices like this one show up in the public track record rather than in a claim about how sophisticated the model is. If a head-to-head term ever earns its place, it will be because it improves those numbers — not because it makes a better sentence in a preview.