Model

How injuries affect football predictions: putting a number on absence

Published

A star striker is ruled out and the previews rewrite themselves within the hour. The question almost nobody asks is the only one that matters: how injuries affect football predictions in numbers rather than adjectives. Not whether the absence is bad news — it obviously is — but how many percentage points it should actually move a match probability, and whether that movement has already been priced by the time you read about it.

The honest answer is smaller than the coverage implies, and it hinges on a variable the headlines almost never mention: who plays instead.

An absence is a goal-rate question

Football probabilities are built from goal expectancies. A model estimates how many goals each side should score, then converts that pair of numbers into 1X2, Over/Under and both-teams-to-score prices. So the only way an absence can change a prediction is by changing a goal expectancy — which means the useful question is how much of a team’s scoring or conceding rate travels with one player.

Work it through with hypothetical but realistic numbers. Take a side whose model expects roughly 1.6 goals per match. Their leading scorer accounts for a quarter of the team’s goals across a full season — a high share, the kind that gets a player called indispensable. Naively, removing him removes 0.40 goals of expectancy.

That figure is wrong in two directions at once, and both corrections push it down:

  • He is replaced, not deleted. The team still fields eleven players. If his deputy scores at 60% of his rate in the same role, the marginal loss is 40% of 0.40, or 0.16 goals.
  • Goals are not created solo. A striker’s finishing sits on top of chance creation by teammates who are all still playing. Some of the production attributed to him is really the team’s, and it stays when he leaves.

So a genuinely important attacker is often worth something like 0.10–0.20 goals of expectancy, not 0.40. That is the number to carry into the probability calculation — and the largest term in it, by some distance, is replacement quality rather than the absentee’s own quality.

How injuries affect football predictions in percentage points

Suppose our hypothetical home side sits at 1.60 expected goals against an away side at 1.10. Feed those into a scoreline grid and you get roughly 49% home / 25% draw / 26% away.

Now remove 0.15 from the home expectancy, to 1.45, and rebuild the grid. The home win drifts to about 45%, the draw ticks up towards 26%, and the away win to roughly 29%.

A four-point move. Real, worth having, and completely unlike the story the previews tell. It does not turn a favourite into an underdog; it turns a solid favourite into a slightly less solid one. Two structural facts about football force this result:

  1. The sport is low-scoring. When a match hinges on one or two goals, outcome probabilities are compressed towards each other and it takes a large shift in expectancy to reorder them.
  2. One player is a fraction of a team. Even the most concentrated attacking side spreads production across a squad, and the bench exists.

The corollary is that an absence severe enough to swing a match probability by ten points or more is genuinely rare. When somebody claims one, the burden of proof is on the claim.

Which absences are actually worth adjusting for

Not all absences are the same size, and the differences are systematic enough to be usable.

Concentration of production. A team whose goals and chance creation funnel through one or two players is more exposed than a team that spreads them evenly, even if the individual is less celebrated. Look at the distribution, not the name.

Replacement quality and squad depth. This is the dominant term. A deep squad replaces a starter with someone close to the same level, so the marginal effect approaches zero. A thin squad replaces him with a fringe player or a teenager, and the loss is close to the full difference. Two clubs losing “their best midfielder” can face entirely different corrections.

Clusters beat individuals. Three absences in the same unit do more than three times the damage of one, because they force a structural change rather than a substitution. Individual player-value estimates are additive by construction and quietly understate this case.

Goalkeepers are their own category. The position is the least interchangeable on the pitch, and reliable public estimates of the gap between a first and third choice are hard to find — which argues for widening your uncertainty rather than sharpening your number.

Suspensions are not injuries. A suspension is known days ahead, publicly and unambiguously; an injury may be doubtful until an hour before kick-off. Those two reach the market on completely different schedules, and that schedule is most of what determines whether the news is still worth anything.

The market has usually moved before you have

Team news is public information, and public information in a liquid betting market has a short half-life. A significant absence confirmed in a press conference is in the price within minutes; official lineups arrive about an hour before kick-off and are absorbed almost immediately.

That reframes the whole exercise. Knowing a key player is out is not an edge — everyone knows. An edge would require valuing that absence more accurately than the market does, and the market’s collective view of it is embedded in the closing price, which research and our own backtesting both find hard to beat (why most “value bets” aren’t).

The practical consequence: if a price has already drifted on team news, the drift is the market’s estimate of the effect. Betting into it because you have just read the headline is paying for information you do not have first.

What our model does with injuries, and what it doesn’t

Our published methodology is a Dixon-Coles rating model with time decay plus an Elo signal, fitted on fixtures, results and standings. It carries no injury, suspension or lineup input — that limitation is listed openly on the how it works page rather than buried, because it changes how the numbers should be read.

Two consequences follow, and they point in opposite directions.

The first is a genuine weakness. A fixture with a major, freshly confirmed absence is priced by our model as though the squad were intact, so the model’s probability for that specific match is worse than one a careful human with team news could produce.

The second is less obvious and partly offsetting. Ratings are estimated from results with time decay, so a long-term absence is absorbed automatically. Once a team has played several matches without a player, those results — worse, if the absence truly hurt — flow into the attack and defence ratings on their own. The model never learns the player’s name and never needs to. It is short-notice absences, where no results have yet been played without that player, that sit in the blind spot.

A checklist for reading team news

  1. Name the replacement, not just the absentee. If you cannot say who plays instead and roughly how good they are, you cannot size the effect at all.
  2. Convert to goals before converting to opinions. Estimate the change in expected goals first, then let a scoreline grid turn it into probabilities. Reasoning directly from “he’s out” to “back the other side” skips the only step that contains any arithmetic.
  3. Check whether it is short-notice or long-standing. An absence several matches old is already inside a results-based rating. Treating it as fresh news double counts it.
  4. Look at the price before the headline. If the line has already moved, you are reading old information.
  5. Expect a small number. Single absences are usually worth low single digits in percentage points. Clusters and goalkeepers earn more; almost nothing earns a double-digit swing.

None of this is a claim that team news is worthless. It is a claim that its value is measurable, usually modest, and frequently already spent by the time it reaches you. We publish the model’s inputs, its limits and every graded prediction on the track record precisely so that gaps like this one stay visible — if adding lineup data ever improves those numbers, that is the evidence that will show it.