How football prediction algorithms actually work
By FootInsights · Published · Updated 25 Sep 2026 · 4 min read
Search for football predictions and you will find two kinds of site: ones that shout “92% win rate” without ever showing their working, and ones that quietly do statistics. This article explains what the statistical kind actually does — including ours, whose entire track record is public.
Step 1: turn results into ratings
Every serious model starts by compressing a team’s history into a small number of ratings. The classic approach gives each club two: an attack rating for how well it scores, and a defence rating for how well it prevents, both estimated from a few seasons of results.
Two details matter enormously.
Home advantage is real and has to be modelled explicitly. Across the 17,495 matches in our dataset, home teams won 44.1% and away teams 30.5%. Ignore that asymmetry and every prediction skews. Its size also varies by competition far more than people assume — from a 7.8-point gap in Serie A to 21.1 in Brazil’s Série A — which is why a single global home term is not good enough.
Recency weighting. A result from two seasons ago says much less about a squad than last month’s. Good models decay the weight of old matches, usually exponentially, so ratings track what a team is now rather than what it was.
Step 2: turn ratings into goals
Ratings become predictions through a goal model. The workhorse is the Poisson distribution: given how many goals a team is expected to score against this opponent at this venue, it returns the probability of exactly 0, 1, 2, 3 and so on. The full explainer is here, including how well it actually fits our data.
Combine the home side’s distribution with the away side’s and you have a grid of every plausible scoreline with a probability attached: 1-0 at around 10%, 1-1 at around 12%, and so on down.
Step 3: turn goals into markets
Once the grid exists, every market is arithmetic on it:
- Match result — sum the cells where home goals exceed away, where they are level, and where away exceeds home.
- Over/Under 2.5 — sum the cells totalling three or more.
- Both teams to score — sum the cells where neither side is on zero.
- Asian handicap — sum by margin of victory instead of by winner.
- Double chance — add two of the three result probabilities together.
This is why one well-built goal model prices many markets consistently, while sites that predict each market separately routinely contradict themselves. Checking whether a site’s markets reconcile is the fastest test of whether there is a model underneath at all.
What the better models add
Plain Poisson treats the two teams’ goal counts as independent, and the Dixon-Coles correction of 1997 adjusts the low-scoring corner of the grid where that assumption strains. The mechanics are in our Dixon-Coles write-up.
Beyond that, models differ in what they add rather than in the Poisson step, which is standard machinery: Elo-style ratings alongside the goal model, the market price itself, expected goals where the data is available and trustworthy, or machine learning over a wider feature set. More sophistication is not automatically better — it has to survive a backtest that the simpler version does not.
What no model can do
An honest list, because this niche rarely offers one.
- Beat the closing market consistently without private information. The market aggregates thousands of sharp opinions, and a model built on public data should expect to sit a little behind it. Where ours sits is measured on the track record rather than claimed, and the reasoning is in why most value bets are not.
- Price a promoted club well. With no top-flight history, any model starts it near the league average and learns on the fly — a limitation we document in what our model knows about promoted clubs.
- See one-off events. A third-minute red card, a cup hangover, a manager sacked overnight: statistical models see none of it until it shows up in results.
- Tell you it is well calibrated. That is measured, not claimed — with a Brier score and a calibration curve, over a sample large enough to mean something.
How to judge any prediction site, including this one
Three questions. Do they show probabilities, or just picks? Do they publish every past prediction, graded, misses included? Do they say what a prediction is and how it is graded?
If the answer to any of them is no, you are reading marketing rather than modelling. A longer version of that test, applied to the industry, is in how accurate are football predictions. We built FootInsights to pass it by construction: predictions are recorded before kick-off and graded automatically, and the track record only ever grows.