Why football stats differ between sites (and what it does to predictions)
By FootInsights · Published · Updated 25 Sep 2026 · 6 min read
Look up the same match on two football statistics sites and you will often find two different matches. Shots differ by one or two. Possession differs by a percentage point. Expected goals can differ by a tenth of a goal or more. Nobody is lying, and neither site is broken. Football stats differ between sites because a “stat” is not a measurement in the way a scoreline is — it is the output of a collection process and a definition, and both vary by provider. That distinction matters enormously if you are building or judging a prediction model, because a model inherits every quirk of the data it was fitted on.
Where the numbers actually come from
Almost nothing beyond the final score is counted by the competition itself. Match statistics are produced by data companies that watch every match and tag events: shot, pass, tackle, duel, foul, corner. Some of that tagging is done by trained human operators, some by automated tracking, most by a mix of the two with a review pass on top.
Each provider builds this on its own event schema — a private dictionary that defines what each event type is and which attributes it carries. Sites then license one provider’s feed and display it. So when two sites disagree, the first question is not “who is right?” but “are these even the same feed?” Two sites showing identical numbers usually share a supplier; two sites showing different numbers usually do not.
On top of the raw feed sit derived metrics: expected goals, expected assists, packing, PPDA, progressive carries. These are not observations at all. They are models, each with its own training data and its own inputs, run over the event feed. A derived metric can differ between sites even when the underlying event feed is identical.
Why football stats differ between sites: three separate causes
It helps to keep the three causes apart, because they have different consequences.
1. Definitions. Providers draw boundaries in different places. Is a deflected shot heading wide “on target”? Does a blocked shot count as a shot at all? Is possession computed as a share of passes, a share of touches, or a share of ball-in-play time? Does a header back to the goalkeeper start a new possession sequence? None of these has a single correct answer, and each choice moves the published number.
2. Models layered on top. Expected goals is the clearest case. An xG model scores each shot on the features it can see. A provider whose feed records defender and goalkeeper positions at the moment of the shot can price a chance differently from one working only from shot location, body part and assist type. The same shot might be worth roughly 0.1 to one model and roughly 0.2 to another — not an error, just two models with different information. Providers also differ on housekeeping: some include penalties in a team’s match xG total, some strip them out and report them separately.
3. Revisions and timing. Feeds are corrected after the fact. A misattributed assist gets reassigned, a duplicate event is removed, a disputed own goal is reclassified. Sites that cache a snapshot at full time and never refresh will keep showing the original tagging; sites that re-pull will show the corrected version. Two sites can therefore be reading the same provider and still disagree, purely because they read it at different moments.
The reassuring part is that these disagreements are largely idiosyncratic rather than systematic. Published comparisons of competing xG feeds find that per-shot differences mostly wash out in aggregate: over a full season, provider totals for the same team track each other very closely. Single-match differences are noisy; season-scale differences are small.
What this does to a prediction model
The practical consequence is stricter than it first appears: a model fitted on one provider’s data is a model of that provider’s data. Its coefficients absorb that provider’s definitions. Swap the feed underneath it and the model is no longer the thing you tested.
This is where models quietly break. A backtest run on one historical feed and a live system running on another will disagree, and the gap will not announce itself — the model keeps producing confident-looking probabilities, just slightly miscalibrated ones. Worse, if a provider changes its xG model version mid-history, as they periodically do, an archive can contain two incompatible definitions of the same column with no flag separating them. Fit across that boundary and you have taught the model a discontinuity that has nothing to do with football.
There is also a subtler trap. Richer data is more useful and more fragile. Every additional derived feature adds predictive potential and adds a dependency on someone else’s modelling choices, which you do not control and cannot audit.
Why we anchor on the least ambiguous number in football
That trade-off is why we are deliberately conservative about what a prediction is built on and settled against. The FootInsights model reads years of results, fixtures and prices, and every prediction it makes is graded against the final score rather than against event data. The approach is set out on how our predictions work.
A final score is the one football statistic with no definitional ambiguity. Everyone agrees on it, it is published by the competition, and it is effectively never revised. That makes a scorecard reproducible: anyone re-scoring the same predictions over the same period gets the same answer, because there is no provider-specific tagging in the chain to reproduce.
The cost of leaning on results is real and worth stating plainly. A model that learns from outcomes cannot, on its own, see that a team lost while creating far more than its opponent, so it updates on results that a chance-quality model would partly discount as variance. That is a genuine limitation, not a virtue. What you get in exchange is a record that is auditable end to end — which is also why the predictions themselves are published as a downloadable open dataset rather than described in prose.
How to compare two sites without fooling yourself
If you are checking one source against another, a short checklist removes most apparent contradictions:
- Identify the provider, not the website. Most sites document their supplier somewhere. Same supplier, same numbers.
- Read the definition before the number. If a site publishes a glossary, the disagreement is usually resolved there in one line.
- Check penalty handling on any aggregate xG figure. It alone explains a surprising share of mismatches.
- Note the snapshot date. Comparing a full-time capture with a refreshed one compares two moments, not two opinions.
- Aggregate before concluding. One match is noise. A season is signal.
And for judging a prediction site specifically, there is a shortcut past all of it: do not compare the stats, compare the probabilities against outcomes. A published forecast can be scored with a proper scoring rule regardless of which feed produced it, which is the whole point of keeping an append-only public track record. Data provenance determines how a model was built; the scorecard determines whether it worked. Only the second one is directly checkable by a reader, and it is the one worth asking any prediction site for.
Not incompetence, just different definitions
Differing statistics across sites are a feature of how football data is produced, not evidence of incompetence. Collection differs, definitions differ, derived metrics are models rather than measurements, and feeds get revised. For a casual reader that means treating single-match stat differences as rounding rather than contradiction. For anyone building a model it means something firmer: know exactly which feed you are fitted to, never mix feeds across a history, and prefer inputs whose definitions cannot drift underneath you.