Transparency

How to compare football prediction sites: a fair test in five steps

Published

Every prediction site publishes a number that flatters it. One quotes 54% accuracy, another quotes a 12% yield, a third quotes a Brier score, and none of them were measured on the same matches. Anyone trying to compare football prediction sites using those headlines is comparing three different experiments and calling the winner. The comparison is not hard to do properly — it just cannot be done from the outside using the figures each site chooses to print.

What follows is the procedure, in the order the steps actually matter.

Why you can’t compare football prediction sites on their published numbers

A published accuracy figure is a function of four things: the matches covered, the market scored, the period, and the filter deciding which predictions got published at all. Change any one and the number moves by more than the gap between a good model and a bad one.

Coverage alone does most of the damage. A site restricted to the five big European leagues faces a different problem from one that also prices second tiers and cup ties, where team strength is less stable and results are noisier. It will report a higher accuracy while doing less well, because its fixtures were easier. Nothing in the headline tells you this happened.

Step one: build one shared fixture set

Take the intersection. Collect every match both sites priced over the same window, discard everything either one skipped, and run the whole comparison on that list.

This is the step that makes everything after it valid, and it is also the step that eliminates most sites from the exercise entirely — because you cannot intersect a fixture list you were never given. A site that shows today’s picks and quietly rotates them away leaves you nothing to intersect. That is why our full prediction ledger, every match with its probabilities and its settled result, is downloadable as a CSV from the open data page: the comparison requires the raw rows, not a summary.

Step two: score one market, not “the picks”

“Accuracy” is not a market. A site mixing 1X2 calls, over/under 2.5 selections and both-teams-to-score tips into one percentage has produced a number whose value depends on the mix, and the mix is under its control.

Over/under 2.5 is roughly a coin flip and lands near 50–55% for almost everyone. Three-way 1X2 is much harder. Shift the blend towards goals markets and the headline rises without the model improving at all. Pick one market, score both sites only on that, and repeat separately for each market you care about.

Step three: score probabilities, not picks — with a worked example

Suppose the shared list is 200 matches and you are scoring 1X2.

Site A publishes picks only and gets 108 right: 54%. Site B publishes probabilities, and its highest-probability outcome came in 102 times: 51%. Site A wins the headline by three points.

Now add the baseline. On those same 200 matches, a rule that requires no knowledge whatsoever — always back the shorter of the two match-odds prices — got 106 right, or 53%. Site A’s advantage over a rule you could apply with no model, no data and no staff is two matches out of two hundred.

The probabilities are where the real difference shows up, and only Site B supplied any. Score them with the Brier score: for each match, square the gap between each stated probability and the outcome that happened, and average across the list. Suppose Site B comes out at 0.588 against 0.651 for a forecast using nothing but long-run league frequencies. That gap is small in absolute terms and it is real information, and Site A cannot produce the equivalent number at any price, because a pick carries no confidence.

This is the asymmetry worth naming plainly: a pick-only site can be beaten but cannot be measured. If one of your two candidates only publishes picks, the comparison is already over — not because it lost, but because it declined to enter.

Step four: check the publication filter

Sites that publish only their confident selections are not showing you a better model. They are showing you the easy tail of an ordinary one.

Take any competent 1X2 model, sort its matches by top-outcome probability, and keep the most confident fifth. Accuracy on that subset will land far above the model’s overall figure — 70% or more is unremarkable — because those are the matches where one side is heavily favoured. The lift comes from the fixtures, not the forecasting. A site quoting 68% on “selected” matches and a site quoting 53% on every match it prices may be running models of identical quality.

The fix is the same as step one: restrict both sites to the matches the more selective one published, and score the other on that same subset. Whatever survives that restriction is a difference in skill rather than in editorial policy.

Step five: put an error bar on the result

Three percentage points over 200 matches is not a finding.

For a proportion near 50%, the standard error over n matches is roughly the square root of 0.25 divided by n. At n = 200 that is about 3.5 points on each site’s figure — larger than the gap you are trying to interpret. A 54%-versus-51% split on 200 matches is entirely consistent with two equally good models and a run of luck.

Two things help. Score probabilities rather than picks, since a Brier score uses the full forecast and settles down faster than a hit rate. And compare paired — for each match take the difference between the two sites’ Brier scores, then average those differences. Pairing cancels the difficulty of the fixture list, which is the largest source of noise in the whole exercise, and leaves only the disagreement between the two forecasts.

Even so, expect to need several hundred matches before a modest gap becomes a claim rather than a possibility.

What a comparable site looks like

The procedure implies a checklist, and it is a short one. Comparable sites publish every prediction they make, not a curated subset; state probabilities rather than picks; timestamp predictions before kickoff; keep losing calls in the record permanently; and let you export the raw rows.

Our track record is built to survive exactly this test, which is the only reason to publish one. It is also why the record is append-only: a prediction, once made, is never edited or removed, whatever it did afterwards. A record that can be pruned cannot be audited, and a site asking to be compared on numbers it can still change is asking for trust it has not earned.

The point of the five steps is not to crown a winner. It is that most comparisons between prediction sites measure their fixture lists, their market mix and their editorial filters — and report the result as a difference in skill.