Model

Why our model can't tell you if Arsenal are better than Flamengo

Published

A reasonable thing to want from a football model is a single list: every team we cover, strongest to weakest, regardless of competition. We can produce that list trivially — the ratings are numbers, and numbers sort. What we cannot do is tell you it means anything, and the reason is worth understanding because it applies to every rating system you will encounter, including the famous ones.

The numbers look comparable

Here are the current Elo ranges for four of the leagues we model:

League Teams Lowest Highest Average
Premier League 22 1472 1788 1601
Bundesliga 27 1334 1877 1500
Championship 34 1289 1753 1473
Brasileirão 30 1296 1728 1500

Read that table naively and it says the best team in the Championship (1753) is almost exactly as strong as the best team in the Premier League (1788), and comfortably stronger than the average Premier League side. Nobody who has watched both divisions believes that for a moment.

The table is not wrong. It is being read wrong.

Why every league averages 1500

Elo is a closed, zero-sum accounting system. Every point one team gains, another loses. Run it over a set of teams that only ever play each other, and the total is conserved: the average of the pool stays fixed at wherever it started, which by convention is 1500.

Our nine domestic leagues are nine separate pools. Teams inside each pool play each other and essentially never play across, so nine independent Elo systems each sit anchored at 1500 by construction. A rating of 1753 means “far above average for this league”, and that is the entire content of the number. It carries no information about how the league itself compares to another.

The same problem exists in a different form for the attack and defence ratings our Dixon-Coles model fits. Those are estimated per league, against that league’s own scoring baseline. An attack rating of 0.50 means “scores half a log-goal more than a typical team in this competition” — and a typical team in the Bundesliga, which averages over three goals a match, is not a typical team in the Brasileirão, which averages under two and a half. The units are the same. The zero point is not.

This is the whole reason the model is fitted per league rather than globally: it makes each league’s numbers internally sharp, at the cost of making them mutually meaningless. That trade is the right one for what we publish, which is match predictions inside a competition.

What would actually fix it

Cross-league ratings are not impossible. They require matches that cross the leagues — and enough of them.

Continental competitions are the standard bridge. When Premier League clubs play Bundesliga clubs in the Champions League, those results tie the two pools together, and a rating system that ingests them can place both on one scale. This is exactly what the well-known global club rankings do, and why they can say something our domestic-only ratings cannot.

We cover one of those bridges ourselves — and deliberately do not use it to merge the pools. The Champions League runs on its own model version, cl-elo, with its own rating ladder; its results are kept out of the domestic Elo snapshots entirely. That is a design choice, not an oversight: letting a handful of European nights re-anchor nine domestic leagues would import the very extrapolation this article is about, and it would do it invisibly, inside numbers that look domestic. The reasoning is on our methodology page.

The catch is that the bridge is narrow. Only a handful of clubs from each league qualify, they are the strongest ones, and they play a few matches a season against each other. So the connection between two leagues is estimated from a small, non-random sample of their best teams — which tells you about the top of each league and extrapolates, with real uncertainty, to the rest. Even the good cross-league systems are less confident than their tidy numbers suggest, and they are far less confident about the bottom of a table than the top.

Between European and South American club football the bridge is narrower still: a single annual fixture between two continental champions, which is essentially no sample at all. Anyone giving you a confident answer on Arsenal against Flamengo is extrapolating, whether they say so or not.

Why we don’t publish the list anyway

We could take the ratings, apply an offset per league guessed from continental results, and publish a global table. It would get traffic. It would also be a number with far more uncertainty than its presentation implied, and the uncertainty would be invisible to anyone reading it.

That runs against the thing this site is for. Our methodology page states what the model knows, and a global ranking would be a claim it does not support. The track record grades predictions we can actually check against results; a cross-league rating has no fixture to be graded against, which is precisely what makes it comfortable to publish and impossible to falsify.

What to do with ratings you find elsewhere

Three questions worth asking of any global club ranking:

  1. What connects the leagues? If the answer is continental competition, the ranking is only as good as that bridge — strong for elite clubs, thin for everyone else.
  2. Is there an uncertainty range? A rating quoted to the integer with no error bar is presenting precision it does not have. This is the same problem as a win rate with no sample size.
  3. Is it being used for anything? A ranking that never gets tested against a fixture cannot be wrong, and things that cannot be wrong are not evidence.

Within a league, our ratings do real work and are graded in public every week. Across leagues, the honest answer to “who is better” is that our model does not know, and neither does anyone else with nearly the confidence they imply.