Model

Why early-season football predictions are the least reliable

Published

There is a period every year when football prediction is at its least trustworthy, and it happens to be the period when interest is highest. A new season starts, everyone wants a view on it, and the models handing out views have less to go on than at any other point in the calendar.

We would rather say this plainly than let it be discovered.

What a rating actually contains in August

Our model rates teams from results, weighting recent matches more heavily than old ones. That design is right for most of the year and awkward at the start of one.

On the opening weekend, the most recent competitive evidence about a club is from last season — a different squad, sometimes a different manager, always a different set of circumstances. The ratings carry over rather than resetting, because starting every team from scratch each August would be worse: it would throw away the genuine, persistent differences between clubs that make a rating useful in the first place.

So the numbers you see in August are a considered estimate built on evidence that is several months old and describes a team that may no longer exist in that form. That is the best available answer. It is not a good one.

The three things that break

Squad turnover is invisible. A club that sold its two leading scorers and rebuilt its defence carries the rating those departed players earned. The model has no transfer data, no squad valuation, no way to know the team changed. It will find out the same way everyone else does — by watching results — and that takes matches.

Promoted teams have no usable history. A side arriving from a lower division has no record in the competition being modelled, so it starts at something close to league-average and gets corrected from there. This is the model’s single weakest ground, and early-season fixtures involving promoted clubs deserve the most scepticism of anything we publish.

Nothing has happened yet to correct anything. Time-decay weighting only helps once there are recent matches to weight. In the first few rounds, “recent” and “last season” are the same thing.

Worth noting what does not break: the league-level base rates. How often a competition draws, how many goals it produces, how strong home advantage is — those are stable across seasons and carry over reliably. It is the team-level ratings that go stale, not the structure they sit in.

How long it takes to settle

There is no clean switch. The ratings improve continuously as results arrive, and the practical shape is that the first handful of rounds move them a lot, and by around the ten-match mark most teams have accumulated enough current-season evidence to dominate the carried-over signal.

Our model also refuses to fit a league at all until it has a substantial body of finished matches to work from — a couple of hundred, in the current configuration. That threshold exists for the same reason: a fit on a thin sample produces confident-looking parameters that are mostly noise, which is worse than no model.

If you want a rule of thumb: treat predictions in the first month of a season as carrying materially more uncertainty than the probability alone suggests, and treat anything involving a newly promoted side as more uncertain still.

What this means for reading anyone’s early-season claims

This is the practical part, because the same constraints apply to every model, and only some of them mention it.

A confident August prediction is a claim about last season. No public model has meaningfully more information than results, and results are old in August. If a site’s early-season probabilities look as sharp as its March ones, that is a statement about presentation, not information.

Early-season hit rates are noisy in both directions. A forecaster having a superb opening month has probably been lucky; one having a terrible month has probably been unlucky. Neither tells you much, for the same reason any short run tells you little in a sport this variable.

Watch out for records that start in August. A track record whose sample begins at a season boundary is measuring the hardest period, or — if it looks unusually strong — has been started at a convenient moment. The verification checklist asks for the complete record precisely to catch this.

The market has the same problem, and less of it. Bookmakers price transfers, pre-season form and team news that our model cannot see, so the gap between a public model and the closing price is widest in August. If you are comparing the two, expect the market to look better early and the gap to narrow as the season provides evidence.

Where we show it

Our track record is append-only and carries sample sizes, so an early-season stretch is visible as exactly what it is rather than being smoothed away. The methodology page lists the promoted-team problem and the absence of lineup data among the model’s known limits, and both bite hardest right now.

The alternative — publishing August probabilities with the same confidence as April ones and saying nothing — is more comfortable and less honest. Treat early-season numbers as the roughest estimates we produce all year, and stake only what you can afford to lose.