International football predictions: why national teams break the model
Published
Every few months the club calendar stops and the fixture list fills with countries. The models that price those matches are usually the same models that price league football, pointed at a different set of teams — and that is the problem. International football predictions are not club predictions with different badges. Almost every assumption a domestic model depends on quietly stops holding, and the failures compound. Here is what breaks, and what an honest forecast for a national team looks like as a result.
Ten matches a year is not a season
Start with the arithmetic. A club side plays roughly forty to fifty competitive matches in a season; a national team plays ten to a dozen. Same calendar, a quarter of the evidence.
That ratio matters more than it first appears, because a rating system does not learn in units of time — it learns in units of matches. Suppose a model weights history with a half-life of thirty matches, a fairly typical choice. For a club, thirty matches is about eight months: the rating tracks a genuine change in team strength within a single season. For a national team, thirty matches is three years or more. By the time the rating has fully absorbed what a side became, the manager has usually gone and a third of the squad has aged out.
Shorten the half-life to compensate and you buy responsiveness with noise: a single bad night against a well-organised opponent now swings the rating hard, and international fixture lists are full of single bad nights. No setting escapes this. The information is not there, and every choice about weighting it is a choice about which way to be wrong.
The squad is not the team
Domestic models treat a club as a persistent object, and the approximation is good: the same eleven or twelve players start most weeks, they train together daily, and the squad turns over gradually across transfer windows.
A national team is not persistent in that sense. It disbands, its members return to other coaches and other systems, and it reassembles months later — often without three or four of the players whose performances produced the rating. Withdrawals arrive in the days before a camp, after the model has already fitted a number to “the team”.
This breaks the unit of analysis rather than merely adding noise to it. When a model says a national side is worth a certain strength, it is averaging over squads that in some cases share barely two-thirds of their personnel. A club model handling an injury adjusts a known team for a known absence — a solvable problem. An international model is often being asked what an unfamiliar combination of players will do together, with no fixtures in which that combination has ever appeared.
Three competitions wearing one label
Group every international fixture together and you are fitting one set of parameters to several genuinely different games.
Friendlies are exhibitions. Coaches rotate freely, use expanded substitution allowances, experiment with shape and accept defeat cheaply. It is a real result, produced under incentives no competitive match shares.
Qualifiers are competitive and lopsided. Confederations pair their strongest sides with their weakest in the same group, so a large share of the fixture list is decided before kick-off.
Tournament matches shift again, and shift within the tournament: a final group game means everything to one side and nothing to another that has already qualified, which is a motivational asymmetry with no real club equivalent. Knockout rounds then add extra time and penalties, which have their own probability structure and are not simply more football.
Pool all three and you are estimating an average over three populations. Separate them and you have cut an already-thin sample into three thinner ones. Neither option is good.
Mismatches feel informative and are not
Qualifying produces a steady supply of six-goal wins, and they are nearly worthless as evidence.
A rating updates on surprise. If a model already expected a comfortable win, a comfortable win confirms what it thought and barely moves the number — which is correct behaviour. So a meaningful fraction of an international team’s small fixture list carries almost no information about its strength. The effective sample is smaller than the raw match count, and it is smallest exactly for the elite sides whose ratings people most want to compare.
Goal-based models have a second problem here. Poisson-style scoring models are calibrated where a team expects one to two goals; asked to price an expectation of four or five, the tail behaves badly. Unless extreme results are damped, a side racking up huge margins against weak opposition ends up rated as an attacking juggernaut on evidence that says mostly “played someone much worse”.
Ratings stop at the border
Every league is effectively a closed pool, so a domestic rating means “strong relative to this league” and nothing more — an argument worked through in full in why our model can’t put two leagues on one scale.
International football inherits that problem and worsens it. Qualifying happens almost entirely within a confederation, so each confederation is its own weakly-connected pool, and the fixtures linking those pools are mostly friendlies — the least representative matches on the calendar — plus a handful of tournament games every few years.
Then the sting: in club football the cross-pool comparison is a curiosity you can decline to make, while a World Cup forces it, at the precise moment the connecting data is thinnest. Any confident cross-confederation number is an extrapolation, and should be labelled as one.
Neutral ground removes a fitted parameter
Home advantage in domestic football is a fitted quantity worth a few tenths of a goal, and it varies by league. Tournament football often takes place on neutral ground, where that parameter does not apply, or in a host nation, where something quite different applies to exactly one team. Qualifying, meanwhile, has more venue effect than league football, not less: long-haul travel, unfamiliar climate, altitude, pitches nothing like the ones the visitors use weekly. The correct adjustment is therefore larger than the domestic one for some fixtures and zero for others, and a model carrying a single league-derived constant across all of them is wrong in both directions on different nights.
What honest international football predictions look like
None of this makes the problem unforecastable. It makes it a problem with wider error bars, and the honesty is in showing them. Three practices separate a serious international forecast from a merely confident one:
- Probabilities pulled toward the prior. When evidence is thin, the right response is less extreme numbers, not the same numbers with a caveat attached. A model as confident about a national side as about a club it has watched for two hundred matches is not accounting for what it does not know.
- Segmented accuracy, always. Internationals are a different population, and mixing them into a club calibration table hides both. One overall figure spanning leagues, cups and internationals describes no competition in particular. Reporting by competition type keeps the weak segment visible instead of averaging it into respectability — the reasoning behind how our own track record is presented.
- A named benchmark. “Better than a coin flip” is no achievement in a market full of heavy favourites. The comparison that matters is the closing market price — and international markets, being thinner and fatter-margined than major league markets, look beatable more often than they are.
Why we don’t publish them
FootInsights models domestic club competitions — the leagues we cover — and does not publish national-team predictions. That is a data decision rather than a lack of interest.
We could point the existing pipeline at international fixtures tomorrow. It would return numbers that looked exactly like our league numbers, with uncertainty several times larger and nothing on the page showing it. Publishing a probability implies you can say how much to trust it; for international football, from the data a domestic pipeline holds, we cannot — and a prediction whose error bar we cannot state is decoration.
The short version: international football is not more random than club football, it is more unobserved. Ten matches a year, a roster that reassembles each camp, three competitions sharing a label and confederations that rarely meet add up to a sport seen through a much smaller window. Any forecast ignoring the size of that window reads as precise, right up until it is checked.