Model watch

Atlético 2–0 Málaga: what one result proves about a model–market gap

Published

Full time in Madrid: Atlético 2, Málaga 0. On Monday we published something prediction sites rarely publish — a post explaining why the betting market’s number was probably better than ours. Our 16 August model run priced the home win at 58.3%; the market had Atlético around 73% with the margin stripped, and it never really moved — the morning-of-kickoff fetch (Paddy Power, 19 August) still showed 1.30, roughly 72% after removing the overround. A fifteen-point disagreement, on the public record, now resolved by ninety minutes of football.

So: who was right?

The honest answer is the reason this post exists. On tonight’s evidence — the market, slightly. On any evidence that actually matters — ask again in a few months.

The market earns the better mark tonight

Both numbers named the same favourite, so in the shallow sense both “got it right”. But probabilities aren’t picks, and the scoring rules we grade ourselves with don’t treat them as picks: when the predicted outcome happens, the more confident call earns the better score. The market said Atlético with more conviction than we did, Atlético won, so tonight’s round goes to the market. No spin makes that otherwise.

The betting reading is the same. Taken at face value, our numbers made the Málaga side of the market look generous — we had Málaga avoiding defeat 42% of the time (22.0% draw plus 19.7% win, 16 August run) against a market pricing the away win below 10%. Monday’s post said plainly that this apparent value was noise, because the model’s Málaga number was a league-average placeholder, not information. Anyone who took the placeholder at face value instead backed a bet that lost.

What one result cannot tell you

Here’s the part the “told you so” framing gets wrong: tonight’s result does almost nothing to tell us whether 58% or 72% was the truer number. Both said Atlético more likely than not. A weighted coin came up on its heavier side — that’s consistent with almost any weighting. A 58% event landing is unremarkable; so is a 72% event landing; so, for that matter, would either of them failing have been. Single results grade picks. They barely scratch probabilities.

The same fixture offers a neat proof of that from the other direction. Our 16 August run had two near coin flips on the goals markets: 53.5% for Over 2.5 and 52.0% for both teams to score. The match finished 2–0 — under, clean sheet — so both landed on the “no” side. Were those misses? Calling them misses is exactly as shallow as calling the 1X2 lean a hit. Numbers in the low 50s land against you almost half the time by construction; that’s what the number means. The only way to find out whether a model’s 53% is honest is to collect hundreds of them and check how often they land — which is precisely what our public track record is for.

What would settle the model-versus-market question from Monday is a cohort, not a fixture: every promoted-club match this season, graded the same way, until the sample is big enough to show whether our cold-start placeholder systematically overrates newcomers against the market’s squad-level knowledge. One August evening, however tidy the scoreline, is not that sample.

What the result does settle

Two things, both small and both real.

First, Málaga now have a match in the dataset our model is fitted on. The league-average placeholder that produced Monday’s 58.3% starts fading tonight, replaced one result at a time by evidence about this actual squad. The fix we described on Monday isn’t a promise — it’s now mechanically underway, and the match page keeps the full prediction-versus-result record for this fixture permanently.

Second, the 58.3% gets graded and folded into the track record like every other call — a “hit”, as it happens. We’d rather you read it as what it was: a number we published, argued against, and left on the record anyway, because deleting inconvenient predictions is how this industry lies and grading all of them is how we’d like to differ.

The rule from Monday survives its first test unchanged: when our model disagrees with the market because it knows less, the gap is noise, not value. Tonight didn’t prove that rule. It just declined to contradict it — and we’ll keep counting.