Transparency

The ranked probability score in football, and when it beats Brier

By · Published · Updated 26 Sep 2026 · 7 min read

Every public football model needs a number that says how good it is, and accuracy is the wrong one — it throws away the probabilities and keeps only the top pick. The ranked probability score in football is one of three serious candidates that don’t, alongside the Brier score and log loss. It is also the one built specifically for a sport whose outcomes sit in an order: a home win, a draw and an away win are not three unrelated labels, and the ranked probability score is the only one of the three that knows it.

That sounds like a technicality. It changes which of two forecasts gets called better — worth understanding before you judge anyone’s record, including ours.

What the ranked probability score in football actually measures

Start with the problem it exists to solve. Take two forecasts of the same match:

  • Forecast A: 50% home, 20% draw, 30% away
  • Forecast B: 50% home, 30% draw, 20% away

The home team wins. Both forecasts had the winner at 50%, so on accuracy they are identical. On the Brier score they are also identical — sum the squared errors and both come to 0.38, because the 10 points of probability that moved from away to draw moved between two outcomes that neither occurred.

But most analysts would say B was the better forecast. It leaned toward the draw, the outcome adjacent to a home win, rather than toward an away win, which is as wrong as a forecast can be. Football results live on a line — away, draw, home — and being wrong by one step is not the same as being wrong by two. The ranked probability score is the scoring rule that encodes exactly that, and the only one of the three that separates A from B.

The formula, worked through

RPS is the Brier score applied to cumulative probabilities rather than individual ones. Order the outcomes (home, draw, away), build the running totals, and compare them with the running totals of what actually happened.

For a three-outcome match:

  1. Cumulative forecast: P₁ = p_home, P₂ = p_home + p_draw (P₃ is always 1, so it contributes nothing).
  2. Cumulative outcome, if the home team won: O₁ = 1, O₂ = 1. If it was a draw: O₁ = 0, O₂ = 1. If away: O₁ = 0, O₂ = 0.
  3. RPS = ½ × [ (P₁ − O₁)² + (P₂ − O₂)² ]

Dividing by two — one less than the number of outcomes — puts the result on a 0-to-1 scale, where 0 is a perfect forecast and 1 is putting all your probability on an away win that never came.

Run the two forecasts through it, with the home team winning:

Forecast P₁ P₂ Squared errors RPS Brier
A — 50 / 20 / 30 0.50 0.70 0.25 + 0.09 0.170 0.380
B — 50 / 30 / 20 0.50 0.80 0.25 + 0.04 0.145 0.380

Brier calls them a tie. RPS gives B the better (lower) score, because its cumulative probability climbed toward certainty faster on the correct side of the result.

The rule is symmetric, which is the test of whether it is doing something real rather than just rewarding draws. Rerun the same two forecasts against an away win and the ordering flips: A scores 0.370, B scores 0.445. B is punished harder precisely because it had drained probability away from the outcome that happened.

Three scoring rules, three blind spots

All three metrics worth using are strictly proper: a forecaster minimises their expected score by publishing exactly what they believe. There is no shading, no hedging toward safe numbers, no way to game the metric by being strategically vague. That is why a published Brier or RPS is a meaningful commitment in a way that a published win rate is not.

Where they differ is how much of the forecast they look at:

  • Log loss (−ln p of whatever happened) is local: it reads only the probability assigned to the actual outcome and ignores how the rest was distributed. On the example above it scores A and B identically at 0.693, and it would do so however the remaining 50% was split. Its other quirk is unforgiving: a probability of zero on an outcome that occurs scores infinity, so any model quoting log loss must floor its probabilities somewhere.
  • Brier is sensitive to the whole distribution but treats the three outcomes as unordered categories. It notices how much went to the wrong outcomes, not which wrong outcomes.
  • RPS is sensitive to the whole distribution and to the ordering. It is the only one that treats a draw as sitting between the two wins.

None of these is a substitute for looking at calibration directly. A scoring rule collapses two distinct virtues — being calibrated and being confident — into one number, and two models can land on the same score by being good at different things.

The scale is narrower than it looks

Scoring-rule values need benchmarks or they mean nothing, and RPS values compress into a tight band that catches people out.

Take the base rates of a typical top division — roughly 44% home, 25% draw, 31% away — and ask what a forecaster scores who ignores the teams entirely and quotes those base rates on every single match. Under those same rates, the arithmetic gives an expected RPS of about 0.230. Now do it for a forecaster who is even lazier and quotes 33/33/33 every time: about 0.236.

Six thousandths. That is the entire distance between “knows the league’s long-run tendencies” and “knows nothing at all”, and every real model — and the closing betting market, the strongest public forecaster there is — has to fit its improvement below that base-rate figure in a space of comparable size.

Two things follow. Treat any large claimed gap with suspicion — on this scale, improvements are measured in hundredths. And a small true edge needs a lot of matches before it separates from noise, which leads directly to the argument against using a distance-sensitive rule at all.

The case against RPS

Distance-sensitivity is not universally accepted as a virtue, and the honest thing is to say so. The published critique of RPS runs roughly like this: the purpose of a scoring rule is to identify better forecasters as quickly and reliably as possible, and rewarding a forecast for being “close” in outcome-space does not obviously serve that purpose. Being one step wrong still means the bet lost, the prediction failed, and the model misread the match.

Worse, in simulation studies log loss tends to distinguish a genuinely better model from a worse one faster — in fewer matches — than RPS does. Since everyone evaluating football models works with limited samples, a rule that converges more slowly carries a real cost that its conceptual elegance does not repay.

There is no settled verdict. RPS remains standard in the academic football-forecasting literature, and in practice the three rules agree most of the time. The reasonable position is that the choice is a judgement, that a service should say which rule it uses and why, and that any of the three beats a headline accuracy figure.

What to demand from a prediction service

The scoring rule is downstream of a more basic question: is the record complete and gradeable at all? A metric computed on a hand-picked subset of predictions is worthless whichever formula produced it. Three things have to be true before the choice of formula even matters — every prediction graded and kept, losers included; the full probability distribution published before kick-off, since you cannot compute RPS, Brier or log loss from a tip; and the scoring rule named alongside its benchmark, because a number without a baseline is decoration.

Our own model publishes probabilities for all three outcomes ahead of every match, and the track record grades every one of them after full time and keeps them, Brier score included. The model’s known limits are written up on how our predictions work.

Whichever scoring rule you prefer, the property that makes a record checkable is the same one: the numbers were published first and nothing was removed afterwards. That is what a proper scoring rule is designed to test, and it only works on a record that stayed honest.