Transparency

Statistical significance in a betting record: the test and its limits

By · Published · 5 min read

Sooner or later every betting record gets defended with the same sentence: “the results are statistically significant.” It sounds like the end of the argument. It is closer to the beginning of one. Statistical significance in a betting record is a specific calculation with a specific null hypothesis, and almost every public use of the phrase gets the null hypothesis wrong, the sample size wrong, or the interpretation wrong — usually all three. This post walks the test itself: what it compares against, how to run it on a record in about a minute, what a p-value does and does not license you to believe, and why a record that passes at the conventional threshold can still be more likely luck than skill.

What statistical significance in a betting record actually tests

A significance test never asks “is this record good?”. It asks a narrower question: if this bettor had no edge whatsoever, how often would chance alone hand them a result at least this good? That probability is the p-value. Small p-value, surprising-under-luck result. Large p-value, unremarkable.

The critical word is edge, and this is where most public claims fall apart. “No edge” does not mean a 50% win rate. It means winning exactly often enough to break even at the prices actually taken. A record of 62% winners looks impressive until you notice the average price was 1.50, where break-even is 66.7% — those results are evidence of a losing system, not a winning one.

So the null hypothesis is always set by the odds:

break-even win rate = 1 / average decimal odds

At average odds of 2.10, break-even is 1 / 2.10 = 47.6%. Everything is measured against that number, not against a coin flip.

The test, in four lines of arithmetic

Take a record of 200 bets at average decimal odds of 2.10, of which 104 won. That is a 52% strike rate and a profit of (104 × 1.10) − 96 = 18.4 units on 200 staked: a +9.2% return. Promotional material writes itself.

Now test it. Under the null, the true win probability is p₀ = 0.476. The standard error of an observed win rate over n bets is:

SE = √( p₀ × (1 − p₀) / n ) = √( 0.476 × 0.524 / 200 ) = 0.0353

The observed rate sits 0.520 − 0.476 = 0.044 above the null, so:

z = 0.044 / 0.0353 = 1.24

A z of 1.24 corresponds to a one-sided p-value of about 0.11. In plain terms: a bettor with no edge at all, betting these prices, produces a 200-bet run this profitable or better roughly one time in nine. A +9.2% return, fully verified and honestly reported, is not significant at any conventional threshold. It is an ordinary outcome for someone with no skill.

How many bets significance actually needs

Rearranging the same formula for the sample size that reaches one-sided significance at 5% (z = 1.645) shows why records rarely clear the bar. All rows assume average odds of 2.10:

Claimed return Implied strike rate Bets needed for p < 0.05
+9.2% 52.0% ~350
+5% 50.0% ~1,200
+3% 49.0% ~3,300
+2% 48.6% ~7,400
+1% 48.1% ~29,800

The relationship is quadratic: halving the edge quadruples the bets required. That is the uncomfortable part, because the top row is not a realistic long-run edge — a sustained 9% return against the closing market would be extraordinary — while the bottom rows are. A genuine, survivable edge is precisely the kind that takes tens of thousands of bets to prove. The records that reach significance quickly are the ones claiming edges too large to be real.

Longer odds make it worse still, since variance per bet scales with the price. This is the same variance arithmetic that governs sample size in betting records, viewed from the hypothesis-testing side rather than the confidence-interval side.

Why p < 0.05 is weaker evidence than it sounds

Suppose a record does clear the threshold. You still cannot conclude there is an edge, for a reason that has nothing to do with the arithmetic and everything to do with how many records exist.

A p-value answers “how likely is this result given no edge?” The question you care about is the reverse: “how likely is an edge given this result?” Converting one into the other requires knowing how common edges are to begin with.

Work it through with round numbers. Imagine 1,000 tipsters. Suppose 2% of them — 20 people — genuinely have an edge, and the tests are powerful enough to detect half of the real ones:

  • 20 skilled tipsters, 50% detected → 10 significant records
  • 980 unskilled tipsters, 5% false-positive rate → 49 significant records

Fifty-nine records now carry a p-value under 0.05. Ten of them are real. A record chosen from that pile has roughly a 17% chance of reflecting genuine skill — meaning that even after passing the test, luck remains over four times likelier than skill. Nothing was cherry-picked and nobody lied; the base rate did all the damage. This is survivorship bias stated as arithmetic: in a large enough population, significance is manufactured by volume.

Two practical consequences follow. First, a significance claim is only meaningful if you know how many strategies, filters and start dates were tried before the reported one — a record tested twenty ways and presented once is guaranteed a small p-value. Second, a threshold of 0.05 is far too permissive for a field this crowded; serious verification uses much stricter thresholds precisely because the population of candidates is enormous.

The faster honest alternative

None of this makes verification hopeless. It makes profit a slow instrument, because each settled bet reveals a single bit of information: won or lost.

A forecaster who publishes calibrated probabilities can be graded on the full number instead. Scoring how far each stated probability sat from the outcome extracts far more information per match than a win/loss tally, so systematic overconfidence becomes measurable in hundreds of predictions rather than tens of thousands of bets. That is the argument for judging a forecaster on a proper scoring rule over a complete, timestamped, append-only history — the structure our own track record uses — rather than on a profit figure and a significance claim attached to it.

So when a record is described as statistically significant, three questions settle it: significant against which break-even rate, over how many bets, and out of how many records that were tried. If any of the three is unavailable, the phrase is decoration. If all three are available, the honest answer is usually that the sample is still too small — including, for most published records, the flattering ones.