Transparency

Sample size in betting: how many bets before a record means anything?

Published

Every betting record you will ever be shown — a tipster’s banner, a model’s backtest, your own spreadsheet — is a sample, and most of them are far too small to mean what they claim. Sample size in betting is the least glamorous concept in the entire subject and the one that quietly decides whether a number is evidence or noise. A 70% win rate over 20 picks is an anecdote. The same 70% over 2,000 picks is a phenomenon. This post puts actual numbers on the difference: how wide the noise band around a win rate really is, how many bets it takes to demonstrate an edge of a given size, and why return on investment is even slower to converge than win rate.

None of this requires more than one formula. But once you have it, most of the records you see online stop being impressive.

Why sample size in betting is nearly everything

Start with a bettor who has no skill at all: they flip a fair coin on even-money bets, with a true win rate of exactly 50%.

Over 20 bets, the binomial distribution says this coin-flipper has roughly a 13% chance of winning 13 or more — a 65% win rate. Roughly one skill-free bettor in eight will show you that record honestly, without cherry-picking a thing. Now imagine a platform hosting a few hundred tipsters. Dozens of them will be sitting on a 65%+ streak over their last 20 picks at any moment, purely by chance, and those are precisely the records that get screenshotted.

This is why a short hot streak proves nothing and — just as important — why a short cold streak proves nothing either. Small samples generate extreme results in both directions. The question is never “is this record good?” but “is this record distinguishable from luck at this sample size?”

The noise band around any win rate

For a record of n bets with true win probability p, the standard error of the observed win rate is:

SE = √( p × (1 − p) / n )

A useful rule of thumb: the observed rate lands within about two standard errors of the truth 95% of the time. Run the numbers for a genuinely 50/50 bettor:

Bets Standard error 95% noise band around 50%
25 10.0 pts 30% – 70%
100 5.0 pts 40% – 60%
400 2.5 pts 45% – 55%
1,000 1.6 pts 47% – 53%
2,500 1.0 pts 48% – 52%

Read the second row carefully: over 100 even-money bets, anything between 40% and 60% is consistent with zero skill. A “58% win rate, verified, 100 picks” banner sits comfortably inside the band a coin flip produces. It takes around 1,000 bets before the noise band tightens to a few points — and most public records never get there.

How many bets to demonstrate an edge

Flip the question around: suppose a bettor has a real edge on even-money bets. How long until it shows? For the observed rate to sit clearly outside the coin-flip noise band — two standard errors above 50% — the sample needs roughly:

  • True 60% win rate (a huge, market-beating edge): about 100 bets
  • True 55% win rate (an excellent professional edge): about 400 bets
  • True 52% win rate (the kind of thin edge that actually survives in efficient markets): about 2,400 bets

And these are optimistic numbers: at exactly those sample sizes, a bettor with the real edge only clears the significance bar about half the time — variance can hide a genuine edge just as easily as it can fake one. Want to reliably detect a 2-point edge? You are closer to 5,000 bets, which at ten bets a week is a decade of records.

The uncomfortable conclusion: the edges most worth having in football betting are exactly the ones that take thousands of bets to prove. Anyone confidently claiming a large edge from a small sample is describing their luck, not their skill.

ROI is even noisier than win rate

Most records lead with return on investment rather than win rate — and ROI converges more slowly, because the odds multiply the variance.

At even money (decimal odds of 2.00), each one-unit bet returns +1 or −1, so a zero-edge bettor’s ROI over 100 bets has a standard deviation of about ±10%. An “8% ROI over 100 bets” record is less than one standard deviation from breaking even — nothing.

Lengthen the odds and it gets worse. At decimal odds of 4.00 with a fair 25% win probability, each bet returns +3 or −1. The variance per bet triples relative to even money, so a longshot-heavy record needs roughly three times as many bets to reach the same statistical precision. A tipster specialising in prices above 4.00 can post a 40% ROI over 150 picks and still be entirely inside the luck band — which is exactly why longshot records dominate promotional screenshots: they generate spectacular small-sample results in both directions, and only the good tails get shown. (They are also the prices where bookmaker margins bite hardest, which makes a sustained longshot ROI doubly implausible.)

Reading a real record with this lens

The math above turns into a short field guide:

  1. Find n first. Before the win rate, before the ROI, before anything. If the sample size isn’t stated, the record is decoration.
  2. Compute the noise band. √(p(1−p)/n), double it. If the claimed performance sits inside the band a zero-edge bettor produces, you have learned nothing — favourable or unfavourable.
  3. Check the odds profile. Longer average odds mean wider bands. Judge a longshot record against a longshot-sized noise band, not an even-money one.
  4. Demand the full history. Sample-size math only works on all the bets. A record that starts at the convenient moment, or quietly drops a bad month, has an effective sample size of zero — the selection did the work, not the skill. This is the core of how to verify a football tipster’s track record, and it is why our own track record is append-only and complete: every settled prediction stays, including the misses, so the n you divide by is the real one.

The shortcut: score the probabilities, not the profit

There is one honest way to shorten the wait. A win/loss record extracts a single bit of information per bet, but a forecaster who publishes probabilities can be scored on how far each stated probability sat from the outcome — which is what a Brier score does. Grading the full probability rather than the binary result extracts more information from every match, so systematic over- or under-confidence becomes visible in hundreds of predictions rather than thousands of bets.

That, ultimately, is the argument for judging forecasters on scored probabilities over a complete public history instead of on profit claims: not that it flatters anyone, but that it is the fastest statistically honest verdict available. Until the sample is large, the only defensible position — about any record, ours included — is patience.