How accurate is Forebet? How to check any prediction site yourself
Published
“How accurate is Forebet?” is one of the most-searched questions in football prediction, and variants of it get asked about every large free tips site — PredictZ, WinDrawWin, the lot. It is exactly the right question. It is also, asked that way, unanswerable: the only people who can produce the number are the ones with an interest in it, and no visitor can check their work. This article is about the version of the question you can answer, and the fortnight of light bookkeeping it takes to answer it.
To be clear about what follows: none of this is a claim that Forebet’s predictions are poor. It is a claim about verifiability, and the same objection applies with full force to FootInsights.
Why a site’s own accuracy figure can’t settle it
A published success rate is a summary produced by the party being summarised. Three things have to be true before it means anything, and a visitor can confirm none of them from outside:
- The predictions were fixed before kick-off. A results table assembled after the fact, from a database its owner can edit, records intentions that cannot be distinguished from hindsight.
- Every prediction is in the denominator. Drop the weekends that went badly, or quietly exclude the leagues where the model struggles, and the surviving figure is arithmetic about a curated subset.
- The definition is stated. “Accuracy” over 1X2 on heavy favourites, over both-teams-to-score, or over correct scores are three wildly different numbers, and a headline percentage rarely says which one it is.
None of these require bad faith to go wrong. A site that reports honestly still can’t prove it reported honestly — which is why the burden belongs with the record’s structure rather than with anyone’s assurance. We wrote the general version of this test in how to verify a football tipster’s track record; what follows is the version you run yourself, on whichever site you actually use.
Decide what “accurate” means before you measure
Most disappointment with prediction sites comes from measuring the wrong thing, so settle this first.
Hit rate is the share of predictions that came in. It is the number everyone quotes and the least informative one available, because it depends almost entirely on which matches got predicted. In our nine-league dataset, simply backing the home team in every fixture lands about 44% of the time, and a model that only ever tips short-priced favourites will beat that comfortably while losing money on every one of them.
Calibration asks whether the stated probabilities are honest: of all the matches given a 70% chance, did close to 70% actually happen? This is the property that separates a model from a guess, and it is measurable with a proper scoring rule like the Brier score. A site that publishes no probabilities at all — just a pick — cannot be scored this way, which is itself a finding.
Return asks whether following the picks made money at the prices available. It is the only one that pays for anything, and it is where the vast majority of tipping records quietly fail against closing odds.
A site can be excellent on the first and hopeless on the third. Decide which one you care about, because the audit below measures whichever you point it at.
The 30-day audit
You need a spreadsheet and about five minutes a day. The design goal is that the record is built by you, before results exist, so nothing in it can be revised later.
- Freeze the picks each morning. Pick one league and one market and keep them fixed for the whole month — mixing markets halfway through is how audits become unreadable. Each day, before the first kick-off, copy the site’s prediction for every fixture into a row: date, teams, the pick, the stated probability if one is given, and the best odds you could actually get at that moment. Recording the odds at tip time is not optional; without them the return column is unrecoverable afterwards.
- Settle honestly the next day. Add the real result and mark each row won or lost. Every fixture you logged gets settled, including the ones you would rather forget and the ones that were postponed or void — those come out of the sample with a note, not silently.
- Compute three numbers, not one. Hit rate is wins over settled picks. Return is your net stake movement at level stakes, expressed as a percentage of total staked. If probabilities were published, average the Brier score across the month.
- Compare against a baseline you compute the same way. This is the step almost everyone skips, and it is the one that makes the exercise mean anything. Run a second column in the same spreadsheet, on the same fixtures, using a rule that requires no skill at all: always back the home side, or always back the bookmakers’ favourite. A prediction service is only worth attention if it beats the free alternative on the metric you chose.
What to expect when you run it
Two warnings, so the result doesn’t mislead you in the other direction.
First, a month is a small sample — usually a very small one. Football’s variance is large enough that a genuinely good model has losing months routinely and a coin-flipper has winning ones. Thirty picks cannot separate skill from luck; a few hundred begin to. We laid out the arithmetic of how badly small samples deceive in sample size in betting records, and a one-month audit should be read as a smell test, not a verdict.
Second, the ceiling is lower than the marketing suggests. Top-flight football is close to the noisiest sport there is to forecast, and no model — statistical, machine-learned or otherwise — escapes that. Advertised rates in the 80s and 90s are a strong signal that the denominator is doing the work. Our honest account of where the ceiling actually sits is a useful reference point before you judge whatever number your spreadsheet produces.
Apply the same test to us
We publish probabilities rather than bare picks, freeze every prediction before kick-off with its model version attached, and settle all of them — misses included — on a public track record scored per league and per market with the Brier score rather than a headline win rate. The methodology page sets out what the model uses and where it is weakest.
We deliberately quote no performance figures in articles, ours or anyone else’s, because a number typed into static text can’t be audited and can’t go stale honestly. The live page is the claim, and the audit above is precisely the instrument for testing it. Run it on us, run it on Forebet, run it on whichever site you were going to trust anyway — and whatever you conclude, stake only what you can afford to lose, because even a real edge is a small one and variance never signs a contract.