How to tell if a prediction was right
A forecast of 70% isn't wrong when the event fails to happen. Here's how calibration and the Brier score actually judge whether a prediction was any good.
Outcomer Team · Jul 19, 2026
Suppose a market prices an event at 70% and the event never happens. Was the market wrong? Most people say yes. The honest answer is: you cannot tell from a single case. A 70% forecast is a claim that things like this happen about seven times in ten — which means it is supposed to fail three times in ten. Judging one probability by one outcome is like judging a weighted coin from a single flip.
That gap between "what felt right" and "what was actually good" is the whole subject of forecast scoring. Two ideas do most of the work: calibration and the Brier score.
Calibration: do your 70% calls happen 70% of the time?
Calibration asks a simple question. Take every time a forecaster — or a market — said "70%". Group them together. In a well-calibrated set, close to 70% of those events actually happened. Do the same for the 30% calls, the 90% calls, and so on. If each bucket lands near its label, the forecaster is calibrated.
This is powerful because it needs no single prediction to be "correct". It only asks that the probabilities mean what they say over many cases. A weather service that says "30% chance of rain" is doing well if it rains on roughly three of every ten such days — no more, no less. Raining on zero of them would be as much a miss as raining on all of them, because it would prove the 30% was really something lower.
Calibration also exposes the two classic failures. Overconfidence is saying 90% when the real rate is 70% — big, dramatic claims that reality does not back up. Underconfidence is huddling near 50% to avoid being wrong, which is safe but useless. A good forecaster is bold and right: sharp probabilities that still land where they should.
The Brier score: one number for accuracy
Calibration tells you whether the labels are honest, but it does not reward sharpness on its own. For a single combined measure, forecasters use the Brier score, introduced by the meteorologist Glenn Brier in 1950 to grade weather forecasts.
The maths is friendly. Take your probability as a decimal, subtract the outcome (1 if it happened, 0 if it did not), square the difference. Average that across all your forecasts. Lower is better.
Say you forecast 0.70 and the event happens. Your error is (0.70 − 1) = −0.30, squared to 0.09. If instead it does not happen, the error is (0.70 − 0) = 0.70, squared to 0.49 — a much bigger penalty for being confident on the wrong side. Forecast a lazy 0.50 and you always score 0.25, win or lose. That is why the Brier score rewards forecasters who move off the fence and get it right, and punishes confident mistakes hardest of all. A perfect seer scores 0; someone who says 50% to everything scores 0.25; the worst possible score is 1.
Because it is a "proper" scoring rule, your best long-run strategy under the Brier score is simply to report your true belief. Shading your number to look bolder or safer only hurts you over time.
Why this matters for a market
A prediction market is a forecaster made of thousands of people, and the same tests apply. Studies of large forecasting platforms consistently find that well-run markets and reputation-based communities are reasonably well calibrated on short-term, clearly defined questions — their 70% prices really do resolve yes about seven times in ten. Calibration weakens on vague or very long-horizon questions, where even the resolution criteria are debatable. Knowing that distinction tells you when to lean on the crowd and when to stay sceptical.
It also changes how you read the board. As we cover in reading the odds, a price is an implied probability, not a promise. And as the wisdom of crowds explains, the number is an average of many judgements — one that earns trust only if it stays calibrated across many resolved markets. For more on how markets stack up against reality, see are prediction markets accurate.
A better way to keep score
The practical takeaway is to stop grading yourself on your loudest hit or your most painful miss. One resolved market says almost nothing. Instead, keep a log: write down the probability before the event, note what happened, and after a few dozen entries check two things — are your buckets calibrated, and is your Brier score drifting down over time? That is the difference between feeling like a good forecaster and being one.
The cheapest place to build that track record is a market where the stakes are points, not money. On Outcomer you can price real questions, watch them resolve, and score your own calls with virtual currency — long enough to see whether your 70% calls really come home seven times in ten.