Evidence
A prediction that cannot be scored was never a prediction
Most public forecasting is built to be unfalsifiable, and the small amount of structure needed to make it scorable is also what makes it worth listening to.
By Samar Bhatia4 min read

The three parts a forecast needs
A statement about the future can be evaluated only if it says what will happen, by when, and how confident the forecaster is. Drop any of the three and there is nothing to grade. Most commentary drops at least two, which is why the same voices can be wrong indefinitely without ever being visibly wrong.
The confidence term is the part people find strangest, because it feels like hedging. It is the opposite: a forecast with a number attached commits the speaker to a rate at which they should be wrong, and being wrong at exactly that rate is the definition of doing it properly. Saying something "could" happen commits to nothing at all and can’t be scored even in principle.
Resolution criteria, written in advance
The second requirement is that somebody could look at the world afterwards and agree on whether it happened. This sounds trivial and is where most attempts collapse, because interesting questions are vague: whether an economy did badly, whether a project succeeded, whether a relationship improved. Each needs a stated source and a threshold before the fact.
Writing the resolution criteria is genuinely difficult, and the difficulty is informative. When you find you cannot specify what would count as the event occurring, you have discovered that the claim you were about to make was not about the world in a checkable way. That is worth knowing before you build anything on it.
The criteria also have to be fixed. Adjusting them after the outcome is known is the forecasting equivalent of choosing the analysis after seeing the data, and it produces the same false impression of accuracy.
Two different things a good forecaster does
Accuracy has two separable components, and confusing them causes most of the bad arguments about who is any good. Calibration is whether your stated probabilities match the observed rates: of the things you said were seventy per cent likely, roughly seventy per cent should happen. Discrimination is whether you distinguish between the things that happened and the things that didn’t — whether you were pushing towards the extremes when justified.
A forecaster who says fifty per cent to everything is perfectly calibrated on a question set that resolves half the time, and useless. A forecaster who is bold but consistently overconfident discriminates well and is badly calibrated. Both properties are needed, and scoring rules exist that reward them together.
How scoring rules work, in outline
The standard approach compares the probability you assigned to what actually happened, and penalises distance from the outcome. Squared-error scoring of this kind has an important property: your best strategy is to state the probability you actually believe. Overstating confidence is punished when you are wrong more than it is rewarded when you are right, and hedging towards the middle costs you when you were justified in being bold.
Rules with that property are called proper, and the propriety is the whole reason to use one. A scoring system without it rewards gaming — saying what scores well rather than what you think — and a forecasting record produced under such a system isn’t evidence about anybody’s judgement.
Any illustration of the arithmetic here would need invented figures, so treat the mechanism rather than the numbers as the point: the score falls as the assigned probability approaches what happened, and rises steeply as it approaches the opposite.
The comparison that gives a score meaning
A score in isolation says nothing, because questions differ enormously in difficulty. A forecaster looks brilliant on a set of easy questions and mediocre on hard ones, and the same person can appear in both states in the same year depending on what they were asked.
So a record needs a benchmark: the same questions answered by the base rate, by a simple rule, or by other forecasters. Beating the base rate consistently across a varied question set is a real achievement and a rare one. Without the comparison, a published accuracy figure is decoration.
What this changes about listening to predictions
The practical filter is short. Does the claim name an event, a date and a probability? Would two people agree on whether it had happened? Is there a record of the speaker’s previous claims of the same kind, and was it kept by somebody other than them?
Almost nothing in public commentary survives that filter, which is not an argument for ignoring commentary — much of it’s explanation rather than prediction and should be judged differently. It is an argument for not treating confident description of the future as evidence about the future. And it is an argument for keeping your own record, since the only forecasting ability you can actually verify is the one you have written down.
Common questions
Editor, Think Twice Today
Samar joined to cover biases, choices, risk and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.





