Risk
A probability about a single event needs a reference class before it means anything
Numbers attached to one-off events are shorthand for a class of similar cases, and until that class is named the figure cannot be checked, compared or acted on.
By Adrian Novak3 min read

A number that has to be about more than one thing
Saying an event has a seventy per cent chance of happening looks like a description of that event. It cannot be, because the event will either happen or not, and no observation of the single case can establish that the figure was right or wrong. The number has to be understood as a claim about a class of situations of which this is one, and the class is almost never stated.
This is not pedantry, because the choice of class changes the answer. The same person can be a member of many groups with very different rates, and each of those rates is a correct answer to a slightly different question. Nothing in the number itself tells you which question was being answered.
The forecast everybody quotes and nobody reads carefully
A chance of rain is the standard example. Depending on the issuing service, it can mean the probability that measurable precipitation falls at a given point in the forecast area, the expected proportion of the area receiving rain, or a combination of the two. Surveys of public interpretation have repeatedly found that people read it in several different ways, including as the fraction of the day it will rain or the confidence of the forecaster.
All of those readings are coherent. Only one matches what the issuer meant, and which one that is depends on the issuer. A number can be perfectly well defined at the source and functionally ambiguous by the time it reaches a decision, which is a communication failure rather than a technical one.
Choosing the class is a judgement, not a formality
The general problem is that any individual case belongs to indefinitely many classes. A project is one of your projects, one of this type of project, one run by this team, one attempted in this economic climate. Rates differ across all of them, and there is no procedure that identifies the correct class from first principles.
What guides the choice is relevance and sample size, which pull against each other. A narrow class matches the case more closely and contains fewer observations, so its rate is noisier; a broad class is better measured and less similar. Most sensible practice moves between the two, checking whether the answer is stable across reasonable choices, and treating a figure that swings wildly as a signal that the reference class is doing more work than the evidence supports.
Verification is what gives a single-event number content
Since one case cannot test a probability, the meaning has to come from the forecaster’s record across many. If everything a person calls seventy per cent happens about seventy per cent of the time, the numbers are carrying information regardless of how any individual call turned out. That is what calibration means, and it is measurable given enough forecasts.
This is why unscored probabilities are close to worthless and why a forecaster with a track record is worth much more than a confident one without. It also means the honest response to a single-event figure from an unknown source is not disagreement but a question: what happened the last hundred times this source said seventy per cent?
How this differs from the individual and population distinction
It is worth separating this from the related point that a probability which is negligible for one person is a reliable count for a population. That concerns two different uses of an agreed number. What is at issue here is prior to that: which set of cases the number was computed over in the first place.
The two interact. A figure derived from a broad population and then applied to a specific individual is a reference class choice, and whether it is a good one depends on how similar that individual is to the population on the features that matter. Disagreements that look like arguments about probability are frequently arguments about membership.
Three questions worth asking of any such figure
What set of cases is this the rate for, and would I recognise a member of that set? How many cases were in it, and were they observed or modelled? And has the source been scored before, on numbers of this kind?
Where the answers are unavailable, the number is not useless, but it should be treated as a rough expression of somebody’s confidence rather than as a measurement. That is a meaningful downgrade, and it is one that the appearance of a percentage tends to hide.
Common questions
Staff writer, Think Twice Today
Adrian joined to cover biases, choices, risk and stayed for the awkward questions and would rather show the working than assert the conclusion.





