Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Risk

A probability about a single event needs a reference class before it means anything

Numbers attached to one-off events are shorthand for a class of similar cases, and until that class is named the figure cannot be checked, compared or acted on.

By Adrian Novak3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A number that has to be about more than one thing

Saying an event has a seventy per cent chance of happening looks like a description of that event. It cannot be, because the event will either happen or not, and no observation of the single case can establish that the figure was right or wrong. The number has to be understood as a claim about a class of situations of which this is one, and the class is almost never stated.

This is not pedantry, because the choice of class changes the answer. The same person can be a member of many groups with very different rates, and each of those rates is a correct answer to a slightly different question. Nothing in the number itself tells you which question was being answered.

The forecast everybody quotes and nobody reads carefully

A chance of rain is the standard example. Depending on the issuing service, it can mean the probability that measurable precipitation falls at a given point in the forecast area, the expected proportion of the area receiving rain, or a combination of the two. Surveys of public interpretation have repeatedly found that people read it in several different ways, including as the fraction of the day it will rain or the confidence of the forecaster.

All of those readings are coherent. Only one matches what the issuer meant, and which one that is depends on the issuer. A number can be perfectly well defined at the source and functionally ambiguous by the time it reaches a decision, which is a communication failure rather than a technical one.

Choosing the class is a judgement, not a formality

The general problem is that any individual case belongs to indefinitely many classes. A project is one of your projects, one of this type of project, one run by this team, one attempted in this economic climate. Rates differ across all of them, and there is no procedure that identifies the correct class from first principles.

What guides the choice is relevance and sample size, which pull against each other. A narrow class matches the case more closely and contains fewer observations, so its rate is noisier; a broad class is better measured and less similar. Most sensible practice moves between the two, checking whether the answer is stable across reasonable choices, and treating a figure that swings wildly as a signal that the reference class is doing more work than the evidence supports.

Verification is what gives a single-event number content

Since one case cannot test a probability, the meaning has to come from the forecaster’s record across many. If everything a person calls seventy per cent happens about seventy per cent of the time, the numbers are carrying information regardless of how any individual call turned out. That is what calibration means, and it is measurable given enough forecasts.

This is why unscored probabilities are close to worthless and why a forecaster with a track record is worth much more than a confident one without. It also means the honest response to a single-event figure from an unknown source is not disagreement but a question: what happened the last hundred times this source said seventy per cent?

How this differs from the individual and population distinction

It is worth separating this from the related point that a probability which is negligible for one person is a reliable count for a population. That concerns two different uses of an agreed number. What is at issue here is prior to that: which set of cases the number was computed over in the first place.

The two interact. A figure derived from a broad population and then applied to a specific individual is a reference class choice, and whether it is a good one depends on how similar that individual is to the population on the features that matter. Disagreements that look like arguments about probability are frequently arguments about membership.

Three questions worth asking of any such figure

What set of cases is this the rate for, and would I recognise a member of that set? How many cases were in it, and were they observed or modelled? And has the source been scored before, on numbers of this kind?

Where the answers are unavailable, the number is not useless, but it should be treated as a rough expression of somebody’s confidence rather than as a measurement. That is a meaningful downgrade, and it is one that the appearance of a percentage tends to hide.

Common questions

Can a single-event probability be wrong?

Not verifiably, from the single event. It can be assessed only as part of a set of similar claims by the same source, which is why calibration is measured over many forecasts rather than judged case by case.

What does a chance of rain actually mean?

It varies by issuing service. Common definitions concern the probability of measurable precipitation at a point in the forecast area, sometimes combined with the expected coverage of that area, and studies of public understanding find several different readings in circulation.

How do I choose a reference class?

Balance similarity against sample size and check whether the answer is stable across a few defensible choices. If it moves a great deal depending on which class you pick, that instability is the most important thing you have learned.

Riskprobabilityforecastinginterpretationevidence
Adrian Novak
Staff writer, Think Twice Today

Adrian joined to cover biases, choices, risk and stayed for the awkward questions and would rather show the working than assert the conclusion.

Read next

More risk →