Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Risk

A one-in-a-million figure was calculated, not observed

Numbers for very rare events cannot come from counting them, so they come from models, and the assumptions inside the model deserve more scrutiny than the digits do.

By Tara Mukherjee3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

You cannot wait long enough to measure it

To observe the rate of an event that happens once in a million opportunities, you need a great many million opportunities, and for the events people care most about — a structural failure, a catastrophic malfunction, a rare adverse outcome — that quantity of experience does not exist. The number therefore cannot be a frequency read off a record. It has to have been produced some other way.

This is worth stating because a probability presented as a bare figure looks like a measurement. Nothing in the notation distinguishes a rate counted from ten million observations from one derived through a chain of assumptions, and the second is far more common in exactly the range where the numbers sound most authoritative.

The three ways such a number gets made

The first is decomposition. Break the failure into components, use measured rates for each component, and combine them according to an assumed structure of how the system fails. The component rates may be well established; the structure is a model of the system, and models are simplifications by definition.

The second is extrapolation. Fit a curve to the events you have observed and read off the value far out in the tail where you have none. This is respectable statistical practice and it depends entirely on the fitted shape continuing to hold in a region where nothing was measured. The third is elicitation: ask people who know the domain and combine their judgements. That is often the only option available, and it inherits everything known about human calibration in the very small range.

Completeness is the assumption that fails

The characteristic failure of a decomposed estimate is not that a component rate was wrong. It is that a failure path was not in the model at all. An estimate produced by summing the paths somebody thought of is bounded below by the paths nobody thought of, and those are exactly the ones that show up in accounts of things going badly wrong.

A related weakness is that the components are usually treated as failing separately, when a shared cause can take several out together. That correlation problem is a large topic on its own; here the point is narrower. Both weaknesses push the estimate in the same direction, which is why calculated tail probabilities tend to be too small rather than too large.

The uncertainty can be larger than the estimate

When a figure results from multiplying several uncertain quantities, the uncertainty compounds, and the plausible range for the answer can easily span more than one order of magnitude. A result quoted as one in a million might be defensible anywhere between one in a hundred thousand and one in ten million on the same evidence.

That is not a criticism of the analysis. It is a description of what the analysis can support, and the failure is in the reporting rather than the method. A single figure with no range attached implies a precision that the chain of assumptions never had, and this is the point at which a careful estimate turns into a misleading one.

What the number is still good for

A great deal, provided it is used comparatively. Estimates built the same way, with the same assumptions, are useful for ranking: this design is safer than that one, this path contributes most of the risk, this component is where an improvement would pay. The systematic errors partly cancel when the comparison is internal.

What such a number will not support is a claim about absolute safety, or a comparison with a rate measured a completely different way. Setting a calculated tail probability alongside an observed frequency from a large record and treating them as the same kind of quantity is the most common misuse, and the two are not commensurable however similar they look on the page.

Reading a very small number

Ask how it was produced. If the answer is a model, ask what the model contains and, more importantly, what would have to be missing from it for the figure to be badly wrong. Ask whether the estimate has ever been compared with subsequent experience, since some domains have accumulated enough events to check the older predictions.

And treat the figure’s stability as the main signal. An estimate that changes by a factor of ten when a reasonable person adjusts one assumption is telling you where the real uncertainty lives. That is the useful output of the exercise, rather than the number printed at the end of it.

Common questions

Are calculated tail probabilities useless?

No, they are useful comparatively — for ranking designs, finding which paths dominate, and deciding where improvement pays. What they do not support is a confident absolute claim, or a direct comparison with a rate measured in an entirely different way.

Why do such estimates tend to be too small?

Because the two main weaknesses point the same way: failure paths that were never modelled contribute nothing to the total, and components assumed to fail independently can be taken out together by a shared cause.

What should be reported alongside the figure?

A range, the assumptions the result is most sensitive to, and how the estimate was produced. A single number with no range implies a precision that a chain of multiplied uncertainties cannot deliver.

Riskprobabilitymodelsestimatesuncertainty
Tara Mukherjee
Contributing editor, Think Twice Today

Tara writes the explanatory pieces on biases, choices, risk and thinks most subjects are more interesting once you know how they work.

Read next

More risk →