Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Risk

A test that is right most of the time can still be mostly wrong

When the thing being detected is rare, the great majority of positive results can be false without anything at all being defective about the test.

By Tara Mukherjee3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two numbers that get confused

Every imperfect test has two accuracy figures pulling in different directions, and almost every argument about testing comes from treating them as one. The first is how often the test flags a case that really is what it is looking for. The second is how often it flags a case that is not. Both are properties of the test itself, measured on cases whose true status is already known.

What a person actually wants to know is a third thing entirely: given that this particular result came back positive, how likely is it that the thing is really present? That number is not a property of the test. It depends on the test and on how common the condition is in the group being tested, and it can swing wildly while the test itself stays exactly the same.

Working it through with counts

Take a hypothetical group of ten thousand people, and suppose the thing being detected is present in one of them in every thousand — ten people in total. Suppose the test catches nearly all real cases, so it flags nine or ten of them. Now suppose it also wrongly flags one in every hundred people who do not have it, which sounds like a very accurate test. That false-alarm rate applies to the nine thousand nine hundred and ninety people who are fine, producing about a hundred false flags.

So roughly a hundred and ten people receive a positive result and about ten of them have the condition. The great majority of positives are wrong. These figures are invented purely to show the shape of the arithmetic and they describe no real test, but the structure holds for any rare target: the false-alarm rate is multiplied by an enormous group and the hit rate is multiplied by a tiny one.

Notice what did the damage. Not inaccuracy — the test in this example is good. Rarity did it.

Why the intuition fails

The mind hears that a test is accurate and converts that into a statement about the result in hand. That conversion is the error, and it is the same one that appears throughout probability reasoning: swapping the chance of the evidence given the condition for the chance of the condition given the evidence. The two are different quantities and they are only similar when the condition is common.

Presenting the problem in counts, as above, rather than in percentages and conditional probabilities, improves performance markedly. That is one of the more robust findings in the study of statistical reasoning and it has been shown with professionals as well as with students. Nobody is bad at this because they are careless. They are bad at it because the standard notation hides the denominators.

What follows in practice

The first consequence is that testing an entire population for something rare produces a large volume of false alarms in absolute terms, whatever the test’s quality, and any system doing so has to plan for that volume rather than treat it as a defect. The second is that a positive result from a broad screen usually means look again with a different method, not conclude.

The third is about who is tested. Applying the same test to a group where the condition is much more common — because of a specific reason for suspicion — changes the arithmetic in favour of the positive result, sometimes dramatically. That is why targeted testing gives more informative positives than universal testing, using the identical test. It is also why deciding whom to test is a substantive decision and not an administrative one.

This is a description of arithmetic, not advice about any particular test. Anything concerning health belongs with a clinician who knows the specific test, the specific person and the reason for testing, and the numbers used above are illustrative only.

The same logic outside the laboratory

The pattern is not confined to medicine. Automated fraud detection, security screening, plagiarism detection, content moderation and hiring filters all involve applying an imperfect indicator to a population where the target is rare, and all of them therefore generate far more false positives than true ones unless something else narrows the population first.

The important design consequence is that the cost of a false positive has to be kept low, because there will be many of them. Systems that attach a heavy penalty to being flagged, with no cheap route to review, will impose that penalty mostly on people who did nothing. That is not a failure of the algorithm. It is a failure to read the arithmetic before deploying it.

Common questions

Does a second test fix the problem?

It helps substantially if the second test is independent of the first, because the group being tested the second time now has a much higher prevalence. If the two tests share a mechanism, they may share their errors too, and the improvement is smaller than it looks.

Why are the numbers above described as invented?

Because they are chosen to make the structure visible, not to describe any real test. Using real figures would require naming a specific test and population, and the point being made is about the shape of the calculation rather than any particular case.

Is a negative result more trustworthy?

For a rare condition, usually yes, and for the same reason in reverse: almost everyone tested does not have it, so a negative result agrees with a very strong prior. The exception is a test that misses many real cases, where a negative carries much less weight.

Risktestingfalse positivesbase ratesscreening
Tara Mukherjee
Contributing editor, Think Twice Today

Tara writes the explanatory pieces on biases, choices, risk and thinks most subjects are more interesting once you know how they work.

Read next

More risk →