Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Risk

Every alarm sets a trade between two kinds of error

Any detector can be made to miss less only by flagging more, and the choice of where to put that threshold is a judgement about which mistake you would rather make.

By Tara Mukherjee3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two errors, and you can’t minimise both

Any system that decides whether something is present makes two kinds of mistake: flagging when nothing is there, and staying quiet when something is. Improving the underlying detector reduces both. Moving the threshold does not — it trades one directly for the other, and the trade is unavoidable rather than a sign of poor design.

This is the distinction that most arguments about alarms, filters and screens fail to make. Complaints about false alarms are often answered with promises of better accuracy, when what is actually being proposed is a threshold change that will cause more misses. Those are different interventions with different costs, and conflating them means the real decision never gets discussed.

The threshold is a value judgement wearing technical clothes

Where to set it depends entirely on what each error costs, and those costs are rarely symmetric. A missed structural failure and an unnecessary inspection are not comparable quantities. A wrongly rejected application and a wrongly accepted one land on different people, in different amounts, at different times.

Because the two costs are usually borne by different parties, the threshold is also a distributive decision. Tightening a filter to reduce one kind of error pushes the burden onto whoever suffers the other kind, and the people who set thresholds are frequently not among them. This is why arguments about detection systems become political so quickly, and why treating the setting as a purely technical parameter is a way of avoiding the argument rather than resolving it.

Rarity makes the arithmetic brutal

When the thing being detected is rare, most flags are wrong however good the detector is, because the false-alarm rate is applied to a very large population of ordinary cases and the hit rate to a tiny population of real ones. That consequence is arithmetic and can’t be designed away.

What it implies for a threshold is specific: in a rare-target setting, small movements towards greater sensitivity produce large absolute increases in false alarms, because the population supplying them is enormous. The same movement in a setting where the target is common barely changes the alarm volume at all. The identical adjustment is cheap in one context and ruinous in the other, which is why thresholds can’t sensibly be copied between applications.

What happens to the people receiving the alarms

A system producing many false alarms trains the people monitoring it to discount them, and the discounting is a rational response rather than negligence. If nearly every alert has turned out to be nothing, treating the next one as probably nothing is correct reasoning from the observed record, and reprimanding somebody for it misdiagnoses the problem entirely.

That response is well recognised in safety-critical fields and has been implicated in serious incidents in several industries. The design implication runs against instinct: adding more alerts to a system that already produces too many reduces the attention paid to all of them, so the marginal alarm can lower the overall detection rate. More warning is not monotonically safer.

Making the trade visible before deploying anything

A useful discipline is to state, in advance and in words, what each error costs and who pays it, before any threshold number is chosen. Then choose the threshold to reflect that statement rather than to make an evaluation metric look good, since a single accuracy figure hides the trade completely and can be improved by moving in either direction.

The second discipline is to plan for the volume rather than treating it as a defect to be fixed later. If a screen will produce many false positives, then the cost of being flagged has to be kept low and the route to review has to be cheap and quick, because most of the people bearing that cost will have done nothing. Systems that attach heavy consequences to a flag are implicitly claiming a precision that rarity makes impossible.

A reader’s checklist

When somebody reports how well a detector works, ask which of the two error rates is being quoted, since a system can be described as highly accurate on either one alone. Ask how common the target is in the population being screened, because that determines what a flag actually means. Ask what happens to a flagged case, and how easily an error is corrected.

And ask who chose the threshold, and what they were optimising. That question is answerable surprisingly often, and the answer explains more about a system’s behaviour than any amount of description of the underlying method.

Common questions

Can a better detector avoid the trade-off?

It shifts the whole trade-off in your favour — every threshold setting gets better — but it does not remove it. At any given quality of detection, reducing misses still means accepting more false alarms, which is why the two improvements should be discussed separately.

Why is a single accuracy figure misleading?

Because it combines both error types into one number that depends on how common the target is. In a rare-target setting a detector that never flags anything scores very well on overall accuracy while being completely useless, which is why the two rates should always be reported separately.

Is alarm fatigue a failure of the operators?

Generally not. Discounting alerts that have almost always been false is a correct inference from experience. The failure is upstream, in a threshold that produces more alarms than the attention available can service, and the remedy has to be a design change rather than an exhortation.

Riskthresholdsdetectionfalse alarmsdesign
Tara Mukherjee
Contributing editor, Think Twice Today

Tara writes the explanatory pieces on biases, choices, risk and thinks most subjects are more interesting once you know how they work.

Read next

More risk →