Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Choices

A crude model usually beats an expert combining the same information by feel

When the inputs are fixed and the task is to combine them, mechanical combination outperforms human judgement with unusual consistency, and the reason is not that the model is clever.

By Gautam Pillai3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two ways of putting the same facts together

Consider a judgement with a small number of relevant inputs — a hiring decision, an estimate of how long something will take, a forecast of whether a project will hold together. One approach is to look at everything and form an overall impression, weighing the factors as the case seems to require. The other is to write down the factors in advance, assign each a weight, and add them up.

The second sounds crude, and the first sounds like what expertise is for. The comparison between them has been run in a great many domains over several decades, and the result is one of the more consistent findings in the whole of applied psychology: the mechanical combination equals or beats the holistic one most of the time.

The finding, and the shape of the evidence

The comparison has an unusually clean design. The same information is given to both, so the model has no informational advantage; the only difference is how the pieces are combined. Reviews pooling many such comparisons have found mechanical prediction ahead in a clear majority of studies, roughly tied in a minority, and behind in a small number.

It is worth being precise about the size of the claim. The advantage is not enormous in every domain, and there are areas where judgement holds its own. What is remarkable is the direction and the consistency — a broad literature in which a very simple procedure keeps winning against trained people who have every reason to be good at the task.

Consistency is doing almost all the work

The explanation is not that the model is insightful. It is that the model is consistent, and people are not. Present a professional with the same case twice, separated enough that they do not recognise it, and the two judgements often differ. Mood, order, fatigue and irrelevant features of presentation all move the output, and none of that variation carries information about the case.

A formula cannot do that. It applies the same weights every time, so it captures whatever valid signal the inputs contain without adding noise on top. This also explains a stranger result in the same literature: models built with roughly equal weights, or with weights derived from the expert’s own stated policy, often perform about as well as statistically optimised ones. Getting the direction of each factor right matters far more than getting the weights exactly right.

What the expert is still required for

Almost everything upstream. Deciding what the question is, deciding which factors belong in the model, defining them so two people would score them the same way, noticing when a case is not the kind of case the model was built for, and interpreting the output for a decision. The comparison is narrow: it concerns combining given inputs, not the whole judgement.

Human judgement is also the only available source for many inputs. A structured interview rating is a human assessment, and a good one is valuable — it simply performs better as a number entering a formula than as a component of an overall impression formed in the room.

The exception clause is real, and it is overused

The standard objection is that a formula cannot know about the unusual fact — the case where something obvious and unmodelled changes everything. The objection is correct in principle. The difficulty is that people identify such cases far more often than they occur, and each override reintroduces exactly the inconsistency the formula was removing.

A practical compromise is to permit overrides and to record them: what the model said, what was done instead, why, and how it turned out. That converts the exception from an untracked privilege into something with a record attached, and a record is the only way anybody ever discovers whether their overrides help.

What the comparison does not settle

It does not settle whether a model should be used in a given setting. A formula trained on past decisions inherits the patterns in those decisions, including ones nobody would endorse if they were stated. It also assumes the relationship it was fitted to is still the relationship that holds, which is a strong assumption in any environment that is changing.

And it says nothing about acceptability. Being assessed by an arithmetic rule feels different from being assessed by a person, and that feeling is a real cost in contexts where people have to accept the outcome. The finding is about accuracy, which is one input to a decision about method rather than the whole of it.

Common questions

Does this apply to any decision?

It applies to repeated judgements with identifiable inputs and outcomes that can eventually be scored. One-off decisions in novel situations have neither the data to build a rule nor the repetition to make consistency valuable.

Why do simple equal weights work so well?

Because most of the achievable accuracy comes from including the right factors with the right sign, and because equal weights cannot be overfitted to a small sample. Precise weights estimated from limited data often perform worse out of sample than crude ones.

Should experts be replaced by formulas?

That is not what the finding supports. Experts choose the factors, produce many of the inputs, and recognise when a case falls outside the rule. What the evidence argues against is having them combine the inputs by impression when a consistent rule is available.

Choicespredictionexpertiseconsistencydecisions
Gautam Pillai
Reporter, Think Twice Today

Gautam writes about biases, choices, risk, mostly the parts other people skip and reads the small print so you do not have to.