Choices
A crude model usually beats an expert combining the same information by feel
When the inputs are fixed and the task is to combine them, mechanical combination outperforms human judgement with unusual consistency, and the reason is not that the model is clever.
By Gautam Pillai3 min read

Two ways of putting the same facts together
Consider a judgement with a small number of relevant inputs — a hiring decision, an estimate of how long something will take, a forecast of whether a project will hold together. One approach is to look at everything and form an overall impression, weighing the factors as the case seems to require. The other is to write down the factors in advance, assign each a weight, and add them up.
The second sounds crude, and the first sounds like what expertise is for. The comparison between them has been run in a great many domains over several decades, and the result is one of the more consistent findings in the whole of applied psychology: the mechanical combination equals or beats the holistic one most of the time.
The finding, and the shape of the evidence
The comparison has an unusually clean design. The same information is given to both, so the model has no informational advantage; the only difference is how the pieces are combined. Reviews pooling many such comparisons have found mechanical prediction ahead in a clear majority of studies, roughly tied in a minority, and behind in a small number.
It is worth being precise about the size of the claim. The advantage is not enormous in every domain, and there are areas where judgement holds its own. What is remarkable is the direction and the consistency — a broad literature in which a very simple procedure keeps winning against trained people who have every reason to be good at the task.
Consistency is doing almost all the work
The explanation is not that the model is insightful. It is that the model is consistent, and people are not. Present a professional with the same case twice, separated enough that they do not recognise it, and the two judgements often differ. Mood, order, fatigue and irrelevant features of presentation all move the output, and none of that variation carries information about the case.
A formula cannot do that. It applies the same weights every time, so it captures whatever valid signal the inputs contain without adding noise on top. This also explains a stranger result in the same literature: models built with roughly equal weights, or with weights derived from the expert’s own stated policy, often perform about as well as statistically optimised ones. Getting the direction of each factor right matters far more than getting the weights exactly right.
What the expert is still required for
Almost everything upstream. Deciding what the question is, deciding which factors belong in the model, defining them so two people would score them the same way, noticing when a case is not the kind of case the model was built for, and interpreting the output for a decision. The comparison is narrow: it concerns combining given inputs, not the whole judgement.
Human judgement is also the only available source for many inputs. A structured interview rating is a human assessment, and a good one is valuable — it simply performs better as a number entering a formula than as a component of an overall impression formed in the room.
The exception clause is real, and it is overused
The standard objection is that a formula cannot know about the unusual fact — the case where something obvious and unmodelled changes everything. The objection is correct in principle. The difficulty is that people identify such cases far more often than they occur, and each override reintroduces exactly the inconsistency the formula was removing.
A practical compromise is to permit overrides and to record them: what the model said, what was done instead, why, and how it turned out. That converts the exception from an untracked privilege into something with a record attached, and a record is the only way anybody ever discovers whether their overrides help.
What the comparison does not settle
It does not settle whether a model should be used in a given setting. A formula trained on past decisions inherits the patterns in those decisions, including ones nobody would endorse if they were stated. It also assumes the relationship it was fitted to is still the relationship that holds, which is a strong assumption in any environment that is changing.
And it says nothing about acceptability. Being assessed by an arithmetic rule feels different from being assessed by a person, and that feeling is a real cost in contexts where people have to accept the outcome. The finding is about accuracy, which is one input to a decision about method rather than the whole of it.
Common questions
Reporter, Think Twice Today
Gautam writes about biases, choices, risk, mostly the parts other people skip and reads the small print so you do not have to.





