Biases
One strong impression contaminates every judgement that follows
Ratings of unrelated qualities move together because an overall feeling forms first and the specific assessments are read off it rather than made independently.
By Varun Krishnan3 min read

The observation that started it
Early work on performance ratings noticed something odd in the numbers: when supervisors rated people on several supposedly independent qualities, the ratings correlated far more than the qualities plausibly did. Somebody judged high on one dimension was judged high on nearly all of them, including dimensions the rater had no real evidence about. The pattern was named the halo effect, and it has been observed in rating data ever since.
The mechanism is not flattery. It is that a global impression forms quickly and cheaply, and the specific questions afterwards are answered by consulting that impression rather than by retrieving separate evidence for each. Where no evidence exists for a particular dimension, the global feeling fills the gap silently, and the resulting rating looks as confident as any other.
Why it is so hard to catch
The judgements do not feel derived. Asked why you rated somebody highly on reliability, you will produce reasons, and the reasons will be real memories — but they were retrieved because they fit the impression, which is confirmation bias operating downstream of the halo. The introspective report is genuine and the causal story it tells is wrong.
It also has a structural advantage: it is usually invisible in the output. A set of ratings contaminated by halo looks like a coherent, consistent assessment, which is exactly what a reader hopes to see. Inconsistent ratings, the sign that somebody actually evaluated each dimension separately, look like carelessness.
Order matters more than it should, too. Whatever information arrives first tends to set the impression, and later information is interpreted in its light rather than weighed against it. Two identical dossiers read in opposite orders can produce different conclusions, which is one reason serious assessment processes fix the sequence in advance and give every case the same one.
Where the damage concentrates
Anywhere a judgement of one thing is used as evidence about another. Attractiveness and vocal confidence influence ratings of competence. A well-designed document raises assessments of the reasoning inside it. A company that has performed well financially gets described as having a strong culture, clear strategy and good leadership — descriptions that would have been written differently, about the same firm and the same people, had the results gone the other way.
That last case is worth dwelling on, because it means much of the published account of why organisations succeed is halo working backwards from the outcome. The attributes aren’t measured independently and then linked to performance; they are inferred from performance and then offered as its cause.
The standard corrective, and its cost
The reliable structural fix is to break the judgement into dimensions, gather evidence for each one separately, and score each before forming or discussing an overall view. Independence is doing the work here: once the global impression exists, dimensional scoring becomes decoration.
In group assessment the same principle means collecting individual judgements before discussion rather than after. This is unpopular, because discussion feels like the part where quality is added, and the group converges to a shared impression which then contaminates every remaining dimension at once. A consensus reached early isn’t stronger evidence than one reached late; it’s usually weaker, because it was reached with fewer independent readings.
The cost is real: this is slower, it produces uncomfortable inconsistencies, and it feels bureaucratic to people who trust their own read. That resistance is the main reason the fix is known and rarely implemented.
What the effect does not mean
It doesn’t mean that impressions are worthless. Global judgements formed by experienced people carry real information, and in domains with fast, clear feedback they can be very good. The problem is specifically the transfer of confidence from a dimension where somebody has evidence to dimensions where they have none.
It also does not mean that correlated ratings are always contaminated. Some qualities genuinely travel together, and a person who is diligent about one thing is often diligent about another. Distinguishing true correlation from halo requires evidence gathered independently, which is precisely what the halo prevents, and this is why the effect is more often reasoned about than measured in any particular case.
What follows is modest but worth having. Treat any assessment that is uniformly positive or uniformly negative across unrelated dimensions as weaker evidence than it appears, ask which dimensions the assessor actually observed, and give more weight to the one or two they can describe in specifics. Detail is the closest available marker of a judgement that was made rather than inherited.
Common questions
Deputy editor, Think Twice Today
Varun writes the explanatory pieces on biases, choices, risk and would rather show the working than assert the conclusion.





