Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Biases

One strong impression contaminates every judgement that follows

Ratings of unrelated qualities move together because an overall feeling forms first and the specific assessments are read off it rather than made independently.

By Varun Krishnan3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The observation that started it

Early work on performance ratings noticed something odd in the numbers: when supervisors rated people on several supposedly independent qualities, the ratings correlated far more than the qualities plausibly did. Somebody judged high on one dimension was judged high on nearly all of them, including dimensions the rater had no real evidence about. The pattern was named the halo effect, and it has been observed in rating data ever since.

The mechanism is not flattery. It is that a global impression forms quickly and cheaply, and the specific questions afterwards are answered by consulting that impression rather than by retrieving separate evidence for each. Where no evidence exists for a particular dimension, the global feeling fills the gap silently, and the resulting rating looks as confident as any other.

Why it is so hard to catch

The judgements do not feel derived. Asked why you rated somebody highly on reliability, you will produce reasons, and the reasons will be real memories — but they were retrieved because they fit the impression, which is confirmation bias operating downstream of the halo. The introspective report is genuine and the causal story it tells is wrong.

It also has a structural advantage: it is usually invisible in the output. A set of ratings contaminated by halo looks like a coherent, consistent assessment, which is exactly what a reader hopes to see. Inconsistent ratings, the sign that somebody actually evaluated each dimension separately, look like carelessness.

Order matters more than it should, too. Whatever information arrives first tends to set the impression, and later information is interpreted in its light rather than weighed against it. Two identical dossiers read in opposite orders can produce different conclusions, which is one reason serious assessment processes fix the sequence in advance and give every case the same one.

Where the damage concentrates

Anywhere a judgement of one thing is used as evidence about another. Attractiveness and vocal confidence influence ratings of competence. A well-designed document raises assessments of the reasoning inside it. A company that has performed well financially gets described as having a strong culture, clear strategy and good leadership — descriptions that would have been written differently, about the same firm and the same people, had the results gone the other way.

That last case is worth dwelling on, because it means much of the published account of why organisations succeed is halo working backwards from the outcome. The attributes aren’t measured independently and then linked to performance; they are inferred from performance and then offered as its cause.

The standard corrective, and its cost

The reliable structural fix is to break the judgement into dimensions, gather evidence for each one separately, and score each before forming or discussing an overall view. Independence is doing the work here: once the global impression exists, dimensional scoring becomes decoration.

In group assessment the same principle means collecting individual judgements before discussion rather than after. This is unpopular, because discussion feels like the part where quality is added, and the group converges to a shared impression which then contaminates every remaining dimension at once. A consensus reached early isn’t stronger evidence than one reached late; it’s usually weaker, because it was reached with fewer independent readings.

The cost is real: this is slower, it produces uncomfortable inconsistencies, and it feels bureaucratic to people who trust their own read. That resistance is the main reason the fix is known and rarely implemented.

What the effect does not mean

It doesn’t mean that impressions are worthless. Global judgements formed by experienced people carry real information, and in domains with fast, clear feedback they can be very good. The problem is specifically the transfer of confidence from a dimension where somebody has evidence to dimensions where they have none.

It also does not mean that correlated ratings are always contaminated. Some qualities genuinely travel together, and a person who is diligent about one thing is often diligent about another. Distinguishing true correlation from halo requires evidence gathered independently, which is precisely what the halo prevents, and this is why the effect is more often reasoned about than measured in any particular case.

What follows is modest but worth having. Treat any assessment that is uniformly positive or uniformly negative across unrelated dimensions as weaker evidence than it appears, ask which dimensions the assessor actually observed, and give more weight to the one or two they can describe in specifics. Detail is the closest available marker of a judgement that was made rather than inherited.

Common questions

Is there a negative version?

Yes, and it works identically. One poor impression drags down ratings of unrelated qualities, and it tends to be stickier, because the negative impression discourages the further contact that might have corrected it.

Does knowing the person longer help?

It helps if the additional contact produces independent evidence on separate dimensions, and it makes things worse if it mainly produces more occasions to confirm the existing impression. Length of acquaintance by itself is not a safeguard.

Why not just average several people’s overall judgements?

Averaging helps only when the judgements are independent. If the assessors discussed the case first, or all read the same summary, their errors are correlated and the average inherits the shared impression rather than cancelling it.

Biaseshalo effectassessmenthiringjudgement
Varun Krishnan
Deputy editor, Think Twice Today

Varun writes the explanatory pieces on biases, choices, risk and would rather show the working than assert the conclusion.