Evidence
The thing being measured is rarely the thing anyone cares about
Research and management both run on proxies, and most disputes about what a number means are really disputes about the gap between the proxy and the quantity underneath.
By Rohan D’Souza3 min read

Nearly every important quantity is measured indirectly
Wellbeing, competence, risk, quality, understanding, health: none of these can be read off an instrument. What gets recorded is something correlated with them and easier to observe — a questionnaire response, a test score, a count of incidents, a time to completion. The measure is a stand-in, and the entire value of the resulting number depends on how good a stand-in it is.
This is not a criticism of measurement. Proxies are unavoidable and often excellent. The failure is in forgetting the substitution happened, after which the proxy quietly becomes the definition and every subsequent argument is conducted about the wrong quantity.
The two ways a proxy fails
It can miss part of what you care about, which makes it incomplete. A test that measures recall but not application still measures something real; it simply leaves out a component, and decisions made on it will systematically favour whatever it happens to capture.
It can also include things you do not care about, which makes it contaminated. A measure of activity picks up genuine effort and also picks up busyness, reporting behaviour and whoever happens to be enthusiastic about filling in forms. Contamination is usually the more dangerous of the two, because it produces movement in the number that has no counterpart in the world.
Self-report is the most-used instrument and the least examined
A large share of what is known about behaviour and experience rests on people describing themselves. This isn’t worthless — for internal states like mood or pain, self-report is arguably the criterion rather than a proxy for it. For behaviour it is much weaker, and comparisons between reported and directly observed behaviour tend to find systematic gaps, generally in the flattering direction.
The gaps are structured rather than random, which is what makes them a problem. Memory for how often something happened is reconstructed rather than counted, socially approved behaviour is over-reported, and answers shift with how the question is worded and who is asking. Any finding that depends entirely on self-reported behaviour should be read with that in mind, and studies that validate self-report against an objective measure are doing something genuinely valuable.
What happens when a proxy becomes a target
Once a measure is used to allocate rewards, effort moves towards the measure. If the proxy and the underlying quality are tightly linked, that’s exactly what was wanted. If they are loosely linked, the effort flows into the gap, and the number improves while the thing it stood for does not.
The general observation that a measure under pressure stops being a good measure is old and well supported by experience across many fields. The mechanism is not usually cheating. It is ordinary prioritisation by people responding to what is counted, doing the parts of the job that show up and dropping the parts that do not, which is behaviour the system explicitly asked for.
How to check a measure before trusting it
Ask what the quantity of interest actually is, in a sentence that does not mention the measurement. Then ask what would make the measure move without the quantity moving, and what would make the quantity move without the measure moving. Both lists are usually easy to produce and both are usually uncomfortable.
A third check is whether the measure has ever been compared against anything else. A proxy validated against an independent measure — even imperfectly — carries far more weight than one adopted because it was available. In many fields nobody has done that comparison, and the measure is in use because it was in use last year.
A fourth is to look at what the measure does at the extremes. Many proxies track the quantity well in the middle of the range and badly at the ends, which matters because the ends are where decisions get made — the worst cases, the outstanding ones, the thresholds that trigger action.
Living with imperfect measurement
The response to all this is not to abandon numbers for impressions, which are proxies too and considerably less examined. It is to use several measures that fail in different ways, to treat agreement between them as meaningful and disagreement as a question rather than an error, and to keep a written description of what the numbers are supposed to represent.
And to accept a limit: some things that matter are measured badly and will continue to be. Acknowledging that a decision rests partly on judgement is more honest than dressing the judgement in a metric, and it leaves the reasoning open to challenge, which the metric does not.
Common questions
Features writer, Think Twice Today
Rohan writes the explanatory pieces on biases, choices, risk and would rather show the working than assert the conclusion.





