Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Risk

The best and worst performers are usually just the smallest ones

Rates measured on small groups swing wildly for arithmetic reasons alone, so any league table sorted by rate fills both ends with the same kind of place.

By Varun Krishnan4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A property of averages that nobody finds intuitive

The average of a small sample bounces around far more than the average of a large one. This isn’t a subtle statistical claim; it follows directly from the fact that one unusual case is a large share of a small group and a negligible share of a big one. Everyone accepts it when stated abstractly and almost nobody applies it when reading a table of results.

The consequence for rankings is specific and severe. If you rank units by some rate — outcomes per head, incidents per year, success per attempt — and the units differ in size, the top and bottom of the list will be dominated by the small units, because they are the ones capable of producing extreme rates. The middle of the table will be crowded with the large ones, whose numbers cannot move much.

The pattern that fooled a generation of policy

A well-known illustration comes from studies of institutional performance where small units appeared to do unusually well. Programmes were built on that observation, on the reasonable assumption that something about being small was responsible. Then somebody looked at the other end of the same table and found small units there too, in disproportionate numbers, doing unusually badly.

Both ends were the same phenomenon. Size was not causing good performance or bad performance; it was causing variability. The distribution of rates is wider for small samples, so the extremes of any ranking are populated by them, regardless of what is actually being measured. That the pattern is symmetric is the tell, and looking at both tails is the diagnostic.

Why the top of a table is the wrong place to look for lessons

A unit that reaches the top of a rate ranking has usually had both some genuine quality and some luck, and the smaller the unit the larger the luck component has to be to get there. Studying the leaders therefore means studying a group selected partly for having had a good year, which is a selection effect masquerading as a research method.

This is also why leaders of small units so often fall back the following period. Nothing has gone wrong. The quality persists and the luck resets, which is a purely statistical movement that gets narrated as complacency, or as the burden of expectation, or as some other story about people.

What to ask of any ranking

Three questions strip most of the noise out. How many events does each unit’s rate rest on — not how many people, but how many of the things being counted. Are the units comparable in size, and if not, does the ranking adjust for it. And what does the bottom of the table look like, since a tail dominated by the same category as the top is a sign that variance, rather than performance, is doing the sorting.

A fourth question is about the interval. A rate estimated from few events comes with a wide range of plausible values, and when those ranges are drawn on a chart, the apparent ordering of most rankings dissolves into a band within which nobody is distinguishable. Publications that show the interval look far less dramatic and are far more honest.

The same trap in ordinary judgement

It appears whenever a rate is computed on a handful of instances. A shop rated on a few reviews, a treatment judged on a few cases, a colleague evaluated on three projects, a route judged fast from two journeys — all of these produce rates with enormous swing, and the human mind treats a rate as a rate regardless of what it was computed from.

The rule of thumb that helps is to think in counts rather than percentages whenever the denominator is small. A success rate of two-thirds sounds like a fact. Two out of three sounds like what it is, which is almost no information at all.

What this does not prove

None of this means that small units are all mediocre or that rankings are useless. Real differences exist, and with enough events behind each estimate they show up clearly and stably. The claim is narrower: the extremes of a rate table, computed on small numbers, are mostly a measurement artefact, and any explanation offered for them is being fitted to noise.

The practical version is to look for consistency across periods rather than position in one. Something that stays near the top of a ranking year after year is unlikely to be doing so by luck. Something that appears at the top once, from a small base, has told you almost nothing, and building a programme on it’s how a decade of well-intentioned effort gets spent chasing a statistical property.

Common questions

How many events are enough for a rate to be meaningful?

There is no single threshold, since it depends on how large a difference you are trying to detect. The useful habit is to look at the count of events rather than the population size, and to distrust any rate resting on a handful of them regardless of how large the underlying group is.

Is this the same as regression to the mean?

They are closely related. Small samples produce extreme measurements, and extreme measurements are followed by less extreme ones — the first explains why the extremes appear, the second explains why they do not persist.

How should a fair ranking be presented?

With the count of events, an interval around each estimate, and ideally some adjustment that pulls small-sample estimates towards the overall average. Presented that way, most tables show far fewer real differences than the ordered list implied.

Risksample sizerankingsvariancestatistics
Varun Krishnan
Deputy editor, Think Twice Today

Varun writes the explanatory pieces on biases, choices, risk and would rather show the working than assert the conclusion.

Read next

More risk →