Any outcome that depends partly on skill and partly on luck will show this pattern: the most extreme results are extreme partly because the luck ran one way, and next time the luck is fresh. So exceptional results are followed by less exceptional ones, and terrible results are followed by less terrible ones, with nobody having done anything.

Francis Galton found it measuring the heights of parents and children and called it regression toward mediocrity. It is the single most reliable source of false causal stories in business, in sport, and in management.

The mechanism, stated carefully

It is not a force pulling things toward average, and it is not a law of averages — nothing is owed and nothing is being corrected.

It is an arithmetic consequence of the fact that any single result is signal plus noise. The more extreme the result, the larger the share of it that was probably noise, simply because large noise is required to produce a large deviation. Take a second measurement and the signal is still there and the noise is redrawn. The second result therefore sits closer to the signal, which is to say closer to average.

Which also gives the condition: the effect is proportional to how much luck is involved. Where outcomes are almost entirely skill, there is almost no regression. Where they are mostly luck, an extreme result tells you almost nothing about the next one. Most business outcomes are somewhere in the middle and are treated as though they were at the skill end.

Why it produces bad management

Kahneman's flight-instructor example is the cleanest demonstration.

Instructors observed that praising a cadet after an exceptional landing was followed by a worse one, and criticizing after a poor landing was followed by a better one. They concluded that criticism works and praise does not, and they taught accordingly for years.

Both sequences are exactly what regression predicts with no effect from either intervention. But the instructors were watching real events in the right order and drawing the obvious inference — and because they never stopped criticizing, they never saw the comparison that would have told them otherwise.

This runs constantly. A bad quarter triggers an intervention, the next quarter is better, and the intervention is credited. A salesperson has an extraordinary month, gets promoted, and looks worse in the new role. A consultant is brought in after the worst period and things improve. In each case the regression explanation is available, costs nothing, and is almost never the one offered.

How to tell the difference

Look at the whole series, not the two endpoints. Regression is invisible in a before-and-after comparison and obvious in a run of numbers.

Ask what would have happened anyway. Not a rhetorical question — it has a defensible answer. If the measure is noisy and the starting point was extreme, the expected next value is closer to the mean before anyone does anything.

Check whether the intervention is always applied at extremes. If action is only ever taken after a bad month, the design guarantees an apparent improvement whether or not the action does anything. This is the same structural problem that makes uncontrolled before-and-after studies nearly worthless.

Count the reverse cases. If the process works, it should also work when applied after an ordinary month. Nobody ever tries it then, which is the tell.

The version that costs real money

Hiring and selection. Candidates are chosen on the strength of an extreme signal — a remarkable recent result, an exceptional interview — and both are noisy measurements. The subsequent performance regresses, the new hire is judged a disappointment, and the selection process is not examined, because the story available is about the person.

The same structure explains why acquisitions of businesses bought on a peak year so often disappoint, and why the year used to set the price is almost always the best one available.

None of this argues against acting on evidence. It argues for using more than one observation, because a single extreme measurement is the least informative kind there is, and it is the kind that most reliably provokes a decision.