Regression to the Mean

Also known as: rtm, regression

Extreme results tend to be followed by ordinary ones for purely statistical reasons, and that drift gets mistaken for an effect.

Share: also:

Structural diagram of Regression to the Mean. An extreme reading is followed by an ordinary one, with or without treatment.
An extreme reading is followed by an ordinary one, with or without treatment.

In plain terms

Most things we measure are part skill and part luck. A quarter that goes badly, a month of migraines, a shooting slump, a record sales week: each combines something stable about the person or process with whatever the conditions happened to be that time.

Pick out the most extreme cases and you have selected, by construction, the ones where the luck ran hardest in one direction. Luck doesn't repeat on cue. So the next measurement usually sits closer to ordinary, with nothing having caused the change.

Francis Galton named the effect in 1886 after noticing that unusually tall parents had tall children who were, on average, closer to the population mean than their parents were. He called it regression towards mediocrity. The children weren't shrinking. The extremes were just less extreme the second time around.

Why it matters

Anything applied because a measurement was extreme will appear to work. Patients see a doctor when symptoms peak, so most treatments look effective. Companies bring in consultants after the worst quarter, so most turnarounds look successful. Schools intervene with the lowest-scoring students, so most interventions look promising. In each case some of the improvement was going to happen anyway.

It cuts the other way too, and this is where it does real damage. Anything applied after an unusually good result will appear to backfire. The Sports Illustrated cover jinx is the folk version: athletes make the cover after a career-best run, then perform worse. No jinx is needed. The cover is awarded for the outlier, and outliers don't hold.

The pattern is easy to miss because the causal story is always available and always more satisfying than "that was partly noise." If you want a rule of thumb: whenever a group was chosen for being extreme, expect movement toward average and don't hand out credit for it until you've ruled it out.

Canonical example

Daniel Kahneman describes teaching flight instructors in the Israeli Air Force about the value of rewarding good performance. One experienced instructor pushed back. In his experience, cadets praised for an outstanding manoeuvre flew worse on the next attempt, while cadets chewed out for a bad one flew better. His conclusion: criticism works, praise doesn't.

The instructor's observations were accurate. His explanation was not. An outstanding manoeuvre is an outlier, and the next attempt tends to be closer to the cadet's usual standard, praise or no praise. Same for a terrible one. The instructors had been rewarding and punishing noise for years, and the statistics had rewarded them with what looked like consistent proof that harshness works.

Kahneman's point, in Thinking, Fast and Slow, is that the world had set up an experiment guaranteed to teach the wrong lesson. Praise reliably preceded decline; criticism reliably preceded improvement. Anyone paying close attention would conclude exactly what the instructor concluded.

Counter-example (not a fallacy)

"We took the 200 patients with the highest blood pressure and randomly assigned half to the new drug and half to placebo. Both groups improved. The drug group improved by 14 points more."

Regression to the mean is fully present here, and the conclusion is still sound. Everyone in the study was selected for an extreme reading, so everyone drifted back toward normal, including the placebo group. Because both arms were selected the same way, the drift affects both equally and cancels out of the comparison. What's left is the 14-point difference.

This is a large part of why control groups exist. Without one, the study reports the drug's effect plus the regression and calls the total the effect. With one, the regression sits in both columns and drops out of the subtraction.

The line: did the comparison group start from the same extreme? If yes, the difference between them is real. If there is no comparison group, some unknown share of the improvement was going to arrive on its own.

How to fix it

If you've been linked here, check one thing: was the group you're studying picked because its numbers were unusual? Worst-performing stores, top scorers, the sickest patients, the month everything went wrong. If so, part of the change you're crediting to your intervention was coming regardless, and the honest version of the claim needs a comparison, either a control group or a similarly extreme group that got nothing. Absent that, you can still report the improvement. You just can't attribute all of it.

If you're on the receiving end, ask how the group was selected before arguing about whether the intervention works: "Were these picked because they were the worst performers?" It's a neutral question, and if the answer is yes, the burden shifts on its own without anyone having to say the word fallacy.