Correlation Is Not Causation
Also known as: cnc, correlation
Treating a statistical association as evidence of cause without ruling out confounders, reverse causation, selection, or chance.
Share: also:
In plain terms
Two measurements move together. When one goes up, the other tends to go up as well, or reliably goes down. That is a correlation, and it's a fact about the data. It isn't yet a fact about the world.
A correlation between A and B is consistent with several different worlds. A causes B. B causes A. Some third factor drives both, which statisticians call a confounder. The sample was assembled in a way that manufactured the link. Or the pattern is noise, and once you measure enough variables, some pairs move together for no reason at all.
This entry is the statistical treatment. False cause is the general form of the error, and post hoc is the version that reads sequence as cause. What follows is the machinery underneath them: what a confounder actually does, what controlling for one buys you, and why randomized trials and natural experiments carry weight that an observed association never does.
Why it matters
A correlation narrows the field of explanations. It doesn't pick one. Reporting an association and then reaching for a causal verb ("raises", "improves", "leads to") promotes one candidate over the others without doing the work that promotion requires.
The work is ruling out alternatives, and there are known ways to do it. Randomized assignment does it by force: if a coin flip decides who gets the treatment, the two groups differ only by chance, so no confounder can be hiding in who signed up. Where randomizing is impossible, a natural experiment can do the same job, using a lottery, an eligibility cutoff, or a rule that sorts people for reasons unrelated to the outcome. Statistical controls help too, but only for the confounders you thought to measure. The ones you never recorded stay inside the estimate, invisible.
Short of that, the case gets built from many angles at once: a plausible mechanism, a dose-response pattern where more exposure means more effect, the right order in time, consistency across different populations, and an association that survives once the obvious confounders are accounted for. Any one of these alone is weak. Together they're how smoking was established as a cause of lung cancer without anyone running a randomized trial on it.
Canonical example
"Women on hormone therapy have much lower rates of heart disease, so hormone therapy protects the heart."
Observational studies found that association, and the association was real. But women taking hormone therapy also tended to be wealthier, more active, and more likely to see a doctor regularly. Every one of those predicts lower heart disease on its own. When large randomized trials finally assigned the treatment by chance instead of by who chose it, the protective effect on heart disease didn't show up, and some arms of the trials found elevated risk.
The confounder wasn't exotic. It was simply that healthier women were more likely to be taking the drug. Collecting more observational data would never have fixed it, because every new observation carried the same distortion baked in.
Counter-example (not a fallacy)
"The tutoring program was oversubscribed, so the district handed out places by lottery. A year later, the students who won a place scored higher than the ones who didn't. The program works."
Nobody ran an experiment here, and the causal claim still holds. The lottery cut the link between who received tutoring and who was already likely to improve. Motivated parents, stronger students and better-resourced families landed on both sides of the draw in roughly equal numbers, so none of them explain the gap. What's left is the tutoring.
The line: did something break the link between who ended up in each group and who was already headed for that outcome, or did people sort themselves in?
How to fix it
If you've been linked here, start by asking how people ended up in the groups you're comparing. If they chose, the choice is probably correlated with the outcome, and that's your confounder. The fix comes in two sizes. The small one is to downgrade the verb: "associated with" is a claim you can defend, "causes" usually isn't. The large one is to go find evidence that rules something out, a randomized trial, a natural experiment, a dose-response pattern, a mechanism that explains how A would produce B. Name the rival explanations yourself and say which ones you've eliminated. That reads as confidence, not retreat.
If you're on the receiving end, ask the two questions that do most of the work: "What else could produce that same pattern?" and "How did people end up in each group?" A solid causal claim has answers ready. A correlation dressed as one usually doesn't.