Ecological Fallacy
Also known as: ef, ecological
Concluding something about an individual from a statistic that describes the group they belong to.
Share: also:
In plain terms
A group average is a summary of many people at once. It is not a description of any one of them. The country with the highest average income contains plenty of people with no income at all, and knowing the national figure tells you close to nothing about the person in front of you.
The ecological fallacy is the step from the group number to the individual claim. It shows up as "their region votes this way, so they probably do," or "that school scores well, so this graduate must be strong," or "our enterprise accounts are worth more, so this one is."
The averages can be perfectly accurate and the inference still fails. What gets lost in an average is the spread, and the spread is usually where the individual lives.
Why it matters
Correlations between groups are often much stronger than correlations between people, and sometimes they point the opposite way. The sociologist William Robinson demonstrated this in 1950 using the 1930 US census. Across the 48 states, the share of residents who were foreign-born correlated positively with literacy, around +0.53. At the individual level, being foreign-born correlated negatively with literacy, around -0.11. Immigrants had settled in states where literacy was already high among everyone else. The state-level number described the states accurately and described the people backwards.
The practical damage is in decisions about individuals. Underwriting, hiring, sentencing, admissions, and credit all involve someone reasoning from a group statistic to a person, and the same arithmetic applies every time: a large difference in group means can coexist with distributions that overlap almost completely. When two groups overlap that much, the group label carries very little information about any specific member.
The mirror error exists too. Generalising from a handful of individuals to a claim about the whole population is hasty generalization in its usual form. Group data and individual data answer different questions, and neither substitutes for the other.
Canonical example
"Our enterprise customers spend four times what small-business customers spend. This lead is an enterprise account, so it's worth four times as much. Give it to the senior rep."
The average is real. The inference is not. Enterprise spending is usually driven by a small number of very large accounts, which drags the mean far above what a typical enterprise customer actually spends. The median enterprise account might be worth well under twice a small-business one, and this particular lead could be anywhere in the distribution.
Same shape, different setting: "the average household in that postcode earns £90,000, so they can afford the premium tier." Averages are pulled around by their tails. Ask for the median and the spread before you attach a number to a person.
Counter-example (not a fallacy)
"Average household income in the district is £90,000, up from £70,000 five years ago. We're opening a second store there."
This is a group statistic supporting a group-level decision. The store isn't serving one household, it's serving the aggregate, and aggregate purchasing power is exactly what the average describes. No individual claim is being made, so no individual claim can be wrong.
Using group data as an explicit starting probability is also legitimate, provided it stays a starting point. "Accounts from this segment renew 80% of the time, so we'll forecast this one at 80% until we learn something specific" is honest, because it's stated as a rate rather than a fact about the customer, and it updates the moment real information arrives. What makes it work is the willingness to be overridden.
The line: is the conclusion about the group, or about a person in it? Group in, group out is fine. Group in, individual out needs individual evidence.
How to fix it
If you've been linked here, ask what the spread looks like, not just the average. If most of the variation sits within the groups rather than between them, the group label barely narrows anything down and the argument needs different evidence. When you only have group data and a decision to make about a person, use it as a prior and say so out loud: "the base rate says X, and I'll revise as soon as I know something about this case." That's a defensible position. Stating the group figure as though it described the individual is not.
If you're on the receiving end, ask about the overlap: "How much do the two groups overlap?" Most people have never looked, and the answer tends to be "almost entirely," which resolves the argument faster than naming the error would.