Small Sample Size
Also known as: ss2, small-n, sample-size
Small groups produce more extreme results by chance alone, so they crowd both the top and the bottom of any ranking.
Share: also:
In plain terms
Averages taken from small groups bounce around more than averages taken from large ones. That's not a flaw in the data, it's arithmetic. Flip a fair coin ten times and you'll get eight or more heads about 5 percent of the time. Flip it a thousand times and you will not see 800 heads, not once in the age of the universe. Same coin, same odds, wildly different range of outcomes.
So when you sort any collection of groups by their average result, the small ones pile up at both ends. They're not better or worse. They're noisier, and noise is what the extremes of a ranking select for.
This differs from hasty generalization and anecdotal evidence, which are about not having enough cases to support a broad claim. Those are sufficiency problems, and the answer is usually "gather more". This one is a directional distortion. It bites even when your dataset is enormous, as long as you rank subgroups inside it by outcome, because ranking by outcome quietly ranks by smallness.
Why it matters
The variability of an average shrinks with the square root of the sample size, not the sample size itself. Compare a school with 25 students in a grade against one with 2,500. The larger school has 100 times the students, so its average score is 10 times more stable. The small school's average will swing across years for no reason other than which handful of kids happened to enroll.
Which means: if you hand out awards to the top performers in any ranking, you're mostly handing out awards for being small. Then you look at the bottom of the same list and the small units are sitting there too, which should be the clue. A real quality effect points one direction. Variance points both.
The consequence is expensive. In the 1990s and 2000s, researchers noticed that small schools were heavily overrepresented among the highest-achieving schools in the United States, and a major philanthropic push went into breaking large schools into smaller ones. Small schools were also overrepresented among the worst performers. The pattern was mostly sample size. Howard Wainer and Harris Zwerling laid out the arithmetic, and the same shape turns up in county-level disease maps, where the counties with the lowest cancer rates and the counties with the highest rates are both the tiny rural ones.
Canonical example
"Six of the top ten schools in the state have fewer than 300 students. Small schools work. Let's fund more of them."
Read the other end of the list. If small schools also fill six of the bottom ten slots, the ranking is measuring enrollment, not education. A grade of 40 students has one average; move two exceptional students in or out and the whole school's number jumps. A grade of 400 barely notices.
The reasoning has a second cost beyond the wasted money. The genuinely excellent small schools get lumped in with the lucky ones, so nobody learns what the good ones were doing right.
Counter-example (not a fallacy)
"We ranked all 380 schools, then plotted each school's score against its enrollment. The spread narrows as enrollment rises, exactly as chance predicts. Three schools sit well outside that envelope. Those three are worth studying."
This is the same ranking, used correctly. The analysis asks how much variation sample size alone would produce, draws that boundary, and only treats results outside it as signal. Small units can absolutely be excellent, and this method can detect it. What it won't do is confuse a noisy average with a good one.
The line: does the result stand out beyond what a group that size would produce by chance, or is it just small?
How to fix it
If you've been linked here, look at the other end of your ranking before you draw a conclusion from the top of it. If the same kind of unit occupies both extremes, you're looking at variance rather than performance. The fix is to attach sample sizes to every figure in the table and check whether the gap between first and last is larger than random variation would produce at those sizes. Where the samples are small, pull the estimates toward the overall average, which is what statisticians do when they can't tell noise from signal. And be careful with year-over-year swings in small units, because a group that spikes one year will usually drift back toward normal the next, no intervention required.
If you're on the receiving end, ask for the denominator and then ask about the bottom of the list. "How many students, patients, or customers is that based on, and what size are the worst performers?" Two questions, and most rankings-driven conclusions come apart on the second one.