Thought Toys · Chance & inference · Exhibit 98
Four small data sets. Same averages, same spread, same correlation, same fitted line — every number a statistics course teaches. Then look at them.
Eleven readings, plotted a reading the fitted line
—
—
Each set is eleven readings of two things: something you measured across the bottom, something you measured up the side. Underneath sit the six numbers you would normally report — the average of each, how spread out each is, how strongly they move together, and the straight line fitted through them.
Switch between the four sets. Watch the numbers. They do not move. The averages hold at 9 and 7.5, the spreads hold at 11 and 4.13, the correlation holds at 0.82, and the fitted line stays at exactly y = 3 + 0.5x in all four. If those six numbers were all you had — and in a report, a paper or a dashboard, they usually are — you would swear you were looking at the same data four times.
Now watch the picture instead. Set I is an ordinary noisy relationship, the one everybody imagines when they hear “correlation 0.82”. Set II is not noisy at all: it is a perfect arch. There is a flawless rule connecting these two quantities, it is simply bent, and the straight line misses it in a way no summary number will ever mention. Set III is a perfect straight line with one point sitting high above it; take that single reading away and the correlation is exactly 1. Set IV is the alarming one: ten of the eleven readings were taken at the very same setting, so they say nothing whatever about a trend, and the entire line — its slope, its correlation, its apparent confidence — rests on one lone point far to the right. Delete that point and no line can be drawn at all.
Then take a reading and drag it. The meter tells you how close your arrangement still is to Anscombe's six numbers. It is hard to keep them, which is the honest surprise waiting at the end: these four sets are not accidents anyone stumbled on. Anscombe built them, on purpose, to make a point that had to be made because people would not otherwise believe it.
The point is not that summary statistics lie. They do exactly what they promise: they compress. A compression throws things away, and these four sets are four different things that survive the compression looking identical. The only reliable defence is the cheapest possible one. Plot it.
improve/verify/98-anscombe.js), and both halves of the claim are
proved, not just the famous one. The estimators are checked first on data whose answers are known — a
perfect line must give r = ±1, a symmetric V must give exactly 0, and the fitted line must
really be the one that minimises squared error, shown by beating a nearby slope. Then the agreement is computed
from the raw numbers, at the precision Anscombe published rather than a rounder one: sets III and IV sit at
variance 4.123, which a careless “all within 0.005 of each other” test would fail by a hair.
The difference is measured too, with statistics blind to the shared summary: a quadratic fit
leaves set II a residual of 2×10−5 against set I's 12.9, so II really is an
exact parabola; set III without its outlier returns r = 1.00000; and every remaining
x in set IV is identical once its one distant point is dropped.
Four negative controls. The match meter must reject a single reading nudged 0.5, and must still
reject a swap of two readings that deliberately preserves the mean — so it is not simply accepting
everything. A sign-flipped fit must score far worse than the real one. And the intuition this exhibit exists to
break — “identical summaries mean identical data” — is refuted by measurement: every set
sits at least 0.67 standard deviations from set I on average and 4.1 standard deviations away at its
farthest point, on data whose six summaries all agree.
Also in Chance & inference: The Galton board →
← the cabinet · Thought Toys — a cabinet of explorable explanations. Exhibit 98.