Thought Toys · Data & inference · Exhibit 48

The wisdom of crowds

Nobody can see inside the jar, so nobody guesses right. Average enough independent guesses anyway, and the crowd lands close to the truth — a real, provable law. But let the guessers hear each other first, and the magic quietly stops working.

Every guess for "how many jellybeans?" — the crowd, one dot each truthcrowd average

The crowd's error as more guessers join (log scale) theorythis jar

your turn — drag ρ up and watch the crowd stop improving

What you're seeing

Nobody here can see inside the jar. Each guesser eyeballs it and blurts out a number — some too high, some too low, all of them wrong in their own particular way. Look at the spray of dots: barely any land near the true count. But average the whole spray together, and the amber mark lands remarkably close to the truth — closer than nearly every individual guesser managed alone. That is the real result behind "the wisdom of crowds": independent errors point in unrelated directions, so they tend to cancel rather than add up.

Drag the crowd size up with ρ at 0 and the error curve keeps falling — more guessers, steadily less noise, no floor in sight. Now drag ρ up first: every guesser is nudged by the same shared hunch — a loud first guess, a headline number, a friend's opinion overheard in line — and the crowd's errors quit cancelling. Watch the error curve go flat. Past that point, adding guessers barely helps, no matter how many more you add: a thousand people who all half-heard the same rumor are barely wiser than ten.

The rule, exactly. Guesser i's relative error is εi = σ(√ρ·Z0 + √(1−ρ)·Zi) where Z0 is one shared "herd" shock felt by everyone and Zi is each guesser's own independent noise (both standard normal). The crowd's error is the average of the εi, and its spread works out to sd = σ·√(ρ + (1−ρ)/N) — at ρ=0 that is the familiar σ/√N, shrinking to zero as N grows; at ρ>0 it floors at σ·√ρ, whatever N becomes. Verified in node (improve/verify/48-wisdom-of-crowds.js): the empirical spread matches this formula at both ρ=0 and ρ>0 across a spread of crowd sizes; the crowd still beats a clear majority of individuals either way; and the shared-factor generator's own pairwise correlation measures out to ρ, as designed. Negative control: at ρ=0.12, growing the crowd 6× (100→600 guessers) shrinks the error by only ≈4% — while the identical 6× growth at ρ=0 cuts the error by ≈60%, so the stall is caused specifically by the herding, not some generic large-crowd effect.

Also in Data & inference: Nobody got worse. The luck just didn't show up twice. →

All 14 in Data & inference
  1. 09Bayes' theorem
  2. 104You always land in the long gap.
  3. 110Five serial numbers. Now guess how many they built.
  4. 113In enough dimensions, nothing is near anything
  5. 121The faster the rating learns, the less it knows
  6. 28Simpson's paradox
  7. 39Zipf's law
  8. 40Benford's law
  9. 48The wisdom of crowds — you are here
  10. 86Nobody got worse. The luck just didn't show up twice.
  11. 87Your friends really do have more friends than you.
  12. 96Plan for the average and you'll be wrong every time
  13. 98The numbers agree. The pictures don't.
  14. 99Unrelated in the crowd. A trade-off inside the gate.

← the cabinet · Thought Toys — a cabinet of explorable explanations. Exhibit 48.