Wisdom of Crowds
Why collective estimates are often accurate, the conditions that accuracy depends on, and when crowds fail.
A large group of people with no special expertise can, under the right conditions, produce a collective estimate more accurate than almost any of its individual members, experts included. The observation is known as the wisdom of crowds. As the intellectual foundation of prediction markets, it is the reason a price formed by thousands of independent traders is worth reading as a probability. The effect is real and has been measured for over a century, and it is conditional. It holds when a specific set of requirements is met and degrades in recognizable ways when they break. Knowing both halves is what separates reading a consensus critically from trusting it blindly.
What is the wisdom of crowds?
The founding observation comes from a livestock fair. At an exhibition in Plymouth, England, in 1906, visitors entered a weight-judging competition. A live ox was on display, and anyone could submit a written estimate of what it would weigh once slaughtered and dressed. Around 800 fairgoers took part, many of them with no professional knowledge of cattle. The statistician Francis Galton collected the cards and worked through the numbers. The middle estimate came to 1,207 pounds, and the ox weighed 1,198. The collective judgment landed within one percent of the truth, closer than the large majority of individual guesses that produced it.
Galton published the analysis in Nature in 1907 under the title Vox Populi, the voice of the people, framing it as a test of popular judgment: "In these democratic days, any investigation into the trustworthiness and peculiarities of popular judgments is of interest." The result did not show that anyone in the crowd was individually skilled; it showed that their errors pointed in different directions, some high and some low, so the misses largely cancelled in the aggregate while the shared kernel of truth remained.
That is the wisdom of crowds in its simplest form. Combine many independent judgments, and the collective estimate tends to be more accurate than most of the individuals who contributed to it, often including the recognized experts. The pattern has since been reproduced across many kinds of estimation tasks. The phrase itself comes from James Surowiecki's book The Wisdom of Crowds, which assembled the evidence and, more usefully, specified the conditions a group must meet before its collective judgment deserves that trust.
Prediction markets are the financial application of the idea. Participants trade contracts tied to a defined event, and the price that emerges functions as the group's live, continuously revised estimate of how likely the event is. The mechanics of contracts and venues are covered in the Prediction Markets guide; this guide stays underneath them, on why aggregated judgment works at all and when it stops working.
What conditions make the wisdom of crowds work?
Surowiecki's analysis distilled the requirements into four conditions, and they remain the standard checklist:
- Diversity of judgment. Members hold different information, models, and perspectives, so their errors point in different directions rather than piling up on one side.
- Independence. Each judgment is formed without copying, anchoring on, or deferring to others.
- Decentralization. Members draw on local and specific knowledge rather than a single shared source.
- Aggregation. Some mechanism exists to combine the individual judgments into one collective output.
Each condition maps to a piece of the statistics behind the effect. Treat every estimate as the truth plus an individual error. When errors are independent and varied, they tend to cancel in the aggregate, and what survives the averaging is the signal the estimates share. When errors are correlated, they survive too, and the collective estimate inherits the group's shared bias.
The role of diversity can be stated exactly. The diversity prediction theorem, formulated by the social scientist Scott Page, ties the quantities together for averaged numeric estimates:
In plain language, the crowd's squared error equals the average member's squared error minus the spread of the estimates. This is an identity rather than a tendency. A crowd outperforms its average member by exactly the amount its members disagree, which is why diversity is load-bearing rather than decorative. A group of well-informed clones gains nothing from aggregation; a group of individually mediocre but genuinely varied estimators can be collectively excellent.
Two details keep expectations honest. The guarantee is relative to the average member, not the best one. A genuine expert can still beat the crowd, and the difficulty is identifying that expert in advance. And the diversity term rewards disagreement rooted in different information, not noise, since random guessing increases the spread and the average individual error together. What improves the collective estimate is adding judgments that are wrong in new and different ways.
How do polls, averaging, and prediction markets aggregate judgments?
The aggregation condition is a design decision, and the mechanism chosen shapes what the collective output means:
| Mechanism | How judgments enter | How they are weighted | Incentive to be accurate | Output |
|---|---|---|---|---|
| Simple averaging | Everyone submits an estimate | Equally | None built in | A single figure at one point in time |
| Opinion polls | A sampled group answers questions | By sampling design and demographic adjustment | None built in | A periodic snapshot |
| Forecasting tournaments | Registered forecasters submit probabilities | By scoring rules and track record | Scores, rankings, reputation | Forecasts revised over time |
| Prediction markets | Participants trade contracts at a price | By capital committed at each price level | A direct financial stake | A price that updates continuously |
Simple averaging is Galton's method. It is transparent and hard to manipulate, and it treats a rancher's estimate and a clerk's as equally informative. That equal weighting is a strength when nobody knows who is informed and a weakness when somebody clearly is. Averages also produce a snapshot, since nothing in the mechanism prompts a revision when the situation changes.
Opinion polls add sampling science to the same idea. A poll constructs a sample designed to represent a population, then weights responses to correct for who actually answered. Polls measure stated views at a moment in time, and question wording shapes the result. Research on election surveys found that asking respondents who they expect to win has historically produced more accurate forecasts than asking who they intend to vote for, a gap attributed to each respondent effectively summarizing the leanings of their entire social circle rather than a single data point (Rothschild and Wolfers). The expectation question quietly turns every respondent into a small aggregator.
Forecasting tournaments keep score. Participants submit explicit probabilities, accuracy is measured with scoring rules, and standing reflects a track record built across many questions. Large tournament studies found that accuracy improves further with training in probabilistic reasoning, teaming, and performance tracking (Mellers et al.). The craft that emerged from that research is the subject of the Superforecasting guide.
Prediction markets replace submission with trading. Participation is voluntary and costs capital, so the mechanism selects for people who believe they know something. Conviction is expressed through position size rather than a checked box, which weights confident, informed judgment more heavily than idle opinion, at the cost of also weighting confident error. Because a trade can happen at any moment, the aggregate revises continuously instead of waiting for the next survey wave.
Aggregation is not limited to estimates a mechanism asks for. The opinions people volunteer in public conversation can be measured and aggregated as well, and that is the territory of Sentiment Analysis.
When does the wisdom of crowds fail?
The recognizable failure modes map closely onto the conditions above, so diagnosing which condition broke indicates how heavily a consensus should be discounted.
Social influence replaces independent judgment
Independence is usually the first condition to break, and its loss has been measured directly. In a controlled experiment, Lorenz, Rauhut, Schweitzer, and Helbing had subjects answer factual estimation questions, then revise them after seeing what others had estimated. Even mild exposure produced three effects. The diversity of estimates collapsed without a matching improvement in collective accuracy, which the authors call the social influence effect. The true value migrated toward the edge of the shrinking range of estimates, the range reduction effect, so the group's spread became a misleading guide to where the truth might sit. And participants grew more confident in the collective answer even though it had not become more accurate, the confidence effect. The combination is what makes this failure dangerous. Consensus tightens and confidence rises exactly as the informational value of the consensus drains away.
Agreement is not accuracy
When members of a group can see one another's views, convergence can reflect imitation rather than independent confirmation. Before treating a tight consensus as strong evidence, it is worth asking whether those agreeing arrived at their views separately.
Information cascades
When judgments are made in sequence and each person can observe the choices of those who came before, the group can enter an information cascade, a state in which it becomes rational for each new participant to follow the visible behavior of predecessors and set aside their own private information. The founding analysis, published in 1992 by Bikhchandani, Hirshleifer, and Welch, showed that once the weight of observed behavior exceeds what any single private signal can outweigh, private information stops entering the pool entirely. From that point the crowd keeps growing while its knowledge does not.
Markets are exposed to cascade dynamics because the price itself is public information that every participant watches. Suppose a contract on a policy decision trades at 34¢ and a wave of buying lifts it to 48¢ over an afternoon. Later traders cannot see why the price moved, only that it moved, and some will buy on the inference that someone else knows something. Their buying moves the price further, strengthening the same inference for the next observer. Each step can be individually rational while the sequence detaches the price from any underlying information. Cascades are also fragile. Because their later stages rest on little private information, one piece of genuinely new public information can unwind them quickly.
Correlated errors and shared blind spots
Independence rarely fails loudly. It fails quietly, through shared inputs. When most of a crowd reads the same coverage, follows the same commentators, and reasons from the same models, errors acquire a common direction, and errors that share a direction do not cancel. Aggregation removes random error and passes systematic error straight through.
The pattern shows up in real aggregates. Employee traders in internal corporate prediction markets displayed a measurable optimism bias about their own company's projects, a shared lean that survived aggregation (Cowgill and Zitzewitz). At larger scale, a calibration study spanning hundreds of millions of trades across two major venues found persistent domain-level patterns, with political contracts clustering toward 50¢ and resolving more decisively than their prices imply. A crowd can be reliably wrong in the same direction for long stretches, and no amount of additional aggregation fixes an error that every member shares.
Thin crowds and questions nobody can answer
Aggregation needs inputs. When participation is sparse, the aggregate is a handful of opinions wearing the costume of a consensus. As one market-structure analysis put it, the wisdom of the crowd only works when you have a crowd. The threshold can sit surprisingly low when participants hold genuinely dispersed knowledge. The same corporate-market research found that even modest internal markets improved on expert forecasts by up to a 25% reduction in mean squared error. What matters tends to be the number of independent information sources feeding the aggregate rather than the raw headcount.
The quietest failure is the question nobody can answer. Aggregation combines information that exists; it cannot conjure information that does not. On questions where the decisive facts are not yet knowable by anyone, a stable consensus figure is a summary of shared ignorance, and its stability says nothing about its accuracy.
Why are prediction markets good at aggregating information?
The idea that prices aggregate knowledge predates any modern venue. In a 1945 essay, the economist F. A. Hayek argued that the central economic problem is coordinating knowledge dispersed across millions of minds, none of which holds the full picture, and that the price system performs the aggregation. A prediction market narrows that mechanism to a single question with a defined end. Contracts pay a fixed amount at resolution if the event occurs and nothing if it does not, and the price at which they trade functions as the group's working estimate of the probability. How to read that price is the subject of Prices and Probabilities.
Measured against the crowd conditions, the design holds up well. Aggregation is native, since the price is the aggregate. Diversity and decentralization arrive through self-selection, because the mechanism is open to anyone who believes they hold relevant information, whatever its source. Independence receives a reward for dissent, something plain crowds never provide. A participant who disagrees with the consensus and turns out to be right is paid by those who were wrong, so standing apart from the group is compensated rather than socially punished. The incentive pushes against herding without eliminating it, because the price remains public information that everyone watches.
The track record supports the design. The canonical survey of the field by Wolfers and Zitzewitz concluded that prices of winner-take-all contracts can be interpreted as event probabilities and that market forecasts have outperformed moderately sophisticated benchmarks. The longest-running evidence comes from the academically operated Iowa Electronic Markets. Across five presidential cycles, the market's forecast was closer to the outcome than 964 contemporaneous polls about 74% of the time, with the advantage largest months before the election rather than on its eve.
Where the wisdom actually comes from remains an active research question. A working paper analyzing the universe of transactions on one large venue attributes much of its accuracy to a small minority of persistently skilled traders, on the order of a few percent of participants, whose profits are funded by the losses of the rest. That reading differs from the classical one, yet both describe aggregation. In one, the market averages away independent noise; in the other, it operates as a discovery mechanism that finds whoever holds real information and lets their capital move the price. Either way, the advantage over a hand-picked expert panel is the same. Nobody has to decide in advance who is informed, because the mechanism discovers them.
None of this makes a market price self-certifying. Practitioner analyses of when prices deserve trust read like the crowd conditions translated into trading terms: precisely defined outcomes, reasonably quick and probable resolution, limited hidden information, and genuine sources of disagreement to sustain two-sided trading. A price from a market that meets those conditions is a serious probability estimate; a price from one that does not is a thin consensus expressed in cents, and the failure modes above describe how to discount it.
The wisdom of crowds is the theory underneath every prediction market; the guides below cover the machinery built on top of it.