Market Accuracy
Why a prediction market's accuracy is measured across many prices and outcomes, not any single one.
A prediction market can price an event at 90%, and the event can then fail to occur. A newcomer reasonably asks whether the market was wrong, and the honest answer is that one outcome cannot settle it. A market price is a probabilistic forecast, and a probability is not proven right or wrong by a single event, any more than a 90% chance of rain is refuted by one dry afternoon. Accuracy, for a prediction market, is a property of many prices judged against many outcomes. It can be measured, it has been measured for decades, and the record is strong in some settings and genuinely weaker in others. This guide covers what accuracy means for a probabilistic forecast, how it is measured, what the research record shows, where market prices have beaten polls and experts, and the conditions under which they stop deserving trust.
What does it mean for a prediction market to be accurate?
Every contract resolves in one direction. It settles at its full value if the event occurs and at nothing if it does not, so at resolution a price of 34¢ or 90¢ collapses into a plain Yes or No. The price beforehand was something different from that binary result. It was the market's estimate of how likely the event was, and the founding survey of the field treats the price of a winner-take-all contract as exactly that, a working probability estimate (Wolfers and Zitzewitz). The Prices and Probabilities guide covers how a price is read as a probability. Judging whether that estimate was accurate is not the same as checking which way the contract settled.
A probability describes a frequency across many similar situations rather than a fact about one of them. A forecast of 90% claims that events described this way occur about nine times in ten, which is also a claim that they fail about one time in ten. When a 90¢ contract resolves No, that lone result is consistent with a perfectly accurate 90% forecast and equally consistent with a wildly overconfident one. Nothing in the single outcome separates the two. A market therefore cannot be graded on one call, and confident claims after the fact that a market was "wrong" because a favored outcome failed tend to misread what the price was ever saying.
Accuracy becomes measurable only in aggregate. Collect every contract a market priced near 90¢ and ask what share of them resolved Yes. If the answer sits close to 90%, those prices meant what they said.
Two meanings of accurate
Picking the winner and being well calibrated are different tests. A market clears the first by leaning the right way on a lopsided question, which is easy and carries little information. It clears the second only if its probabilities match how often events actually occur, which is the harder standard the research uses and the one the rest of this guide follows.
A market can call the winner and still be poorly calibrated, and it can miss a coin-flip outcome while having been perfectly calibrated all along.
How is prediction market accuracy measured?
Three measurement approaches recur in the research, and they answer different questions:
| Measure | The question it answers | How to read it |
|---|---|---|
| Calibration | Do contracts priced near 70¢ resolve Yes about 70% of the time? | Judged across many markets, never a single one |
| Brier score | How far did prices sit from outcomes, on average? | Zero is perfect; 0.25 is the coin-flip baseline |
| Comparative benchmark | Did the market land closer than a poll, model, or expert? | Only as informative as the benchmark it beats |
Calibration is the most direct of the three, because it tests the probabilities against outcomes without any external yardstick. The method is simple to state. Take a large set of resolved contracts, group them by the price they traded at, and for each group measure the share that resolved Yes. Plotted with predicted price on one axis and realized frequency on the other, a perfectly calibrated market traces the diagonal, where contracts at 30¢ resolve Yes 30% of the time and contracts at 80¢ resolve Yes 80% of the time. Forecasting platforms score their participants on the same principle (Metaculus).
When the curve bends away from the diagonal, its shape names the problem. Prices that sit too far from the middle, so that 90¢ contracts resolve Yes only 80% of the time, indicate overconfidence. Prices that huddle too close to the middle, so that events resolving Yes 80% of the time never traded above 70¢, indicate underconfidence.
A calibration curve shows the pattern but withholds a single grade. The Brier score supplies the grade, reducing a market's entire history to one number by averaging the squared distance between each price and the outcome it resolved to. Lower is better, a flawless forecaster scores zero, and a market that answered 50¢ to every question would score 0.25. The Superforecasting guide gives the formula and shows how the same score grades individual forecasters.
One number can hide two very different virtues, which a standard decomposition of the Brier score separates (scoring-rule theory):
In words, a low score requires two things at once. Reliability is calibration error, and lower is better. Resolution rewards prices that commit, moving decisively toward 0¢ or 100¢ rather than hovering near the base rate. Uncertainty is the difficulty baked into the questions, which no forecaster controls. A market that quoted the same base rate on every contract would be trivially well calibrated and completely useless, because it would have no resolution. Calibration keeps a market's probabilities honest, and resolution is what makes them worth reading.
The third approach steps outside the market and sets its forecasts beside an external benchmark such as a poll, a statistical model, or an expert panel. That comparison answers a different and more contested question, taken up two sections on.
Are prediction markets accurate in practice?
Run the calibration test on real markets and the prices hold up well, with important exceptions. A well-calibrated result looks like the pattern below, shown here with illustrative figures rather than data from any one venue:
| Price bucket | Contracts in bucket | Share that resolved Yes |
|---|---|---|
| near 10¢ | 640 | 11% |
| near 30¢ | 910 | 32% |
| near 50¢ | 1,180 | 49% |
| near 70¢ | 870 | 71% |
| near 90¢ | 720 | 88% |
Prices land close to the frequencies they imply, and the buckets hold hundreds of contracts each, which is what makes the comparison meaningful. Large studies of the major venues report roughly this pattern. An analysis of more than 100 million Polymarket trades found that prices closely track the probabilities that actually played out, and slightly outperform bookmaker odds on the same events (Reichenbach and Walther). Play-money platforms run the same test on themselves, and Manifold's public calibration chart plots its prices against realized outcomes, a record that has stayed close to the diagonal.
The usual explanation for that record is aggregation under incentives. A market pools the judgments of everyone trading, weights them by the capital each is willing to commit, and pays the participants who correct a mispricing, which tends to pull prices toward frequencies that hold up. The Wisdom of Crowds guide covers that mechanism and the conditions it depends on.
Calibration is not uniform, though. It varies with how far the event sits in the future, with the domain of the question, and with how much was traded. Those variations are where the honest limits begin, and the next sections work through them.
Are prediction markets more accurate than polls and experts?
The comparative question has the longest track record, because the natural benchmark for a market forecast is whatever it might replace.
Against polls, the deepest evidence comes from the Iowa Electronic Markets, an academic real-money market launched in 1988. Across five US presidential elections, its prices sat closer to the eventual result than 964 contemporaneous polls about 74% of the time, with the advantage largest months before the vote rather than on its eve. Raw prices and raw polls both carry biases, and once each is statistically debiased, market-based forecasts still edged out poll-based ones early in the campaign and in uncertain races, including against the poll aggregators of the day. The Markets vs Polls guide compares the two forecasting methods in full.
Against expert and consensus benchmarks, the pattern recurs. The founding survey concluded that market prices outperform moderately sophisticated forecasting benchmarks (Wolfers and Zitzewitz). More recently, Federal Reserve staff evaluated Kalshi's macroeconomic contracts against the Bloomberg survey of professional economists and found the market's forecasts had significantly smaller errors on inflation prints and were never significantly worse, while correctly identifying the most likely rate decision at every meeting studied (Diercks, Katz, and Wright). The effect even appears inside companies, where internal markets at firms including Google and Ford improved on their own experts' forecasts by up to a 25% reduction in mean squared error (Cowgill and Zitzewitz).
The advantage is not universal. When teams of elite forecasters were tested head to head against prediction markets inside the same government tournament, the forecasters came out ahead. Those markets were small and thinly traded, so the result argues that well-run aggregated judgment can rival a market rather than that markets always win, and it points at the one thing markets most depend on, a deep pool of participants.
When are prediction market prices least accurate?
The record above comes mostly from deep, liquid, heavily studied markets. The prices that most deserve suspicion are the ones formed under the opposite conditions.
| Condition | Why accuracy suffers | What to check |
|---|---|---|
| Thin trading | The price reflects a few positions, not a considered consensus | Volume and open interest behind the price |
| Extreme prices | Very cheap contracts can drift from their true frequency | How the market handles the long tail |
| Manipulation exposure | A thin book can be moved by a single funded trader | Whether a sharp move survived or reversed |
| Long horizons | Capital locked until a distant resolution keeps traders away | How far off resolution sits |
Thin trading is the first and most common. A market's forecast is only as good as the participation behind it, and a price backed by a handful of positions is a small sample wearing the costume of a consensus. A market-structure survey put it plainly: "the wisdom of the crowd only works when you have a crowd." Skeptics of political markets have long noted that thin, low-volume contracts can post confident-looking prices that a modest amount of money could move (Yale Insights).
The extremes of the price range carry their own distortion. In parimutuel racetrack pools, very cheap long-shot outcomes have historically been overbought relative to how often they occur, an effect known as the favorite-longshot bias (Snowberg and Wolfers). Whether it holds on the newer venues is genuinely unsettled. One large Polymarket study found no general bias at the market level (Reichenbach and Walther), while the complete first-generation Polymarket dataset documented the reverse pattern at the token level, with low-probability contracts overpriced (Qin and Yang). A 3¢ price is a reliable signal that the market considers something unlikely and an unreliable guide to whether the true figure is 3% or 1%.
Thin markets are also the ones most exposed to deliberate manipulation, since a book a single order can move is a book a single trader can push. How that works, and how much it actually distorts prices, is the subject of Market Manipulation.
A final case is subtler, because the price can be wrong while nothing is broken. Correcting an obvious mispricing on a long-dated contract means locking up capital until a distant resolution, and that cost keeps some traders out, so a price can sit visibly wrong for weeks. One trader's first-person account of an election market describes prices he considered clearly mistaken staying that way for months, because moving them meant tying up money in a near-certainty for an uncertain stretch of time.
Weight a price by its market
The accuracy record is strongest for deep, liquid, heavily traded markets and weakest for thin ones, yet both quote prices in the same confident cents. Before treating a price as a reliable probability, check the participation behind it and the domain it sits in.
How reliable is the evidence that prediction markets are accurate?
The evidence base is strong enough to take seriously and new enough to hold loosely. Several of the recent large-scale findings above are working papers and preprints rather than peer-reviewed results, so their exact magnitudes may shift under scrutiny.
Three complications carry the most weight. The first is that reported trading volume can overstate real activity. A network analysis of Polymarket flagged roughly a quarter of its volume as wash trading, self-dealing that inflates the appearance of liquidity, rising above half of weekly volume in some stretches (Sirolly and coauthors). A separate decomposition of the 2024 election market found that removing share creation and other non-trading activity cut reported volume by more than half (Tsang and Yang). Since accuracy rests on genuine participation, inflated figures make some markets look deeper, and therefore more trustworthy, than they are.
The second is that the accuracy may not come from broad participation at all. A study of the full transaction history of one large venue attributed its forecasting accuracy to roughly 3% of traders who are persistently skilled, funded by the losses of everyone else (Gomez Cram and coauthors). On that reading a price is accurate because informed money has pushed it there, not because the errors of many independent participants averaged out, and that is a narrower and more fragile foundation than the usual account implies.
The third is selection. The flattering calibration studies tend to run on the deepest, most liquid markets, while most listed contracts are thin. One assessment of several thousand markets found that accuracy had plateaued and that the large majority of volume sat in sports and crypto rather than the civic questions markets are most often praised for forecasting (Schwarz).
The 2024 US election, frequently cited as a vindication, shows how much the chosen standard matters. Markets did name the winner before the race was called, and on the stricter tests they were noisier. Accuracy varied sharply across platforms in the final weeks (Clinton and Huang), option-implied estimates were steadier than market prices through the campaign (Saiegh), and a broader review found markets did little better than statistical models on the popular vote and the Electoral College while faring poorly down-ballot (Undark). Calling the winner and forecasting accurately are not the same achievement, as the opening section noted.
A last subtlety is that a price does not mean the same thing everywhere. A calibration study spanning hundreds of millions of trades found that the relationship between price and outcome frequency differs by domain, with political contracts in particular clustering toward 50¢ and resolving more decisively than their prices implied (Le). A 70¢ contract on an economic release and a 70¢ contract on an election may not describe the same 70%.
None of this argues against reading prices as probabilities. It argues for reading them as serious estimates with error bars rather than as certified measurements, weighted by the liquidity behind them and the domain they sit in.
Related guides
Wisdom of Crowds
Why aggregated judgment from many participants can be accurate, and the conditions it depends on.
Superforecasting
How individual forecasters are scored with Brier scores, and how the craft carries into markets.
Markets vs Polls
How market prices and opinion polls compare as forecasts, from mechanism to track record.
Prices and Probabilities
How a contract price reads as a probability, and where the conversion needs care.