Superforecasting
What the Good Judgment Project found about accurate forecasters, and how far the craft carries into markets.
Every prediction market price is assembled from individual judgments, which raises a question the price alone cannot answer: can individual judgment be measurably, repeatably good? A series of government-sponsored forecasting tournaments says it can. A small fraction of forecasters predict consistently better than everyone else, their advantage persists year after year, and the habits behind it appear learnable rather than innate. Those people came to be called superforecasters, and the practice distilled from studying them is superforecasting. The research is worth knowing in some detail, because a prediction market is, in effect, a forecasting tournament with prices attached, and the craft that wins tournaments is the closest thing there is to a studied version of the judgment a trader brings to a contract.
What did the Good Judgment Project find?
In 2011, IARPA, the research agency of the US intelligence community, launched a multi-year forecasting tournament called the Aggregative Contingent Estimation program. Competing research teams answered the same stream of geopolitical questions with probability estimates, and those estimates were scored against real outcomes once the questions resolved. The questions were concrete and time-bound: whether a leader would remain in office past a given date, whether a treaty would be signed, whether an indicator would cross a threshold. The Good Judgment Project, led by psychologists Philip Tetlock and Barbara Mellers at the University of Pennsylvania, entered with thousands of volunteers, collected over a million forecasts, and won the tournament decisively.
Accuracy was measured with the Brier score, the standard yardstick for probability forecasts, and forecasting platforms build their scoring systems on the same family of rules.
The formula takes each forecast probability, subtracts the outcome (recorded as 1 if the event happened and 0 if it did not), squares the difference, and averages the result over all forecasts. Lower is better. A perfect forecaster scores 0, and answering 50% on every binary question scores 0.25. A 70% forecast on an event that happens contributes 0.09 to the average, while the same forecast on an event that does not happen contributes 0.49. That asymmetry is the point of the design. Confident misses are expensive, so the score rewards people who know how sure to be.
The project's findings gave the field its vocabulary. The top 2% of forecasters each season, the group the researchers came to call superforecasters, were not simply having a lucky year. Roughly 70% of them kept their status from one season to the next, and individual performance correlated at about 0.65 across years, a pattern that reads as skill rather than chance. The advantage was also large. By the program's own account, relayed in the same evidence review, its aggregate forecasts beat a prediction market inside the intelligence community, whose participants were professional analysts with access to classified information, by roughly 25 to 30%. And the skill proved trainable. A randomized experiment within the tournament found that a cognitive-debiasing module lasting under an hour, covering base rates, belief updating, and common biases, improved accuracy by 6 to 11% relative to a control group, with the effect repeating in each of four tournament years. The full story reached general readers through Superforecasting: The Art and Science of Prediction, the 2015 book by Tetlock and Dan Gardner.
Superforecaster is a measured title
Within the research program, superforecaster referred to the top 2% of participants by Brier score, a status most of them retained in later seasons. Outside that context the word gets used loosely, so when someone claims the label, the useful question is whether a scored, public track record sits behind it.
What do superforecasters do differently?
The research is unusually specific about the habits that separate superforecasters from everyone else, because the tournament recorded behavior alongside accuracy: how often people updated, how fine-grained their probabilities were, how they engaged with opposing views. The profile that emerged is a set of practices rather than a personality type.
| Habit | What it looks like in practice |
|---|---|
| Start with the outside view | Establish how often events of this kind happen before weighing the details of this case |
| Decompose the question | Break one large unknown into smaller conditions that can each be estimated |
| Use granular probabilities | Distinguish 60% from 65% instead of rounding to "likely" |
| Update in small steps | Revise by a few points as evidence arrives, saving large jumps for genuinely large news |
| Keep score | Record every forecast, compare it with the outcome, and study the misses |
| Stay actively open-minded | Treat each belief as a hypothesis and seek the strongest case against it |
The outside view deserves the most attention because it is the least intuitive. A base rate is the historical frequency of an event class: how often incumbents lose comparable races, how often projects of a given size finish on time, how often a company under investigation ends up penalized. Superforecasters tend to anchor on that frequency first and adjust for the specifics of the case second, a sequence that protects the estimate from the pull of a vivid story. The set of comparable past cases is called a reference class, and choosing it well is a large part of the skill.
Granularity is the most visible habit. Where most people think in coarse steps such as "unlikely," "toss-up," and "probably," the tournament data showed superforecasters working in fine increments, and their precision carried real information rather than false confidence. Updating followed the same pattern of many small revisions, with occasional large moves when the evidence justified them. Tetlock later condensed the practice into ten guidelines for aspiring superforecasters, which include triaging effort toward questions where work actually improves the answer and balancing underreaction against overreaction to news.
The habits form a loop rather than a checklist, because scoring feeds the next forecast.
None of this requires a tournament invitation. Good Judgment Open runs Brier-scored public forecasting challenges, and Quantified Intuitions offers calibration exercises built around the same loop. Working through either for a season supplies the feedback the habits depend on, since calibration only improves when forecasts are recorded and checked.
Do teams and aggregation improve forecast accuracy?
The tournament treated collaboration as an experiment rather than an assumption. Forecasters were assigned to work alone or in teams, and teaming reliably improved accuracy, with teams of superforecasters performing best of all. The mechanism appears straightforward. Teams surface more information, divide research across members, and challenge one another's reasoning. The benefit depends on norms that keep disagreement alive, since a deferential team converges on its loudest member instead of the evidence.
Aggregation was the second lever, and the stronger one. The project's winning submissions were not any single person's judgment but a pooled forecast that weighted individuals by track record and recency, then pushed the combined probability away from 50% to correct for the caution that averaging introduces, a step known as extremizing. By the program's account, this weighted, extremized aggregate beat unweighted crowd averages by reported margins above 60%. Why pooled judgment tends to beat the individuals inside it, and the conditions under which the effect breaks down, is a subject of its own, covered in Wisdom of Crowds.
The aggregation effect extends beyond people. A study in Science Advances found an ensemble of twelve language models statistically indistinguishable from a crowd of 925 human forecasters over a three-month tournament, even though no individual model stood out. And when researchers built ForecastBench, a benchmark that scores machine forecasters on genuinely unresolved future questions, they chose superforecasters as the human reference class. At the benchmark's launch, the expert humans still led the best model. The designation earned in one tournament became the measuring stick for another field.
Prediction markets belong to this same family of aggregation mechanisms, with one structural difference: instead of averaging stated probabilities, a market aggregates positions, weighting each participant's view by the capital they are willing to put behind it.
What are the criticisms and limits of superforecasting?
The findings hold up well within their frame, and the serious criticisms are mostly about where the frame ends.
The horizon is short. Most tournament questions resolved within about a year, and the demonstrated edge lives at that range. Tetlock's own caution, quoted in the AI Impacts evidence review, is that there is no evidence geopolitical or economic forecasters can predict anything ten years out "beyond the excruciatingly obvious." Claims about superforecasting are best read as claims about questions measured in weeks and months, not as a case for seeing the far future.
Only scoreable questions get asked. A tournament can score a question only if it resolves crisply, which is why forecasting platforms apply standards like Metaculus's question-writing guidelines, under which a well-formed question could be settled by a clairvoyant without further clarification. That discipline filters out vague but important questions, and even carefully drafted ones can diverge from what the asker really wanted to know, a family of failure modes cataloged in Rethink Priorities' taxonomy of specification problems. Demonstrated forecasting skill is therefore skill on operationalized questions, which is narrower than foresight in general.
Parts of the method are contested. How much the extremizing step contributed to the winning aggregate has been debated since the tournament, and the superforecaster group is small enough that some sub-analyses rest on limited samples. Persistence and trainability are the sturdiest parts of the record; precise magnitudes deserve a looser grip.
Accuracy is not automatically useful. Dan Schwarz's data-driven essay in Asterisk argues that calibrated probabilities are increasingly easy to produce, while decision-makers who change course because of them remain scarce, which makes demand the binding constraint on forecasting's value. A forecast nobody acts on is a scored opinion. Traders are an unusually direct audience in this respect, since a market participant acts on a probability every time they trade.
Does superforecasting transfer to prediction markets?
The craft transfers better than the track records do. The forecasting community and the trading community have developed largely in parallel. Calibration, scoring rules, and question operationalization are taught on reputation platforms and in tournaments, while market venues teach mechanics, and relatively few people move between the two worlds. The tournament evidence does not translate one-for-one, because the settings differ structurally.
| Dimension | Forecasting tournament | Prediction market |
|---|---|---|
| What you submit | A probability estimate | A position at a price |
| What gets scored | Distance from outcomes (Brier score) | Profit and loss after costs |
| The competition | Other forecasters' scores | The current market price |
| Cost of participating | Time and attention | Spread, fees, and capital at risk |
| Reward for accuracy | Score and standing | Proceeds at resolution |
The deepest difference is the opponent. A tournament forecaster beats the field by sitting closer to outcomes than other entrants on average. A trader faces a single opponent, the price, and it is a formidable one because it already reflects the pooled judgment of everyone trading the contract. That pooled judgment has a strong record. The field-defining survey by Wolfers and Zitzewitz treats binary contract prices as workable probability estimates that outperform moderately sophisticated forecasting benchmarks, and across five presidential election cycles, the Iowa Electronic Markets sat closer to the final result than 964 contemporaneous national polls about 74% of the time. Profit requires a judgment that differs from that price in the right direction by more than the cost of trading, and it requires this repeatedly rather than once. Suppose your considered probability on a contract is 70% and it trades at 66¢. The gap is 4¢ of expected value per contract before costs, and the spread plus fees can claim much of it unless the judgment is genuinely better than the market's.
Calibration is not the same property as edge
A well-calibrated forecaster, whose 70% calls come true about 70% of the time, can still lose money trading, because the market may already price those events near 70¢. Returns depend on the gap between your probability and the price, measured against spread and fees. Calibration keeps your probabilities honest; edge is having them differ from the market's in the right direction.
The tournament staged one direct comparison, and it favors the forecasters with caveats attached. Tetlock and Gardner report that teams of superforecasters outperformed prediction markets run within the same tournament. Those markets were small and thinly traded compared with the large public venues that came later, so the result shows that elite aggregated judgment can compete with market prices, not that it would beat a deep, liquid market. Market data suggests the tournament's central pattern does reappear inside trading venues. A working paper analyzing the universe of Polymarket transactions attributes market accuracy to roughly 3% of traders who are persistently skilled, with their profits funded by the losses of everyone else, and an independent analysis of 72 million Kalshi trades found that participants who cross the spread underperform by roughly 1% per trade on average while liquidity providers collect the mirror image. A small, persistent minority appears to carry the signal in both settings; in markets, the rest of the volume pays for it.
What carries over is the craft, applied to a different scoreboard:
- Base rates before narratives. Asking how often comparable events have happened tends to be a better anchor than the day's coverage, in a market as in a tournament.
- Granular probabilities. The difference between 60% and 65% is the difference between a trade and no trade when the contract sits at 62¢.
- Small, frequent updates. Prices move continuously, and the habit of revising in increments maps naturally onto re-evaluating a position as evidence arrives instead of reacting once to a headline.
- Keeping score. Every trade implies a probability judgment, so a record of entry prices and outcomes is a Brier-style track record in disguise. Prices and Probabilities explains the implied-probability reading this depends on.
- Reading the question. Tournament forecasters learn that the question text is the contract. The same discipline applies to event contracts, which settle on written resolution criteria rather than on what a headline seems to say.
Much of the remaining work is information: finding what bears on the question, weighing it against the base rate, and noticing when the conversation around an event shifts before the price does. The fair summary of the evidence is that superforecasting practices tend to raise the quality of the judgment a trader brings to a price, while the price remains a strong opponent that already contains most of what is publicly known.
Related guides
Wisdom of Crowds
Why pooled judgment tends to beat individuals, and the conditions aggregation needs to work.
Prices and Probabilities
How contract prices translate into implied probabilities, and what to check before trusting one.
Core Concepts
How trading on event outcomes turns individual judgments into a public forecast.