Base Rates
Why probability estimates start from the frequency of comparable past cases, and when that anchor fails.
Every contract price is a probability claim about a single event, and single events do not arrive with probabilities attached. The discipline that keeps such claims honest begins with a question about the past rather than the case at hand. How often do events of this kind happen? The answer is called a base rate, the long-run frequency of an outcome across a class of comparable situations, and forecasting research has repeatedly found that people who anchor on that frequency before engaging with a case's details tend to predict better than people who reason from the details alone. Superforecasting presents the tournament evidence behind that finding. This guide is about the habit itself: what base rates and reference classes are, how to choose a reference class when every event belongs to many, how to update away from the anchor as evidence arrives, and where the method stops working. For a trader the material is practical, because a contract price read as a probability can be checked against the frequency of its event class before any capital moves.
What is a base rate in forecasting?
A base rate is the frequency of an outcome across a class of similar past situations, counted with the case at hand set aside. The set of past cases doing the counting is called the reference class, and the two ideas travel together, since a base rate is only defined relative to the class it was counted over. The move looks the same in any domain:
| Question being forecast | A reference class | The base-rate question |
|---|---|---|
| Will this incumbent win another term? | Incumbents in comparable races | What share of them won? |
| Will this project ship on schedule? | Projects of similar size and type | What share finished on time? |
| Will this contract resolve Yes by its deadline? | Similar contracts with similar time remaining | What share resolved Yes? |
What makes the base rate worth naming is how reliably people skip it. In a long series of experiments, Daniel Kahneman and Amos Tversky showed that when people hold both a class frequency and a vivid description of the specific case, the description tends to crowd out the frequency even when it carries little real information, a pattern known as the base rate fallacy. They attributed it to a mental shortcut called representativeness, under which a case is judged by how well it resembles a stereotype rather than by how often cases of its kind occur. A candidate who looks like a winner reads as a winner, whatever the record of similar candidates shows.
One habit, three names
The same move goes by different names. Psychology calls the neglected quantity the base rate, Bayesian statistics calls it the prior, and the judgment literature calls consulting it taking the outside view. All three terms point at the frequency of the reference class.
A trader meets base rates on both sides of a position. Their own estimate needs an anchor before case evidence means anything, and the market's price can itself be measured against the class frequency. When a contract trades well away from the base rate of its event class, the distance amounts to a claim that this case differs from its relatives, and the rest of this guide is about deciding when that claim deserves belief. Whether prices in aggregate match the frequencies they imply is a different, measurable question, covered in Market Accuracy.
What are the inside view and the outside view in forecasting?
The vocabulary for the two ways of building a forecast comes from Kahneman's work with the management scholar Dan Lovallo. The inside view assembles the forecast from the particulars of the case: the plan, the people, the obstacles in sight, the scenario that seems to follow from them. The outside view sets the particulars aside and asks how things turned out across the reference class. Their argument to executives is that the inside view feels rigorous while quietly overweighting the story in front of the forecaster, and that the corrective is to locate the case in the distribution of outcomes for comparable cases.
The founding example is disarmingly small. In the 1970s Kahneman assembled a team to write a high-school curriculum on judgment and decision making. After a year of steady progress he asked each member to estimate the time to a finished draft, and the estimates clustered around two years. He then asked the one member who had watched many curriculum teams how comparable efforts had fared. The answer, produced reluctantly by a man whose own guess had been two years, was that roughly forty percent of such teams never finished, and that the ones that did took seven to ten years. Nobody in the room had thought to reason from that record, including the person who held it. By Kahneman's account, the draft took about eight more years to complete.
The pattern received a name in 1979, when Kahneman and Tversky described the planning fallacy, the tendency of forecasts built from the inside to underestimate time and cost even when the forecaster knows that similar efforts have run long. In one frequently cited study, students predicted an average of just under 34 days to finish their theses and took just over 55, with about 30% finishing by their predicted date. The corrective became an engineering practice. Under the name reference class forecasting, Bent Flyvbjerg turned the outside view into a formal procedure that forecasts a new infrastructure project's cost and schedule from the realized outcomes of completed projects of the same type. The method he framed as getting risks right was adopted into official UK transport appraisal guidance in 2004.
The sequence is the point. Run inside-first and the story sets the anchor, leaving the base rate a technicality to argue away. Run outside-first and the burden reverses, with every step away from the class frequency needing evidence behind it. Tournament research backs the ordering. Heavy use of comparison classes was among the habits that distinguished the most accurate forecasters in the Good Judgment Project, and a randomized training experiment in the same tournament found that a short module leading with base rates improved accuracy by 6 to 11% in each of four seasons. Tetlock's published guidance frames the skill as balancing the two views, since the outside view supplies the anchor and the inside view supplies everything the class cannot see.
How do you choose a reference class for a forecast?
Choosing the reference class is where most of the judgment lives. Every event belongs to many classes at once. A sitting governor seeking another term is an incumbent, a member of a party, a regional politician, and a candidate in a particular economy, and each framing counts a different set of relatives. Philosophers of probability named the difficulty first. John Venn observed in 1876 that any single thing belongs to "an indefinite number of different classes," and the puzzle of which class should supply the probability for a given case is known as the reference class problem. Forecasting practice manages the problem rather than solving it, through a trade-off between similarity and sample size.
Suppose the question is whether a sitting governor wins re-election.
| Candidate reference class | Typical sample | What it buys | What it costs |
|---|---|---|---|
| All incumbent governors seeking re-election | Hundreds of races | A stable, well-estimated frequency | Averages over eras and contexts that may matter |
| Same-party incumbents in comparable states | Dozens of races | Political context closer to the case | A noisier frequency |
| Recent races in this state with an incumbent running | A handful | Maximum similarity | Too few cases to estimate anything |
Each class returns a different number. Suppose the broad class puts the frequency near 75%, the middle class near 60%, and the narrow class offers three wins in five tries. None of these is the true base rate. The class is a modeling choice, and the frequency it returns inherits that choice.
Practice has produced workable habits. The broadest class that still resembles the case in the ways that plausibly drive the outcome tends to serve best, since breadth is what makes a frequency trustworthy. Computing several defensible classes turns the spread itself into information. Agreement among them makes the anchor sturdy, and divergence means the class choice is doing the real work, so the estimate deserves wider error bars. Defining the outcome precisely matters as much as choosing the class, since a class of projects that "succeeded" cannot be counted until success has a definition, the same discipline forecasting platforms apply to question wording. Where someone else has already done the counting, that work is worth using, and forecasting teams maintain public collections of base rates precisely because a shared, precomputed class tends to beat one improvised in the moment.
How do you update from a base rate as new evidence arrives?
An anchor is where an estimate starts, and cases genuinely differ from their classes. The question is how far the evidence at hand justifies moving, and the arithmetic that answers it is Bayes' rule, most usable in its odds form:
Read aloud, the formula says the odds after seeing a piece of evidence equal the odds before, multiplied by how much more often that evidence appears when the event goes on to happen than when it does not. The multiplier is called the likelihood ratio, and it is the only part the news of the day controls. The base rate sets the prior odds; the evidence can only scale them.
Suppose the chosen class frequency for an incumbent's race sits near 60%, which is prior odds of 3 to 2. Polling then moves toward the incumbent, and in this example the record shows movement of that size appearing about twice as often in races the incumbent went on to win as in races they lost. The likelihood ratio is 2, the odds update from 3 to 2 up to 3 to 1, and the estimate becomes 75%. Evidence pointing the other way divides instead of multiplying. What the arithmetic never does is discard the prior in favor of the story.
The likelihood ratio is also what separates diagnostic evidence from loud evidence. A dramatic debate moment that occurs about as often in campaigns the incumbent wins as in campaigns they lose has a ratio near 1 and, however heavily it is covered, justifies almost no movement. Coverage volume and diagnostic strength can be unrelated, and they feel identical in the moment, which is much of why base rates get abandoned mid-forecast. The craft habit that guards both flanks is revising in many small steps as evidence accumulates, with Tetlock's guidance naming the underlying skill as balancing underreaction against overreaction to news.
Do not count the same evidence twice
Whatever narrowed the reference class is already inside the base rate. If the class was incumbents running during strong economies, the strong economy cannot also justify an upward adjustment. The market version of the mistake is reusing evidence that has already moved a price as a reason it should move further. An input can define the class or drive the update, never both.
Prices give this loop a public counterpart. A contract trading well above its class base rate is asserting, in effect, that case evidence multiplies the prior odds by a large factor, and the working question for a trader is whether evidence of that strength exists. Sometimes it does and the price is ahead of the class; sometimes the distance rests on a compelling story with a likelihood ratio near 1. Whether adjustments made this way were earned shows up later, in how often the resulting estimates match outcomes, the territory of Calibration.
When do base rates fail in forecasting?
The method's failure modes are as recognizable as its successes, and most trace back to a reference class that no longer deserves the name.
No real class exists. For an event with no meaningful relatives, the outside view has nothing to count. A base rate produced anyway is an analogy dressed up as a frequency, and a forced class can be worse than none, since it adds confidence without adding information. The honest treatment of a truly novel question is a wide prior, more weight on case reasoning, and less certainty.
The world stopped resembling the sample. A base rate assumes the past cases and the present one were generated by roughly the same machinery. When the machinery changes, through new rules, new technology, or a changed information environment, frequencies computed across the old regime quietly lose their claim on the new one. Incumbency illustrates the risk, since a class spanning many decades mixes political eras that may share little, so the stability of the count can conceal the instability of what was counted.
The sample is too small to be a frequency. A base rate computed from five cases is an anecdote with a denominator. Narrow classes produce exactly these, and a precise-looking percentage can lend a handful of observations unearned authority. When the class is small, the honest anchor is a range rather than a point.
The class was chosen to flatter the conclusion. Venn's indefinitely many classes give a motivated forecaster room to locate one whose frequency supports the position they already hold. Choosing the class before computing the rate, and the rate before forming the estimate, is what keeps the method honest. Hunting for a rate after the estimate is rationalization with better notation.
The far ends of the probability scale deserve separate caution. Rare events produce base rates with only a few occurrences in the numerator, so the difference between a 1% class and a 3% class can be statistically invisible while tripling a contract's fair value. Prices show related trouble in the same region. In the historical racetrack pools where it was first documented, rarely occurring outcomes persistently traded above their long-run frequency, an effect known as the favorite-longshot bias, and low single-digit prices remain the range where class frequencies, and the prices meant to reflect them, are hardest to pin down.
Used with those limits in view, the base rate remains the cheapest defense a forecaster has against narrative. Anchoring costs one question about the past and removes the story's power to set the starting point. A forecaster who never departs from the class is uninformative, and one who departs on every story is unanchored. The practice lives in the earned movement between the two. For a trader it compresses to a single check before any position. What frequency does this price imply for events of this kind, and does the evidence at hand honestly cover the distance?
Related guides
Superforecasting
What the Good Judgment Project found about accurate forecasters, and how far the craft carries into markets.
Calibration
How to measure whether stated probabilities match the frequencies that actually play out.
Prices and Probabilities
How contract prices translate into implied probabilities, and what to check before trusting one.
Brier Score
How probability forecasts are scored against outcomes.