Glypse
Sign inSign up
TrendingTop MoversCryptoFinanceElections
OverviewWisdom of CrowdsSuperforecastingCalibrationBrier ScoreBase Rates

Calibration

What it means for probability forecasts to match outcome frequencies, and how the skill is measured and trained.

Every probability forecast makes a checkable promise. A trader who concludes that an event is 70% likely is claiming, in effect, that judgments of this kind come true about seven times in ten, and the claim can be tested once enough of those judgments resolve. Calibration is the property being tested. A forecaster is well calibrated when events happen about as often as their stated probabilities say they should, at every level of confidence. The concept anchors most serious evaluation of forecasts. Market prices are read as probability estimates on exactly this logic (Wolfers and Zitzewitz), forecasting platforms score their communities against it, and the research record on venue prices applies the same test at scale in Market Accuracy. That guide asks whether markets are calibrated. This one stays with the forecaster: what calibration means, how it differs from accuracy, how a calibration curve reads, what research finds about overconfidence, and how the skill is measured and improved.

What is calibration in probability forecasting?

A single probability is a statement about one event, but its meaning comes from a class. Saying 80% places the event among all the things you claim 80% confidence about, and the class is what can be graded. Collect every forecast a person made at a given confidence level, wait for the resolutions, and compare the share that came true with the level that was claimed. When the two match across confidence levels, from rare events to near-certainties, the forecaster is calibrated. Their 60% is a real 60% and their 90% a real 90%, in the only sense a probability can be real for events that happen once.

The word borrows from instrumentation, and the metaphor is exact. A thermometer is calibrated when its readings match the temperatures that produce them, and a forecaster is calibrated when stated probabilities match the frequencies that follow. Miscalibration likewise has direction. A forecaster whose 85% calls come true two thirds of the time is overconfident, claiming more certainty than the judgment delivers. One whose 60% calls come true three quarters of the time is underconfident, and that is a genuine failure too, since probabilities that systematically undersell what a forecaster knows mislead a reader just as surely as inflated ones.

For a trader the property is practical rather than academic. Every position implies a probability judgment, since buying Yes at 62¢ only makes sense if the event looks more likely than 62%, a reading Prices and Probabilities covers in full. A trader whose stated 80% convictions come true just over half the time pays for that gap on every trade built on one of them, whatever else is going right in their analysis.

Calibration is a property of a record

A 70% forecast that misses is not evidence of miscalibration. A well-calibrated forecaster expects exactly that result three times in ten. The property only becomes visible across many resolved forecasts, which is why measuring it takes a written log rather than a memory of having been mostly right.

What is the difference between calibration and accuracy in forecasting?

Scoring theory grades a forecast record on three separable properties, and they answer different questions:

PropertyThe question it answersWhat it cannot show alone
CalibrationDo events happen as often as your probabilities claim?A forecaster can be calibrated while never saying more than the base rate
ResolutionDo your probabilities move decisively toward what actually happens?Decisive probabilities can claim far more certainty than the record supports
AccuracyHow close do forecasts land to outcomes, averaged over a scored record?A single combined number cannot show which of the other two failed

Calibration alone can be earned by refusing to commit. In a category where roughly 30% of events of a given kind occur, a forecaster who answers 30% to every question will grade as perfectly calibrated for as long as the base rate holds, while contributing nothing that a table of historical frequencies would not. What makes forecasts worth reading is resolution, sometimes called discrimination. A forecaster with resolution assigns 85% to the cases that go on to happen and 10% to the ones that do not, instead of quoting the category average every time.

The two properties pull against each other in practice. Moving probabilities away from the base rate creates the room to be badly miscalibrated, and hugging the base rate protects calibration at the price of saying nothing. Proper scoring rules are designed to grade both at once, so that neither retreat is rewarded. The standard one for probability forecasts is the Brier score, which has its own guide in Brier Score, and a standard decomposition of that score splits it into a calibration term, a resolution term, and the irreducible difficulty of the questions asked.

One more boundary keeps expectations straight in a trading context. Calibration measures your probabilities against outcomes, not against the market's price. A calibrated forecaster whose honest 70% meets a contract trading at 70¢ has no trade, because profit needs a justified disagreement with the price rather than a well-tuned agreement. How that distinction plays out in practice is part of Superforecasting.

How do you read a forecaster's calibration curve?

A calibration curve is the standard picture of the property. Group a record of resolved forecasts into buckets by stated confidence, compute the share of each bucket that came true, and plot realized frequency against stated confidence. A perfectly calibrated record traces the diagonal, where the 60% bucket resolves true 60% of the time and the 90% bucket 90% of the time. Departures from the diagonal are the diagnosis, and the same information reads naturally as a table. The record below is synthetic, and its shape is the one research most often finds in untrained judgment:

Stated confidenceResolved forecastsShare that came true
near 55%6053%
near 65%8059%
near 75%9066%
near 85%7072%
near 95%4079%

Read the last column against the first. Near the middle the record is close to honest, and the gap widens with each step of claimed certainty until the near-certainties, which came true a little less than four times in five. Nothing in any single row is damning, since a 95% call is allowed to miss. The pattern across the rows is the finding.

Curve shapes come in recognizable families, catalogued in the expert-judgment literature (Koehler, Brenner, and Griffin). When high-confidence buckets under-deliver and low-confidence buckets over-deliver, as above, probabilities sit too far from the middle, a shape called overextremity and usually summarized as overconfidence. The mirror image, probabilities huddled toward 50% while outcomes run more decisive than claimed, is underextremity, or underconfidence. A record can also run entirely to one side of the diagonal, with events at every confidence level happening less often than forecast. That shape is overprediction bias, and studies reviewed in the same chapter tend to find it where people forecast outcomes they are personally invested in, such as their own projects succeeding.

Two cautions keep curve-reading honest. Buckets need volume before their frequencies mean much, since a 90% bucket holding ten forecasts can land at seven hits or ten without saying anything reliable about the forecaster. And a curve summarizes the past rather than guaranteeing the next forecast, so the usable signal is a gap that persists in the same region of the curve over time, not any single bucket's stumble.

The same construction applies when the forecasts under test are market prices rather than personal judgments, and Calibration City publishes venue-level curves across the major platforms. How those market-level results read, and where they break down, is the territory of Market Accuracy.

What does the research show about overconfidence in forecasting?

The laboratory record is one of the oldest in judgment research. Across the studies collected in the classic survey by Lichtenstein, Fischhoff, and Phillips, people answering general-knowledge questions showed stated confidence running consistently ahead of hit rates, with the gap widest at the highest confidence levels. Later work gave the headline structure. The strength of the effect tracks question difficulty. Hard question sets produce marked overconfidence, very easy ones can flip the record into underconfidence, and the regularity became known as the hard-easy effect (Koehler, Brenner, and Griffin review the evidence, including a survey of 25 task sets that found exactly that split).

Overconfidence is therefore better read as a tendency with structure than as a universal constant, and the strongest evidence for that reading is the experts who escape it. Weather forecasters issuing precipitation probabilities are the canonical case, with a calibration record described in the literature as an existence proof that the skill is reachable (Koehler, Brenner, and Griffin recount assessments ranging from "superb" to "champion"). Their working conditions appear to explain much of it: the same question type every day, an outcome nobody can influence, explicit base rates at hand, and feedback that is fast, unambiguous, and impossible to argue with. Expert bridge players show a similar profile for similar reasons. Miscalibration tends to thrive where feedback is slow, ambiguous, or easy to reinterpret, and to shrink where the environment keeps score.

Markets complicate the picture in a useful way. A price is a crowd forecast rather than one person's, and at venue scale calibration appears to vary by domain rather than running uniformly overconfident. A calibration study spanning hundreds of millions of trades across two major venues found political contracts clustering toward 50¢ and resolving more decisively than their prices implied, a lean toward underconfidence rather than its opposite (Le), and the pattern has drawn sustained practitioner debate. Individual judgment and pooled prices can miss in different directions, which is one more reason to grade your own record rather than assume it inherits the market's virtues.

Where overconfidence gets expensive

The costly region is the extremes. Suppose your record shows that the judgments you call 95% come true about four times in five. Acting on the next one by paying 90¢ for a contract is, on your own evidence, paying 90¢ for exposure worth closer to 80¢, and the loss repeats every time the pattern does. Trimming certainty at the edges of your range tends to be worth more than sharpening estimates near the middle.

How do you measure your own calibration as a forecaster?

The measurement itself requires nothing beyond discipline. Write down the probability at the moment of the forecast, attach a resolution criterion specific enough that a stranger could settle it later, and record the outcome when it arrives. The written number is the load-bearing part, because unrecorded forecasts tend to soften in memory toward whatever ended up happening. The care that platforms apply to question wording belongs in a personal log for the same reason, and Metaculus's question-writing guidance is the standard reference for criteria that resolve cleanly.

Log each forecast Record the outcome Group by confidence level Compare frequency with confidence Review the gaps
The calibration feedback loop: each forecast is logged with an explicit probability, outcomes are recorded as they resolve, and realized frequencies are compared with the confidence levels that claimed them before the next round of forecasts.

Once enough forecasts resolve, the record grades itself. Group entries into confidence buckets, measure each bucket's realized frequency, and weight the gaps by how much of the record sits in each bucket:

Calibration gap  =  ∑knkN∣pˉk−rk∣\text{Calibration gap} \;=\; \sum_{k} \frac{n_k}{N} \left\lvert \bar{p}_k - r_k \right\rvertCalibration gap=k∑​Nnk​​∣pˉ​k​−rk​∣

In plain language, the formula takes, for each bucket, the gap between the average stated probability and the share of forecasts that came true, weights that gap by the fraction of all forecasts sitting in the bucket, and adds the results, with zero as a perfect score. The number captures the calibration half of forecast quality only, and a proper score such as the Brier score folds in resolution as well.

Sample size deserves respect here. With twenty forecasts in a bucket, ordinary chance can move the realized frequency by ten points or more, so early curves are rough sketches, and the direction of a persistent gap matters more than any exact figure. For a trader, part of the record already exists. Every entry implies a probability, and writing your own estimate next to the entry price turns each position into a scored forecast, with the difference between the two numbers as the edge you claimed at the time. Prices and Probabilities explains the implied reading this rests on.

A small set of free tools removes most of the friction:

  • Fatebook is a purpose-built prediction log with resolution dates and calibration charts.
  • Quantified Intuitions offers calibration exercises, and its pastcasting mode scores you on already-resolved questions while restricting your research to sources from before the resolution date. Its pastcasting FAQ also explains why calibration and overall forecasting skill are different measurements.
  • Good Judgment Open runs Brier-scored public forecasting challenges, so a season of participation produces a scored record with no setup at all.
  • Metaculus's scores FAQ documents how a large forecasting platform applies these same ideas to its community, useful as a model for grading a personal log.

How do you improve your calibration as a forecaster?

The research reads as good news, because calibration appears to respond to deliberate practice. Four levers carry direct evidence.

Close the feedback loop, quickly and often. The expert-judgment literature includes a clean demonstration in which professional forecasters received a year of detailed probabilistic feedback on their own predictions, and the following year's calibration curve moved close to the diagonal (Koehler, Brenner, and Griffin describe the study). The weather-forecaster conditions generalize into practice. The loop in the figure above trains nothing while it stays open, which favors many small forecasts on questions that resolve in days or weeks over a handful of grand ones that resolve in years.

Take structured training seriously, because the measured effect is real. In the Good Judgment Project's forecasting tournaments, a randomized experiment found that a probabilistic-reasoning module taking under an hour, covering base rates, belief updating, and common biases, improved accuracy by 6 to 11% relative to controls, with the effect repeating in each of four tournament years. The base-rate component does particular work at the confident end of the range, since anchoring on how often events of a kind actually happen restrains the probabilities a vivid story would otherwise push toward the extremes. Base Rates develops that method on its own.

Commit to numbers, not words. A forecast of "probably" cannot be scored, so it cannot teach anything. Tournament data showed the most accurate forecasters working in unusually fine probability increments, and their precision carried real information rather than false exactness (Mellers and colleagues), a habit documented in depth across the Good Judgment Project evidence. Granular numbers create the gradable record that vague language quietly prevents, and the surrounding habit set is the subject of Superforecasting.

Argue against your own forecast before locking it in. In a classic experiment reviewed in the same expert-judgment chapter, participants asked to list reasons their chosen answer might be wrong became measurably less overconfident, while listing supporting reasons changed nothing, though later replications of the effect have been mixed. Tetlock's guidelines for aspiring superforecasters fold the same move into standing practice as a balance between prudence and decisiveness.

Calibration makes your probabilities mean what they say, and a market then asks a harder question. Profit at resolution comes from probabilities that differ from the price in the right direction, so a well-kept record is the foundation of an edge rather than the edge itself. The craft of carrying that judgment into a market, and the record markets themselves compile against the same test, sit in the guides below.

Related guides

Market Accuracy

How the calibration test is run on market prices at scale, and where the record holds up or thins out.

Brier Score

The scoring rule that grades calibration and decisiveness together in a single number.

Superforecasting

What the Good Judgment Project found about accurate forecasters, and how the craft transfers to trading.

Base Rates

Anchoring a forecast on how often comparable events happen before weighing the case at hand.

Superforecasting

What the Good Judgment Project found about accurate forecasters, and how far the craft carries into markets.

Brier Score

What the standard accuracy score for probability forecasts measures, and what counts as a good one.

On this page

What is calibration in probability forecasting?What is the difference between calibration and accuracy in forecasting?How do you read a forecaster's calibration curve?What does the research show about overconfidence in forecasting?How do you measure your own calibration as a forecaster?How do you improve your calibration as a forecaster?Related guides
Glypse

The AI research engine for prediction markets

Guides

  • Prediction Markets
  • Forecasting Craft
  • Signals & Analytics

Company

  • Terms of Use
  • Privacy Policy

Content on this site is provided for informational purposes only. It is not investment, financial, or trading advice, and it is not a recommendation to buy or sell any prediction market contract or other instrument. Analytics are generated by automated systems, including AI models, and may contain errors or omissions. Trading prediction market contracts involves risk, and you can lose some or all of the amount you commit. We recommend that you do not trade based on this information alone; do your own research and verify anything you read on this site before acting on it. Glypse is not a prediction market, exchange, broker, or trading advisor, it does not execute trades or hold funds, and it is not affiliated with, endorsed by, or sponsored by Polymarket, Kalshi, or any other prediction market. All trademarks belong to their respective owners. You alone are responsible for your decisions, based on your own objectives, financial circumstances, and risk tolerance, and for complying with the laws of your jurisdiction. Consult a qualified professional regarding your specific situation. See the Terms of Use for more information.

Copyright © 2026 Glypse, Inc. All rights reserved.

Trending
Top Movers