Calibration
- Categories
- Decision Making
The match between stated confidence and actual hit rate: a well-calibrated forecaster is right about 70% of the time on the things they call 70% likely. Calibration can be measured over many predictions (the Brier score is one such measure) and improved with feedback, separating the feeling of confidence from real accuracy.
Why it Matters
Confidence and accuracy come apart, especially where feedback is weak, so subjective certainty is a poor guide to truth. Measuring calibration replaces that feeling with a track record, the only way to know whether judgment is actually any good and the only way to get better at it.
Signals
- Confident predictions with no record kept, so accuracy is never checked.
- A long history of "sure things" that miss far more often than claimed.
- Confidence that rises with the vividness of a story rather than the weight of evidence.
Benefits
An objective read on judgment quality, a basis for trusting some forecasters over others, and a feedback loop that turns practice into improvement rather than repetition.
Risks
Optimizing the score instead of the judgment (hedging toward 50% to avoid being wrong); assuming calibration in one domain transfers to another where feedback is absent.
Tensions
Good calibration needs frequent, clear feedback, yet the highest-stakes judgments are often rare and slow to resolve, exactly where calibration is hardest to build.
Examples
A weather service whose "70% chance of rain" rains about 70% of the time; a forecaster who keeps score and discovers their confident calls are no better than chance.