What Forecasting Tournaments Say About Trusting Your Gut

TL;DR

  • Philip Tetlock's Good Judgment Project — originally sponsored by IARPA, the U.S. intelligence community's research arm — ran a multi-year forecasting tournament that measured, rather than theorized, what separates good judgment from bad.
  • A small subset of ordinary volunteers, dubbed "superforecasters," consistently and substantially outperformed both chance and trained intelligence professionals on real geopolitical and economic questions.
  • Superforecasters shared identifiable habits: breaking big questions into sub-questions with knowable base rates, updating beliefs in small frequent increments instead of dramatic reversals, and — the standout finding — tracking their own calibration over time.
  • Almost no growth or product leader logs their confidence in a hypothesis _before_ seeing results and checks back later against what actually happened. Tetlock's data says that single habit is the most powerful lever available for improving judgment.
  • The fix is lightweight: log a confidence percentage on every major test hypothesis before you look at results, then review the log quarterly against outcomes — a personal Brier score for your own business calls.

Every growth and product leader has a mental highlight reel of calls they got right, told with total conviction: "I said this would work, and it did." Almost none of them can tell you, with any precision, how many times they said something would "definitely" work and it didn't — because almost nobody keeps that ledger. The confident calls that missed quietly disappear from the story. The ones that hit become the professional legend. That asymmetry isn't dishonesty. It's the absence of a measurement system, and it is exactly the problem Philip Tetlock spent twenty years and a large-scale forecasting tournament designing a fix for.

The Good Judgment Project: measuring judgment instead of theorizing about it

In the early 2010s, IARPA — the intelligence community's advanced-research arm — ran a multi-year forecasting tournament pitting several research teams against each other on real-world geopolitical and economic questions: Will a specific country default within six months? Will a particular ceasefire hold? Will a named leader remain in power through year-end? Tetlock's team, the Good Judgment Project, recruited thousands of ordinary volunteers with no special access to classified information and had them submit probability forecasts on hundreds of these questions over multiple years.

The result that made the project famous: a small subset of these ordinary volunteers — "superforecasters" — didn't just do well. They substantially and consistently outperformed both random chance and, in the published results, trained intelligence analysts with access to classified information on the same questions. That finding is the whole reason the project matters here. It means good judgment under uncertainty isn't primarily about domain expertise or access to better information — it's a measurable, learnable discipline, separable from subject-matter knowledge, and the tournament identified exactly what that discipline consists of.

The habits that separated superforecasters from everyone else

Three findings from the tournament translate almost without modification into a business context.

Breaking big questions into sub-questions with knowable base rates. Superforecasters rarely answered the headline question directly. Asked whether a leader would remain in power, they'd decompose it: how often do leaders in structurally similar situations get removed within a year? What's specifically different about this case relative to that base rate, and how much should that shift the estimate? The vague, compelling version of the question ("is this leader in trouble?") invites a gut read; the decomposed version invites an actual calculation.

Updating in small, frequent increments — not big dramatic reversals. The best forecasters revised their probability estimates often, and usually by a few points at a time, as new information arrived. They didn't sit on a confident call and then flip it entirely when disconfirming evidence became undeniable. This is the practical, business-legible version of Bayesian updating: your confidence in a hypothesis should move a little with each new piece of evidence, continuously, rather than staying frozen until the evidence forces a dramatic reversal you can no longer avoid.

Tracking their own calibration, explicitly, over time. This is the finding with the most direct business application, and it's the one almost nobody outside the forecasting-research world has adopted. Superforecasters didn't just make predictions — they scored themselves against outcomes using a Brier score, a standard measure of forecast accuracy that penalizes both overconfidence and underconfidence. Someone who says "90% chance" should be right roughly nine times out of ten on similar calls — not seven, not ten. A Brier score, tracked over enough predictions, tells you whether your gut runs hot, cold, or well-calibrated, and it's the only way to actually know rather than assume.

Why this is a business problem, not just a forecasting curiosity

Here's the pattern that shows up constantly in growth and product organizations, and it's worth naming directly: leaders make confident calls on hypotheses constantly — "this redesign will convert better," "this pricing change will hold retention," "this feature is what's blocking activation" — and almost never go back and check the calibration of those calls as a _body of work_. Individual wins and losses get remembered anecdotally. The aggregate pattern — was I actually right 80% of the time when I said I was 80% sure, or was I closer to 50% — never gets computed, because the log that would let you compute it was never kept.

That absence matters more than it looks like it should, for a specific reason: confidence and accuracy are two different things, and without a calibration log, you only ever get feedback on the first one. A leader who is consistently, wrongly, 90% confident sounds exactly as convincing in the room as a leader who is genuinely well-calibrated at 90% — right up until someone tracks the record. Most organizations never track the record, which means conviction gets rewarded as a proxy for judgment, with no actual check on whether the two are the same thing for this particular person on this particular type of call.

The lightweight practice: a personal Brier score for business calls

The tournament's finding doesn't require adopting a research program — it requires one recurring habit, cheap enough to actually stick.

StepWhat it looks likeWhy it matters
1. Log confidence before resultsBefore a test or major bet launches, write down a specific percentage — "I'm 70% confident this beats control by a meaningful margin" — not a vague "I think this will work"Vague confidence can't be scored later. A number can be checked against what actually happened.
2. Resist the urge to round to 50/50 or to extremesUse real numbers across the range — 60%, 70%, 85% — rather than defaulting to "I'm sure" (95%+) or hedging everything to "who knows" (50%)Superforecasters used fine-grained probabilities (as specific as single percentage points); most people compress everything toward the extremes or the middle, which destroys the signal in the score
3. Don't touch the log once it's writtenThe forecast is locked before results come in — no retroactive editing once you know the outcomePost-hoc "I sort of knew that" edits are exactly the bias (hindsight bias) that makes untracked judgment feel better calibrated than it is
4. Review quarterly, in aggregate, not case by casePull every logged prediction from the quarter and check: of the calls you rated 80% confident, did roughly 80% of them actually hit?A single hit or miss tells you nothing about calibration. The pattern across dozens of calls tells you whether your gut runs overconfident, underconfident, or well-tuned
5. Adjust your stated confidence, not just your decisionsIf the review shows your "90% confident" calls only hit 60% of the time, the fix is to distrust your own top-end confidence going forward, specifically — not to become vaguely more cautious everywhereCalibration is often uneven — someone might be well-calibrated at 60% confidence and badly overconfident at 95%. Only a logged record reveals where specifically the miscalibration lives

The mechanism connects directly to how a bet should be sized in the first place: a hypothesis you've logged at 55% confidence and one you've logged at 90% confidence should not be getting the same size of resourcing commitment, and a running calibration score is what tells you whether your stated 90% has actually earned that trust historically. That's the same discipline behind the Confidence Tier Model — sizing the bet to the evidence — except here the "evidence" includes your own historical track record of confident calls, measured rather than assumed.

The uncomfortable part, stated plainly

The likely result of actually running this for a year, for most leaders, is discovering that their gut runs more overconfident than they'd assumed — particularly on the calls that felt most obvious going in. That's not a flattering finding, which is exactly why almost nobody generates the data that would reveal it. Tetlock's forecasters weren't more naturally gifted than intelligence professionals with more information and more experience; they were more willing to keep score on themselves and adjust based on what the score actually showed, rather than what their conviction in the moment suggested. That's a discipline, not a talent, and it's available to anyone willing to write the number down before they know the answer.

FAQ

How is a Brier score actually calculated?

It's the squared difference between your stated probability and the outcome (scored as 1 for "happened," 0 for "didn't"), averaged across all your predictions — lower is better, and a score of 0 means perfect calibration. You don't need to compute it precisely to get the benefit; even an informal quarterly review of "how often did my 70%-confidence calls actually hit" captures most of the value without the formal math.

Isn't this just journaling with extra steps?

The difference is the number and the review discipline. Journaling captures reasoning; a calibration log captures a falsifiable claim (a specific percentage) that gets checked against reality on a schedule. Most business journaling never gets scored against outcomes, which is exactly the step that produces the learning.

What if I don't have enough major decisions to build a meaningful sample?

Log smaller calls too, not just the headline bets — feature hypotheses, minor test predictions, even internal forecasts about team output. Superforecasters built their calibration on volume across many modest questions, not a handful of huge ones. The skill transfers; a well-calibrated gut on small calls tends to be a well-calibrated gut on big ones.

Does this replace statistical rigor in testing?

No — it operates one level up. Statistical testing tells you whether a specific result is likely real. Calibration tracking tells you whether your own pre-test intuition about which hypotheses will pan out is trustworthy, which affects how much weight you should give your gut when deciding what to test in the first place, and how you read ambiguous or underpowered results in the meantime.

How long before the calibration log actually becomes useful?

Meaningful patterns usually need a few dozen logged predictions before the aggregate signal is stronger than noise — roughly a couple of quarters of consistent logging for someone making regular test and product calls. The point isn't a fast verdict; it's building the first real dataset you've ever had on your own judgment.

Related reading: Why Most "Wins" Don't Replicate: The Winner's Curse, The Meta-Analysis Your Experimentation Program Is Missing.

Bottom line

Tetlock's tournament didn't find that some people are naturally gifted forecasters and everyone else isn't — it found that good judgment is a trackable, improvable discipline, and the single habit that improves it fastest is one almost no business leader practices: writing down a real confidence number before you know the answer, and checking it later. It costs a few minutes per major hypothesis and a quarterly review. What it buys is the only honest answer to the question every confident call in a strategy meeting implicitly asks: how much should anyone actually trust this gut, based on its track record, not its tone.

If you're building or auditing an experimentation program and want an outside read on this, get in touch.

Share this article
LinkedIn (opens in new tab) X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified in behavioral economics. Led 100+ in-house experiments at NRG in 2025, with project evidence and limits documented in the case studies.