What Intelligence Analysts Know About Evidence That Growth Teams Don't
TL;DR
- Intelligence analysts almost never get a single conclusive source. Their entire discipline is built around weighing partial, sometimes-contradictory evidence — which is exactly the position most growth and product teams are in, whether they admit it or not.
- Source reliability decays. A source that was accurate a year ago isn't automatically accurate now. The direct business parallel: a "learning" from an old test is a source whose reliability needs to be re-checked, not a permanent fact.
- The most dangerous input isn't a weak signal — it's a single strong, coherent, uncorroborated one. Analysts have a name for the trap of trusting a story because it's compelling rather than because it's verified.
- Analysis of Competing Hypotheses (ACH), a real structured technique from CIA tradecraft, forces you to list every plausible explanation and hunt for evidence that disconfirms each one — before you let the most obvious explanation win by default.
- None of this requires more data. It requires a different discipline for weighing the data you already have — which is the same idea behind the Confidence Tier Model, applied to how a single piece of evidence earns trust in the first place.
Intelligence analysts make consequential calls off incomplete, contradictory, and often unverifiable information as a matter of daily routine. Nobody hands them a controlled experiment with a clean p-value. They get a defector's testimony, a satellite image, an intercepted communication, and a human source who may or may not still be reliable — and they have to produce a judgment anyway, one that someone downstream will act on.
That's not a metaphor for growth work. It's the same job. A growth team gets a viral social post, a handful of enthusiastic session replays, a support ticket pattern, and an underpowered test — and has to decide whether to act. The intelligence community has spent decades building formal tradecraft for exactly this situation, most of it documented in the open literature, and almost none of it has made its way into how growth teams talk about evidence. It should.
Evidence weighing is the discipline of deciding how much trust a piece of information has earned — from its source, its independence, and what it would take to disprove it — before you let it change a decision.
Why intelligence tradecraft is the right model, not just an analogy
The comparison holds up for a structural reason, not just a stylistic one: both fields are defined by the same constraint. You cannot re-run history to get a clean, controlled answer. An analyst can't rewind Cuba in 1962 and try again. A growth team, most of the time, can't wait the six months a fully powered test would take before the market has already moved on. Both are forced into acting on convergent, imperfect evidence rather than waiting for a single decisive source that will never arrive.
The foundational text here is Richards Heuer's _Psychology of Intelligence Analysis_, published by the CIA's Center for the Study of Intelligence and still used in analyst training (the full text is also freely available via the Internet Archive). Heuer's central argument is that most analytical failures aren't failures of missing information — they're failures of _cognitive processing of the information analysts already had_. The mind forms an impression early, and then every subsequent piece of evidence gets fit to confirm it. That is precisely the failure mode of a growth team that decides a hypothesis is true after one good week of data and spends the rest of the test looking for reasons to believe it.
Source reliability decay: why an old "learning" isn't a permanent fact
Every intelligence service maintains formal reliability ratings for its sources — codified in NATO and U.S. intelligence doctrine as the "Admiralty Code," which rates a source's _track record_ (A through F, "completely reliable" to "reliability cannot be judged") separately from the _plausibility_ of a specific report (1 through 6). The rating on the source is not permanent. A human source who was consistently accurate for years can become compromised, start being fed disinformation, or simply lose access to what they used to know — and a good analyst re-evaluates that rating on a schedule rather than assuming yesterday's reliable source is still reliable today.
Growth teams treat "learnings" the exact opposite way. A test result from eighteen months ago gets cited in a strategy deck as settled fact, with no one asking whether the market, the user base, the competitive landscape, or the product itself has changed enough to have quietly invalidated it. The learning was accurate — at the time, in that context, with that source of traffic. None of those conditions are guaranteed to still hold.
The practical fix: treat every stored "learning" in your test-and-learn backlog as a rated source, not a fact, and put an explicit decay clock on it. A pricing-sensitivity finding from a high-inflation period should get re-checked before being reused in a low-inflation one. A UX preference documented pre-mobile-redesign shouldn't quietly anchor a decision on the current mobile experience. This is a five-minute discipline — "when was this learned, and has anything changed that would affect it" — that almost no team runs formally, and it's the single most valuable habit borrowed from source-handling tradecraft.
The single-source trap: when a compelling story is the danger sign, not the reassurance
Heuer's most cited finding is about vividness, not volume. Analysts — like everyone — are disproportionately swayed by information that's concrete, detailed, and tells a coherent story, regardless of whether that information has actually been corroborated. A single vivid human-intelligence report describing exactly what leadership expected to hear can outweigh a dozen ambiguous signals intelligence pointing the other way, purely because the vivid report is more _cognitively satisfying_, not because it's more _true_.
This has a precise growth-team equivalent: the viral social post, the one customer call where someone said exactly what you hoped to hear, the single session replay that perfectly illustrates the UX problem you already suspected. These are dangerous not because they're worthless — they can be genuinely useful directional signal — but because their vividness makes them feel far more conclusive than their evidentiary weight actually supports. A coherent story is not the same thing as a verified one. The tell is emotional: if a single anecdote makes a room go quiet and nod, that's the moment to slow down and ask what corroborates it, not the moment to greenlight the roadmap change.
The corrective isn't to distrust vivid evidence. It's to explicitly separate two questions that get merged by default: "does this feel true" and "is this corroborated." A single compelling source should raise your interest and lower your certainty at the same time — it tells you where to look next, not what to conclude yet.
Weighing a strong uncorroborated signal against several weaker converging ones
This is where the tradecraft gets genuinely counter-intuitive, and where it diverges most sharply from how most teams instinctively weigh evidence.
| Evidence pattern | Intelligence tradecraft treatment | Growth/CRO equivalent |
|---|---|---|
| One strong, vivid, uncorroborated source | Treated with active suspicion — flagged for corroboration before use in a judgment, regardless of how convincing it reads | One viral post, one glowing customer call, one dramatic session replay — high emotional pull, low standalone weight |
| Several weak, independent, converging signals | Often treated as stronger evidence collectively than one strong source alone — independence of origin matters more than individual strength | A support-ticket pattern + a soft directional test result + a sales-call theme, all pointing the same way despite each being individually inconclusive |
| Several weak signals that all trace back to one common origin | Explicitly discounted — this isn't corroboration, it's the same information counted multiple times | Multiple "signals" that are actually all downstream of one influential customer, one viral thread, or one internal advocate repeating the same anecdote |
The key discipline in that middle row is checking for genuine independence. Three session replays that all came from users who arrived via the same viral post aren't three data points — they're one data point wearing three costumes. Analysts are trained to trace a signal back to its origin before counting it as corroboration, precisely because the appearance of multiple sources is one of the easiest ways to be fooled into overconfidence. A growth team doing the same check — "are these three signals actually independent, or downstream of the same root cause" — closes one of the most common gaps between what feels like triangulated evidence and what actually is.
Analysis of Competing Hypotheses: forcing the alternative explanation onto the table
Heuer's tradecraft also produced a specific structured technique now taught throughout the intelligence community: Analysis of Competing Hypotheses (ACH). The method is deliberately mechanical, because Heuer's research found that unaided human judgment reliably jumps to the first plausible explanation and then spends its remaining effort confirming it rather than testing it. ACH is designed to short-circuit that instinct.
The technique, generically: list every plausible hypothesis that could explain the evidence — not just the obvious one — before looking at any evidence in detail. Then build a matrix of evidence against every hypothesis, and instead of asking "what supports my favorite explanation," ask "what evidence is _inconsistent_ with each hypothesis." The hypothesis that survives the most attempts at disconfirmation, not the one with the most supporting evidence, is the one that wins. Supporting evidence is cheap — almost any hypothesis can collect some. Evidence that actively rules a hypothesis out is what actually discriminates between competing explanations.
Translated to a growth context: when conversion drops on a page, the obvious hypothesis ("the new checkout flow hurt conversion") gets all the attention by default. A one-page ACH exercise forces you to also seriously list "seasonality," "a tracking regression," "a traffic-mix shift from a paid campaign change," and "a competitor promotion" — and then actively hunt for evidence that would rule each one out, rather than stopping at the first story that fits. Most teams stop as soon as they find a hypothesis that's _consistent_ with the data. The diagnostic catch that separates a rigorous read from a fast one is asking which competing explanation the data would look like if it were true instead — and checking whether you can actually rule that one out.
Why this matters more as evidence gets noisier, not less
There's a reason this tradecraft was built for intelligence work rather than laboratory science: it's designed for domains where you can't control the conditions, can't rerun the event, and have to act anyway. That description fits more of the modern growth environment every year — fragmented attribution, thinning third-party data, faster market cycles that don't allow for the traffic volume classical testing assumes. Teams under those constraints are already living in the intelligence analyst's world, whether or not they've adopted the analyst's discipline for it. Adopting the tradecraft without the underlying uncertainty would be overkill; refusing to adopt it while facing the same uncertainty is just as much of a mismatch, in the other direction.
FAQ
Isn't this overkill for a growth team that isn't making life-or-death decisions?
The stakes are different, but the evidentiary problem is identical: incomplete, sometimes-contradictory information, and a decision that has to get made anyway. The tradecraft is proportional — a one-page ACH matrix for a genuinely ambiguous conversion drop takes twenty minutes, not a formal intelligence workflow. Use the rigor where the decision size warrants it, same as any other methodology choice.
How do I check whether "converging signals" are actually independent?
Trace each one back to its origin before counting it. Ask where the traffic, the customer, or the anecdote came from. If three data points all lead back to the same viral post, the same influential account, or the same internal advocate repeating a story, that's one source, not three — treat it accordingly and go find a genuinely separate signal before upgrading your confidence.
What's the fastest version of Analysis of Competing Hypotheses I can actually use?
List three plausible explanations for the pattern you're seeing, not just the obvious one, before you look closely at any of them. Then for each, write down one piece of evidence that would have to be true if that explanation were correct — and go check whether it is. The discipline is in seriously trying to rule hypotheses out, not in generating a long list for its own sake.
Does this replace statistical testing, or sit alongside it?
Alongside it. Statistical testing tells you how confident you can be in a specific comparison; evidence-weighing tradecraft tells you how to combine that test result with everything else you know — old learnings, qualitative signals, anecdotes — without being fooled by vividness or false corroboration. They answer different questions.
How does source reliability decay interact with a test result that already reached significance?
A significant result describes what was true in the conditions it was measured under. If the underlying context has shifted meaningfully since then — a different acquisition channel mix, a redesigned adjacent flow, a changed macro environment — the statistical confidence in the original result doesn't transfer automatically to the current moment. Re-verify before reusing, the same way an analyst re-rates a source before trusting its next report.
Related reading: Triangulation Over Isolation, Calibration Training.
Bottom line
Intelligence analysis and growth work share the same core constraint: you rarely get one clean, conclusive source, and you have to make a call anyway. The discipline that's been built for that constraint — rating source reliability instead of assuming permanence, treating vivid single anecdotes with more suspicion rather than less, checking that converging signals are actually independent, and forcing yourself to seriously consider the competing explanation before committing to the obvious one — costs almost nothing to adopt and catches a specific, recurring class of mistake. None of it requires more data. It requires trusting the data you already have more precisely.
If you're building or auditing an experimentation program and want an outside read on this, get in touch.