The Meta-Analysis Your Experimentation Program Is Missing
Meta description: Most programs audit individual tests, almost none audit the program itself. A quarterly portfolio audit answers what leadership actually wants asked.
TL;DR
- A product manager says: "Users want better deals." A brand marketer says: "TV is driving more direct demand." A performance marketer says: "This channel has a strong ROAS." Finance says: "But is this incremental?" Product says: "Will this hurt user trust?" Leadership says: "Should we scale this?" — every one of those claims is being made about an individual initiative. Almost nobody is asking the equivalent question about the _program_ making all these calls.
- Most experimentation teams rigorously audit individual tests but never audit the portfolio of tests as a system, the way a portfolio manager audits a fund's overall performance rather than any single trade.
- The Experimentation Portfolio Audit is a named framework with four components: calibration tracking, pipeline health, hit-rate trend, and idea-source diversification.
- This is the level of question CEOs, CFOs, and CMOs actually want answered about a testing program, but rarely know how to ask for explicitly — "is this program getting better at making decisions" rather than "did this test win."
- Run it quarterly. It's a different cadence and a different unit of analysis than a test readout, and conflating the two is why most programs never see their own blind spots.
A product manager says: "Users want better deals." A brand marketer says: "TV is driving more direct demand." A performance marketer says: "This channel has a strong ROAS." Finance says: "But is this incremental?" Product says: "Will this hurt user trust?" Leadership says: "Should we scale this?"
Every one of those is a claim about a single initiative, and every experimentation program in existence has a process for adjudicating exactly this kind of claim — run the test, read the result, make the call. What almost no program has is a process for asking the same set of questions about _itself_. Is our program's hit rate actually improving, or does it just feel that way? Is finance's skepticism about incrementality something we should be more worried about than we currently are, across the portfolio, not just on this one campaign? Nobody schedules that conversation, because it doesn't fit the cadence of any single test readout. It needs its own cadence, and its own framework.
The blind spot: everyone audits the trade, nobody audits the fund
Picture how a portfolio manager operates versus how most experimentation programs operate. A portfolio manager doesn't just ask "did this trade make money." They track, across every position: was the original thesis right, and by how much? Is the pipeline of new investment ideas healthy, or is the fund running out of genuinely new theses and just re-trading old ones? Is the win rate trending up, flat, or down over the last several quarters? Are ideas coming from diversified research, or is the whole fund making correlated bets off one analyst's worldview?
Now look at how most experimentation programs operate. Team runs test. Test wins or loses. Team reports the individual result. Team moves to the next test. Multiply that by a hundred tests a quarter, and what you have is a hundred individually-audited trades and a completely un-audited fund. Nobody is asking whether the program itself is improving as an instrument for making decisions — only whether this quarter's crop of tests happened to win.
That's the actual gap. It's not that companies lack a testing culture. Many have a mature one at the level of the individual test — proper randomization, reasonable sample sizes, clean reporting. What's missing is the equivalent discipline one level up: treating the whole set of tests as a portfolio with its own health metrics, separate from any individual result.
The Experimentation Portfolio Audit
This is the named framework: a structured, quarterly review of the program as a system, built from four components.
1. Calibration tracking
Before every test ships, someone had a belief about how it would go — a predicted lift, or at minimum an implicit probability of success ("I'm pretty confident this wins"). Calibration tracking means logging that prediction _before_ the result comes in, then comparing it systematically against what actually happened, across every test in the portfolio, over time.
This is the direct, practical payoff of the "winner's curse" phenomenon well-documented in experimentation literature — the pattern where the estimated effect size of a test that clears significance is, on average, inflated relative to the true effect, precisely because it cleared a significance threshold partly by chance as well as by real effect. Most programs know about the winner's curse as a concept. Almost none of them have measured _their own_ inflation factor. Calibration tracking answers a concrete, useful question: when your team predicts a moderate lift, does reality typically come in lower? By how much, roughly, and is that gap shrinking or growing over time? A program that has never measured this has no idea whether its own forecasts — the ones finance is quietly building projections on top of — are honest or systematically optimistic.
2. Pipeline health
Win rate is the metric everyone tracks. Almost nobody tracks whether the pipeline of hypotheses feeding that win rate is healthy. A program can maintain a perfectly respectable win rate while quietly running out of genuinely novel ideas — re-testing minor variants of the same three hypotheses because that's what's easy to generate, rather than sourcing new ones.
Pipeline health asks: over the last few quarters, what fraction of new tests represent a genuinely new hypothesis, as opposed to a small variation on something already tested? A declining rate of novel hypotheses is a leading indicator of trouble that a stable win rate will hide for a surprisingly long time — because a program that keeps re-testing safe, well-understood variants can sustain a good win rate right up until it has nothing left to test that matters.
3. Hit-rate trend over time
Not the win rate in isolation — the _trend_ of the win rate, tracked over enough quarters to see a direction. Is the program actually getting better at picking winners as it accumulates institutional knowledge about what works in this specific business, or is it flat, or quietly declining?
A flat or declining hit-rate trend, even with an acceptable absolute win rate, is worth investigating on its own. It can mean the easy wins have already been captured and the program hasn't adjusted its hypothesis quality to compensate, or it can mean the same idea sources are being mined past the point of diminishing returns (which connects directly to component four).
4. Idea-source diversification
Where are hypotheses actually coming from? A healthy portfolio draws from a mix: quantitative funnel analysis, qualitative research (support tickets, sales calls, user interviews), and competitive observation. An unhealthy one is quietly dominated by one source — often whichever team or individual is loudest, or whichever data is easiest to pull — and the correlated nature of that source means the whole hypothesis pipeline is making a version of the same bet repeatedly, even when it looks diversified on a roadmap slide.
| Component | Question it answers | Why win rate alone misses it |
|---|---|---|
| Calibration tracking | Is our forecasted lift honest, or systematically inflated? | Win rate says a test won; it says nothing about whether the predicted magnitude was realistic |
| Pipeline health | Are we still generating genuinely new hypotheses? | A shrinking pipeline can sustain a fine win rate for several quarters before it shows up as a problem |
| Hit-rate trend | Are we getting better at this, or coasting? | A single quarter's win rate can't show a trend; only a multi-quarter view can |
| Idea-source diversification | Are we making one correlated bet repeatedly? | A long roadmap can look diverse while every hypothesis traces back to the same source |
Why this is the question leadership actually wants asked
CEOs, CFOs, and CMOs rarely ask for "a portfolio audit of the testing program" in those words, because most of them have never seen the framework named. What they do ask, constantly, in slightly different phrasing, is some version of: "Is this program actually getting smarter, or are we just running a lot of tests?" That is precisely the question a pile of individual test readouts cannot answer, no matter how many of them you hand over. A hundred clean individual readouts tells leadership that a hundred decisions were made carefully. It tells them nothing about whether the underlying instrument making those decisions is improving, plateauing, or drifting toward false confidence. For more on why individual test rigor doesn't guarantee program-level rigor, _Trustworthy Online Controlled Experiments_ — written by the team that ran experimentation at Microsoft, Bing, and LinkedIn — remains the standard reference for what rigor looks like at scale.
This is also the natural extension of the Confidence Tier Model: that framework governs how much certainty a single bet needs before you size it. The Portfolio Audit is the same discipline applied one level up — not "how confident should we be in this test," but "how confident should we be in our own confidence, given how this program's predictions have actually performed over time."
Running it quarterly
The cadence matters as much as the content. A quarterly audit is frequent enough to catch a declining pipeline or a calibration drift before it becomes a credibility problem with finance, but infrequent enough that you're looking at a real trend rather than noise from a handful of recent tests. Trying to run this monthly usually just re-measures the same few tests repeatedly and mistakes short-term variance for a trend. Running it annually means a full year can pass with a quietly degrading pipeline before anyone notices.
In practice, the audit is a standing quarterly agenda item, distinct from any individual test readout: pull every prediction and result from the quarter for calibration tracking, tag each new hypothesis by source and novelty for pipeline health and diversification, and plot the trailing win rate against the last several quarters for the trend. None of the four components require new tooling most programs don't already have — the tests were already logged. What's missing is almost never data. It's the standing habit of looking at the portfolio as a portfolio.
FAQ
How is this different from a normal quarterly business review?
A normal QBR typically reports outcomes — revenue impact, notable wins, roadmap for next quarter. The Portfolio Audit is specifically about the health of the decision-making system itself: is it calibrated, is its pipeline healthy, is its hit rate trending in a good direction, is it drawing from diverse sources. It can feed into a QBR, but it's a different unit of analysis than a results summary.
What's a healthy hit rate, and does this framework require a specific number?
No, deliberately. A "good" hit rate varies enormously by vertical, maturity of the product, and how bold the hypotheses being tested are — a program running only safe, incremental tests should have a higher win rate than one deliberately testing bolder bets, and a higher win rate in that case isn't necessarily better. The framework cares about the trend and the calibration, not a universal benchmark number.
We don't have enough historical data to see a trend yet. Is this still worth doing?
Start the calibration logging and source-tagging now, even before you have enough history to see a trend — the audit's value compounds. A program that starts tracking predicted-versus-actual lift this quarter will have a genuinely useful calibration read within two or three quarters. The programs that never start never get that visibility at all.
Doesn't tracking calibration create an incentive to sandbag predictions to look good later?
It's a real risk, which is why the predicted lift needs to be logged before the result is known, ideally in a system the predictor can't quietly edit after the fact, and why the audit should be framed as improving the instrument rather than grading individuals. The goal is an honest read on the program's forecasting accuracy, not a performance review of whoever made the prediction.
Who should own running this audit?
Whoever owns the experimentation program's methodology, not whoever owns any individual test. It needs someone with visibility across the whole portfolio and enough seniority that the findings — especially an uncomfortable calibration gap or a shrinking pipeline — get acted on rather than filed away.
Related reading: Why Most "Wins" Don't Replicate: The Winner's Curse, Calibration Training.
Bottom line
Every program already knows how to audit a test. Far fewer know how to audit themselves, because the muscle for it doesn't get built by running more tests — it gets built by deliberately stepping back from any single result and asking whether the whole system is improving. The Experimentation Portfolio Audit is that step back, made concrete: four questions, one quarterly cadence, and an honest look at whether the program is actually getting better at making bets or just staying busy making them.
If you're building or auditing an experimentation program and want an outside read on this, get in touch.