Atticus Li has led enterprise experimentation portfolios where “win rate” only became useful after the team fixed its denominator and decision rules. This guide shows how to report that process metric without turning one company’s history into an industry benchmark.
The most dangerous question a VP can ask about your experimentation program is: "What percentage of your tests win?"
It is not a bad question. It is an incomplete one.
A reported win rate changes with the denominator. Did the team count every launch, only tests that reached their sample target, directional results, guardrail failures, and iterations of the same idea? It also changes with the decision rule, traffic, metric variance, and the size of effect the program is equipped to detect.
That is why I no longer compare a program with a universal industry benchmark. Two teams can report the same percentage while operating with very different evidence quality.
Why Most Tests Lose (And That's Fine)
Experimentation is a search process. You're systematically testing hypotheses about what will change customer behavior. Most hypotheses are wrong. That's not a flaw in the program — that's the nature of trying to change human behavior in complex systems.
Think about it from first principles. If you could reliably predict which changes would improve conversion rates, you wouldn't need to test. You'd just implement them. The whole point of experimentation is that you don't know. You have informed guesses, but you're operating in uncertainty.
A positive result means the test met the program's pre-agreed decision rule on its primary outcome. It does not automatically mean the change produced durable revenue, and a null result does not prove that the two experiences are identical.
A null or negative result can still improve a decision when the test was adequately powered, the implementation was valid, and the team had agreed what it would do with the answer. Without those conditions, calling every loss a “learning” is just as misleading as calling every positive metric movement a win.
Why Stakeholders Misunderstand Win Rates
The misunderstanding comes from a reasonable but incorrect mental model. Stakeholders think of experiments like projects. Projects have a success rate. If 85% of your projects failed, you'd have a serious execution problem.
But experiments aren't projects. They're questions. And in science, most questions yield null results. That's not failure — that's the process of elimination that makes the eventual discoveries valuable.
The other mental model problem is comparison to marketing channels. If your paid ads have a positive ROAS, that feels like a "win rate" close to 100%. So why can't experimentation do the same? Because paid ads operate on known mechanics — you're buying attention, not discovering new knowledge. Experimentation is research and development. The uncertainty is the whole point.
The useful reframe for executives is: “We are not trying to maximize a percentage. We are trying to improve the quality and economics of the decisions this portfolio supports.”
Portfolio Thinking: The Right Frame for Leadership
The way to present win rates to leadership is through portfolio thinking. Don't talk about individual test outcomes. Talk about the cumulative impact of the portfolio.
If you use a real portfolio example, keep it on its named proof surface with the company, period, counting rule, evidence class, and limitations attached. In a general guide, the transferable lesson is the reporting structure—not the percentage or dollar total.
I do not count every stopped idea as revenue or cost avoidance. A counterfactual requires evidence: the decision that would otherwise have been made, the population exposed, the plausible downside, and confidence that the organization would have shipped it. When that trail is missing, the honest label is “decision informed,” not “money saved.”
Present realized revenue, internally validated modeled impact, and risk reduction as separate categories. Win rate then becomes one diagnostic input rather than the headline.
What a Portfolio Rate Can—and Cannot—Say
A portfolio rate is useful for understanding one program over time when its counting rules remain consistent. It does not prove that a framework caused the rate or that a different portfolio should target the same number.
The program invested in customer data analysis, session recordings, qualitative feedback, competitive analysis, and prior test evidence before launch where those inputs were available. That work made hypotheses easier to scrutinize and measurements easier to defend.
It may also filter the pipeline toward ideas with stronger observed signals. But that is a process explanation, not a causal claim about why the final percentage landed where it did.
This is why I push back on “just run a quick test” when the decision, metric, sample requirement, or implementation path is unclear. More launches do not create more learning if the evidence bar falls as volume rises.
The operating goal is more qualified experiments at the same evidence bar—not the highest possible win rate.
What an Unusually High Win Rate Should Trigger
An unusually high reported rate is not proof of quality or manipulation. It is a reason to audit the denominator, power, peeking rules, metric selection, exclusions, and whether bug fixes were counted as experiments.
The portfolio should include decisions with meaningful upside, but “bolder” is not automatically better. Expected value depends on probability, magnitude, implementation cost, customer risk, and the quality of evidence the test can produce.
That is why I review a value-weighted portfolio: what decision was exposed, what effect was detectable, what happened downstream, and what uncertainty remains.
How to Communicate This to Your Organization
Here's the talk track I've refined over years of presenting to leadership.
Start with the portfolio result and its evidence label: the period, denominator, primary-outcome rule, and whether the financial number is observed, modeled, or recognized after launch.
Explain the mechanism. “Tests that met the pre-agreed positive-primary-outcome rule are separated from stop, iterate, inconclusive, invalid, and measurement-repair decisions.”
Reframe the win rate. “This figure is an internal process diagnostic under our counting rules, not an industry score.”
Address the implicit question. “A higher rate would not, by itself, prove a better program. We would inspect the value, risk, power, and downstream evidence behind it.”
End with the decision. “Here is what we scaled, stopped, repaired, or learned—and which numbers are observed, modeled, or still uncertain.”
Win rate is a process metric. Revenue impact is the outcome metric. Keep the conversation on outcomes, and the win rate takes care of itself.
_Want to make sure your experiments are properly sized to detect real winners? GrowthLayer's A/B test calculator helps you plan and analyze tests without losing the decision context._
FAQ
What is a good experimentation win rate?
There is no universal rate. A useful comparison requires the same denominator, primary-outcome rule, evidence threshold, test mix, and treatment of invalid or inconclusive results.
Should a null result count as a loss?
Label it as null or inconclusive according to the pre-specified rule. Collapsing every non-positive result into "loss" hides whether the test ruled out a meaningful effect.
How should leaders judge a program instead?
Review decision quality, evidence coverage, customer risk, implementation cost, and reconciled business outcomes. Microsoft's research on common interpretation pitfalls helps explain why a single rate is inadequate.
Contact me if you need help defining a decision-ready scorecard for an experimentation portfolio.