Atticus Li has led enterprise experimentation programs where post-hoc metric selection threatened the credibility of otherwise sound analysis. This guide explains the failure mode and the pre-commitment system used to prevent it.
I need to tell you about the thing that almost killed my experimentation program. Not a bad test. Not a platform failure. Not budget cuts. It was something much more insidious: the slow, invisible erosion of credibility that happens when your team starts shopping for winning metrics after seeing results.
Post-hoc metric shopping is when you change your success metric after seeing the data to make a test look like a winner. The behavior can emerge without anyone recognizing it as a problem because every individual metric may still be real.
What Metric Shopping Actually Looks Like
Here is an illustrative scenario. You run a test on a landing page redesign. Your primary metric is form submissions. After the planned multiweek window, the result is flat—no meaningful difference in form submissions between control and variant.
But someone on the team notices that time on page moved in the low-teens. Another person sees that scroll depth increased. Someone else pulls up a secondary click metric that moved in the high-single-digits.
Suddenly the readout becomes: “The redesign did not affect form submissions, but engagement moved. We recommend shipping based on the secondary signals.”
That's metric shopping. And it feels completely rational in the moment. The test clearly did _something_. The data is real. Nobody is fabricating numbers. But the conclusion is manufactured.
You decided the test was about form submissions. The form submission rate didn't move. The test lost. Everything else is noise until you design a follow-up to test those engagement hypotheses explicitly.
Why Every Team Does It
I've watched this pattern play out in multiple organizations. The recurring cause is pressure.
Stakeholder pressure. Someone senior championed this redesign. They allocated design and development resources. They told _their_ boss it was happening. Now you're telling them it didn't work? That's a hard conversation. It's much easier to find a metric that moved and spin a softer narrative.
Sunk cost. The test took six weeks to design, build, QA, and run. The team invested real time and energy. Calling it a loss feels like all that work was wasted. Finding a winning metric — any winning metric — makes the investment feel justified.
Program survival instinct. This is the most dangerous one. When you're building an experimentation program, especially in the first year or two, you feel enormous pressure to demonstrate value. Every test that "fails" is ammunition for someone who thinks experimentation is a waste of resources. So you start unconsciously looking for wins in every dataset.
I understand all of these pressures. I've felt every one of them. But the solution to organizational pressure is not to compromise on truth. Because that path leads somewhere much worse than a flat test.
How It Destroys Programs Over Time
Metric shopping doesn't blow up your program in one dramatic failure. It erodes it slowly, like termites in a foundation. Here's the progression I've seen play out multiple times.
Stage one: Selective storytelling. The experimentation team starts presenting results in the best possible light. Losing tests get reframed as "learnings with positive engagement signals." Win rates magically climb because you're counting these reframed tests as wins.
Stage two: Stakeholder confusion. Different teams start reporting different "wins" from the same test. Marketing says the test won on engagement. Product says it was flat on conversion. The data team can't reconcile the narratives. People start asking which metric actually matters, and nobody has a consistent answer.
Stage three: Loss of trust. Executives notice that the experimentation team always seems to find a win. They start questioning the methodology. When you present a genuine, high-confidence win — the kind that should drive major investment — it gets the same skeptical reception as your reframed engagement metrics. You've trained people to discount your results.
Stage four: Program marginalization. The experimentation team becomes "the team that always says their stuff works." You've become the center of spin instead of the center of truth. Budget gets harder to justify. Headcount requests get denied. The program dies not because it was bad, but because it destroyed its own credibility.
I've watched this progression happen in earlier roles. The timing varies, but recovery requires an explicit reset of the decision rules, reporting ownership, and incentives that allowed the spin.
The Organizational Dynamics Nobody Talks About
Here's what makes this problem so intractable: metric shopping often starts with the people who should be preventing it.
The CRO manager wants to show their boss that the program is working. The director of marketing wants to report wins to the VP. The VP wants to show the CMO that digital experimentation is driving results. At every layer, there's an incentive to find the most favorable interpretation of the data.
Early in a program's life, when you're trying to build proof of concept, the temptation is overwhelming. "We just need a few early wins to secure next year's budget, then we'll tighten up the methodology." I've heard this exact rationalization. It never works. You can't build a rigorous culture on a foundation of expedient narratives.
The other dynamic is cross-functional friction. When the experimentation team owns the methodology but a stakeholder owns the strategy, you get competing incentives. The product manager who requested the test wants it to win because they have a roadmap to defend. The experimentation analyst wants to report the truth because their professional credibility depends on it. Guess who usually loses that argument?
The One Rule That Prevents It
After all the programs I've built and consulted on, I've landed on one rule that prevents post-hoc metric shopping. It's brutally simple and non-negotiable.
One primary metric, locked before launch, documented in the test brief.
Every test has a single primary metric. It's chosen during the hypothesis phase, before anyone sees any data. It's written into the test brief alongside the hypothesis, the sample size calculation, and the expected runtime. Once the brief is approved and the test launches, the primary metric cannot change.
Everything else — engagement metrics, secondary conversions, segment-level cuts — is explicitly labeled as exploratory. Exploratory findings are interesting. They can inform future hypotheses. But they do not determine whether this test won or lost.
In the illustrative readout, there is one primary conclusion: “The test targeted form submission rate. The low-single-digit point estimate did not cross the pre-agreed decision threshold, and its uncertainty includes no meaningful improvement. The test is inconclusive on its primary metric.” Then, and only then, do I mention exploratory findings, clearly labeled as such.
This rule does three things. First, it forces the team to think harder about what actually matters before building anything. If you can only have one primary metric, you'd better choose the right one. Second, it removes the temptation to spin after the fact because everyone agreed on the success criterion upfront. Third, it builds the credibility that makes your real wins land with impact.
How to Handle Losing Tests Honestly
If you lock your primary metric and report honestly, you're going to report negative, null, and inconclusive results. There is no universal “healthy” win rate: it changes with the denominator, decision rule, traffic, effect sizes, and what a program chooses to test.
The key is framing losses correctly. A test that doesn't move the primary metric isn't a failure — it's a decision. You now know that this particular change, at this particular point in the funnel, does not meaningfully impact the metric you care about. That's valuable. It prevents you from investing further in a direction that doesn't work.
I structure every test readout in the same format, whether the test won or lost. Hypothesis. Primary metric result. Confidence level. Business implication. Recommended next action. The format is identical for wins and losses. This normalization is crucial because it removes the emotional charge from "losing" tests.
When a stakeholder pushes back — "But the engagement metrics went up, shouldn't we ship it?" — I have a consistent answer: "Those are exploratory findings. If we believe scroll depth is the metric that matters, let's design the next test with scroll depth as the primary metric and power it appropriately. I'm not willing to retroactively change the success criterion because that undermines every result we report."
That conversation is uncomfortable exactly once. After that, people understand the standard.
Why This Matters for Program Survival
Credibility is the experimentation team's most valuable asset. More valuable than the testing platform. More valuable than the traffic volume. More valuable than the analysts on the team.
When an experimentation team has credibility, a result with a transparent annual-impact model can inform real investment. Without credibility, even a well-labeled estimate gets discounted because stakeholders no longer trust the classification process.
The reporting system matters because positive, negative, null, and inconclusive outcomes follow the same structure and because the primary metric cannot change after results arrive. Named company totals belong on the corresponding project proof surface with their evidence limits.
Trust takes far longer to build than metric shopping takes to damage.
Build the System That Prevents the Temptation
The best way to prevent metric shopping isn't willpower — it's process. Build the system so the temptation never arises.
Lock the primary metric in the test brief. Make test briefs a required artifact before any development begins. Store every brief in a test repository that's accessible to anyone in the organization. When you present results, reference the original brief. Make the audit trail visible.
If you're running more than a handful of tests, you need tooling that enforces this workflow — intake, hypothesis documentation, metric locking, and results tied back to the original brief. That's exactly what we built GrowthLayer to do: automate the test brief process so the primary metric is locked, documented, and visible before the first line of test code gets written.
The organizations that build lasting experimentation programs are the ones that choose honesty over optics from day one. It's harder in the short term. It's the only thing that works in the long term.
FAQ
Is it wrong to inspect secondary metrics after a test?
No. Secondary and exploratory analysis can generate useful hypotheses. The error is presenting a metric chosen after results as if it were the pre-specified success criterion.
How should multiple metrics be handled?
Name the primary decision metric, list guardrails, and define any multiplicity adjustment or hierarchical rule before launch.
What should happen when an exploratory metric looks promising?
Document the finding, effect interval, and number of analyses inspected, then run a confirmatory test. Microsoft Research describes metric-interpretation and multiple-testing pitfalls.