Decide What Counts as a Win Before You Test

Meta description: Medicine proved that picking your primary metric after seeing the data is a structural bias. The five-minute fix most experimentation programs skip.

TL;DR

  • A test finishes. The metric everyone agreed to watch is flat. But average order value moved, or a secondary funnel step improved, and the room quietly reframes what "the win" was — after seeing which number happened to move.
  • The cleanest large-scale evidence that this specific sequence is a real bias, not a harmless judgment call, comes from medicine: large NIH-funded trials that hadn't registered a primary outcome in advance reported a significant benefit far more often than the same kind of trials after registration became mandatory.
  • The fix isn't fewer metrics. It's naming, in writing, which one metric decides ship or kill, what bar it has to clear, and what happens for each outcome — before the test starts, when no result yet exists to make one candidate metric look more flattering than another.
  • Every other metric still matters. It just can't retroactively become the proof for a test it wasn't named to decide — it can only generate the next hypothesis.
  • The actual skill isn't precommitting itself. It's choosing, in advance, the metric that's genuinely closest to the business decision at stake — not the one that happens to be easiest to move.

A test wraps. The metric everyone agreed mattered when the test was planned — checkout conversion, say — is flat, within noise, a clean null. In the same readout, average order value ticked up. Or time-to-first-value improved. Or a downstream retention cohort looks slightly better, though it's early. Someone reasonably says: "isn't that actually what we cared about?" The room nods. The test gets written up as a win, on the metric that happened to move, and nobody in the room did anything that feels like misconduct. That's exactly what makes this failure mode durable — it doesn't require anyone to act in bad faith, only to make a completely natural decision in the wrong order.

The cleanest evidence this is a real bias, not a judgment call

Medicine ran, unintentionally, one of the largest natural experiments available on this exact question. Before the year 2000, researchers running large clinical trials were not required to publicly register which outcome they considered primary before the trial began. After 2000, prospective registration on a public registry — ClinicalTrials.gov — became standard practice for major trials. Kaplan and Irvin's 2015 analysis in PLOS ONE, "Likelihood of Null Effects of Large NHLBI Clinical Trials Has Increased over Time," found that 17 of 30 trials (57%) published before 2000 reported a statistically significant benefit on their primary outcome — compared to just 2 of 25 (8%) published after 2000.

The biology being studied didn't change in 2000. The treatments, the patient populations, the underlying science — none of it shifted because a registry requirement went into effect. What changed was narrower and more specific: researchers lost the ability to look at everything they'd measured and, after the fact, designate whichever outcome looked best as "the" primary result. Prospective registration doesn't make anyone more honest in the moment they're analyzing data. It removes the moment of choice from after the data arrives and forces it to before, which is the entire mechanism. The authors are careful to note this is a strong association rather than an experimentally proven cause — other things about trial conduct changed across those decades too — but the size and direction of the shift, on exactly the outcome you'd predict from the mechanism, makes this some of the best real-world evidence available for a bias that's otherwise mostly discussed in the abstract.

The same mechanism, running quietly in a normal analytics stack

Almost no CRO or growth program operates with anything as formal as the pre-2000 clinical-trial process — but most operate with its structural equivalent. A test launches against a loosely stated goal ("improve checkout"), and a modern analytics stack surfaces a dozen or more numbers that moved by the time it's read out: conversion, AOV, time on page, scroll depth, a funnel step two clicks downstream, a retention cohort that's too early to trust but is right there in the dashboard anyway. Nothing was pre-registered as _the_ metric. So when the obvious one is flat, there's no structural obstacle to the room settling on whichever of the other dozen moved in the direction everyone was hoping for. That's not a smaller, informal version of what pre-2000 trials did. It's the identical mechanism, running by default because nobody removed the moment of choice from after the results arrived.

It's worth being precise about how this differs from a related bias already worth watching for in any testing program: an already-selected "winning" metric being inflated relative to its true effect, simply because it was chosen as the best-looking result out of several noisy candidates. That's a real and separate phenomenon. This one happens earlier and is arguably more consequential — it's not about the size of the metric that won being overstated, it's about _which_ metric gets to be called the win in the first place, chosen after the fact from whichever one happened to move. In practice the two compound: outcome-switching selects a flattering metric to call primary, and the standard selection-inflation problem then overstates how large that metric's true effect actually is. Two separate biases, stacking in the same direction, from two different moments in the same process.

The fix: name it before you look

The fix costs nothing to run and doesn't require new tooling — it requires deciding three things in writing before the test launches, not after the readout:

  1. The one metric that decides ship or kill. Not the one metric you'll look at — the one metric whose result determines the decision. Everything else is still worth watching; it just doesn't get a vote on this test's verdict.
  2. The bar it has to clear, stated as a specific direction and magnitude, not "it should go up."
  3. What happens for each outcome, decided for all three realistic cases: it clearly clears the bar, it clearly doesn't, or it lands in the ambiguous middle. Naming the action for the ambiguous case in advance matters most, because that's the case most likely to trigger exactly the after-the-fact metric search this whole practice exists to prevent.

None of this requires ignoring the other numbers on the dashboard. A secondary metric moving in an interesting direction is a legitimate, valuable thing to notice — it's a candidate hypothesis for the _next_ test, where it can earn the same precommitment treatment before being trusted. What it can't do is retroactively become the proof for the test that was actually designed and powered around a different question. Keeping that boundary explicit is the entire practice. Blurring it is how a null result quietly becomes a win without anyone deciding, on purpose, that it should.

The judgment call underneath the mechanics

Precommitting a metric is not, by itself, a hard skill — anyone can lock in a number in a planning doc. The actual skill is choosing, in advance, which metric is genuinely closest to the business decision this test is supposed to inform, rather than whichever one is easiest to move or most likely to look good.

A team optimizing for a good-looking readout will gravitate toward a shallow, upstream metric — a click, a scroll, a micro-conversion — because those move easily and often, and an easy win is more pleasant to report than an honest null. A team optimizing for a decision that will actually hold up gravitates toward the metric closest to the outcome the business cares about, even when that metric is noisier, slower to move, and more likely to come back flat. That choice has to be made before the test runs, precisely because it's much harder to make honestly once a shallow metric has already moved in a flattering direction and a deeper one hasn't. The senior version of this practice isn't the discipline of precommitting. It's the judgment of knowing which metric deserved the commitment in the first place.

FAQ

Doesn't locking in a metric early just mean I might miss something important?

It doesn't mean missing it — it means categorizing it correctly. Anything unexpected that moves is still visible, still worth discussing, and still a legitimate seed for the next test. Precommitment only determines what's allowed to close out _this_ test's verdict. An interesting secondary movement earns its own precommitted test before it gets to be called a finding rather than an observation.

What if the primary metric is flat but something big and unexpected shows up elsewhere?

Treat it exactly as valuable as it is: a strong new hypothesis, not a result. Write it down, and if it's promising enough, design a test where it's the named primary metric next. What it shouldn't do is get reported as this test's outcome — that's the precise substitution this practice exists to prevent, and it's also usually the moment where a program's reported win rate quietly stops meaning what people assume it means.

How is this different from choosing when to stop a test?

They're separate decisions about separate questions. Stopping-rule methodology — covered in depth elsewhere in this series — governs _when_ you're allowed to look at a result without inflating your false-positive rate. This is about _what you're allowed to call the result once you do look_ — which of possibly many measured numbers gets to be the verdict. A test can have a perfectly rigorous stopping rule and still fall into outcome-switching if the primary metric wasn't named until after the data arrived.

How does this relate to deciding how much evidence a bet needs?

They answer different questions at different points in the process. The Confidence Tier Model helps size a bet against however much evidence you actually have, once you know what the evidence says. This practice determines what's allowed to count as "what the evidence says" in the first place. Skipping this step doesn't just risk a wrong tier assignment — it risks tiering a metric that was never the right one to be measuring the decision against.

Does this only apply to formal A/B tests?

No — it applies to any decision where more than one plausible metric could be used to declare success, which is most consequential business decisions, not just controlled experiments. A new hire's first-quarter review, a feature launch, a pricing change rolled out without a holdout — all of them have multiple candidate metrics available after the fact, and all of them benefit from the same fix: name the one that decides the verdict before you have a result that could bias the choice.

Related reading: The Confidence Tier Model, Sequential Testing and the SPRT, Why Most "Wins" Don't Replicate: The Winner's Curse.

Bottom line

Medicine didn't fix its outcome-switching problem by asking researchers to be more honest. It fixed it by moving one decision — which metric counts — to before the data existed to bias it, and the reported result changed by a wide enough margin, at a large enough scale, to make the mechanism hard to argue with. Nothing about a growth or product team's version of this problem is different in kind, only in formality. Naming the metric, the bar, and the action for every outcome, in writing, before the test starts, is a five-minute habit with the same effect: it doesn't make anyone more honest in the room. It just removes the moment where honesty would have been the only thing standing between a flat result and a win that was chosen, not found.

If you're building or auditing an experimentation program and want an outside read on this, get in touch.

Share this article
LinkedIn (opens in new tab) X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified behavioral economist (1 of ~1,000 worldwide). 200+ A/B tests across energy, SaaS, fintech, e-commerce, and marketplace verticals.