Some experiments become uninterpretable before launch because the plan confuses an eventual recovery rate with the conversion baseline for the population the treatment can actually reach.

The mistake is easy to miss because the arithmetic can be correct while the population and time window are wrong. The dashboard renders, the number sounds plausible, and the resulting test may be badly sized.

Before blaming the creative or hypothesis, verify that the baseline describes the eligible population and decision window.

"Most new analysts make this mistake: they use what is easy to measure instead of what is correct to model."
— Atticus Li

The Setup That Traps Everyone

Here is an explicitly synthetic planning example. An analyst is asked to plan a test on a multi-step flow—a checkout, application, or signup. They pull a 90-day window and write:

  • Users who started the flow but did not finish: 4,000
  • Users who eventually completed the flow: 2,400
  • Calculated baseline: 2,400 / 4,000 = 60%

Then they write "baseline conversion rate = 60%" in the test plan doc and move on.

On paper this is correct arithmetic. In practice, that 60% number is not a baseline conversion rate. It is a 90-day eventual recovery rate, and the difference is the reason the experiment will miss.

What That 60% Actually Measures

The 60% figure answers a different question than the analyst thinks it does. It answers: _"Of the users who started the flow in the last 90 days and did not finish in-session, what fraction eventually crossed the finish line at any point during the window?"_

That is a historical recovery number. It includes users who:

  • Came back three hours later on a different device
  • Returned four weeks later after a retargeting email
  • Bookmarked the page and finished on their lunch break
  • Received a discount code from a separate lifecycle campaign
  • Completed for reasons that have nothing to do with the page you are about to test

None of those users will be affected by the button color, the headline, the form field count, or whatever else is in the variant. Your test touches a single moment. The 60% number rolls up every moment across three months.

When you use that number as the expected conversion rate of the control, you are implicitly assuming the variant will also get credit for every retargeting email, every lifecycle campaign, and every long-tail return visit. The variant will not. The assumption is the trap.

The Mental Model That Fixes It

The clean way to think about this is to break the observed behavior into the actual steps it contains:

Abandon → Return → Restart → Convert

Your warehouse data, in most setups, shows you only the ends of that chain — _abandon_ and _convert_. The middle two steps, the return rate and the conversion-after-return rate, are compressed into a single aggregate number. That aggregate number is the 60%. And it is not what your experiment influences.

Your experiment influences something smaller. Usually it influences one of these:

  • The probability of finishing in-session (no return required)
  • The probability of finishing after returning via your specific surface
  • The probability of finishing after seeing a specific variant of a specific screen

The correct baseline is the conversion rate of the slice your test can actually reach — not the blended three-month recovery rate of everyone who ever touched the flow.

An Illustrative Worked Example

The numbers below are synthetic scenario inputs, not a client or employer readout and not an industry benchmark.

The raw data (90 days, one checkout flow):

MetricCount
Total users in the system1,000,000
Users who entered the checkout6,400
Users who abandoned the checkout4,000
Users who eventually completed (any channel, any day)2,400
Weeks in the window13

The naive analysis:

baseline = 2,400 / 4,000 = 60%
traffic  = 1,000,000 users in the system

Plugged into a power calculator, that inflated traffic input can make a small relative effect look detectable on an implausibly short schedule. Everyone is excited. The test ships.

What is actually wrong:

First, the traffic number is completely off. The test cannot touch 1,000,000 users. It can only touch users who are in the checkout flow, because that is where the variant lives. So the correct traffic denominator is:

eligible traffic = 4,000 abandoners / 13 weeks ≈ 308 users/week

That is three orders of magnitude smaller than what the power calculator was fed. Once the eligible traffic replaces total system users, the detectable effect and runtime must be recalculated under the team's approved error and power targets.

Second, the 60% includes lifecycle and retargeting conversions that the test cannot influence. For this illustration, suppose path-level measurement supports the following scenario range:

BoundEstimateMeaning
Upper60%Everyone who eventually converts, from any source, over 90 days
Midpoint35–45%Users who return and convert through surfaces the test can reach
Lower25%Users who convert in-session or immediately after the test variant

The planning inputs should come from the eligible weekly traffic and measured path-level baseline. In this synthetic example, the analyst might stress-test:

baseline scenario ≈ 40%
eligible traffic  ≈ 300 users/week
effect size        = locally justified MDE or minimum business effect

Rerun the power calc with those three numbers and the picture changes completely. The detectable effect size is much larger. The runtime is much longer. And crucially, the analyst now knows that before shipping — not after.

The Two Errors Are Different, and Both Matter

Notice that there were two separate mistakes in the naive analysis, and they compound:

  1. Wrong numerator for the baseline. Conflating eventual recovery with testable conversion. This inflates the baseline and makes the variant look like it has less room to move than it does.
  2. Wrong denominator for the traffic. Using all users in the system instead of users eligible for the test. This inflates the traffic estimate by three orders of magnitude and makes the runtime look tiny.

Either error can mis-size a test. Together they increase the risk of an underpowered or misinterpreted result.

The fix belongs in the planning document before the variant consumes traffic.

The Framework I Use for Baseline Estimation

When I sit down to plan an experiment and the data is imperfect (which is almost always), I walk through four steps in order.

Step 1: Anchor on what you can measure

Start with the number you actually have. In the example, that is 60%. Do not throw it away — it is useful information, it is just not the answer. Treat it as the _upper bound_. The true testable conversion rate is almost certainly lower than this, because eventual recovery includes sources your test cannot influence.

Step 2: Discount for causal reach

Ask what fraction of the 60% is actually within the influence of the test. Retargeting emails go away. Cross-device returns go away. Lifecycle discount codes go away. Anything that would have converted the user regardless of your variant goes away.

This is not a universal discount calculation. Estimate the testable fraction from identity resolution, path analysis, return source, treatment eligibility, and the outcome window. If those inputs are unavailable, present a sensitivity range and state that the baseline remains uncertain.

Step 3: Build a range, not a point estimate

Write down a lower bound, midpoint, and upper bound supported by the available path data. Run the sample plan across the range instead of treating the midpoint as truth.

lower    = 25%   → conservative, use for minimum runtime
midpoint = 40%   → use for actual test sizing
upper    = 60%   → use only to verify sanity

In this synthetic example, the range makes the assumptions visible. A production plan should replace it with locally measured inputs where possible.

Step 4: Sanity-check the output

Before you ship, take the power calculator's output and ask two questions:

  1. Is that traffic actually going to show up? (Check seasonality, marketing spend, release calendar.)
  2. Is the chosen effect size justified by the business threshold, prior evidence, mechanism, and implementation magnitude?

If either answer is "probably not," go back and fix the inputs.

For more on how I think about decisions with incomplete data, the same framework applies at the decision stage — you are always working in ranges, not point estimates.

The Traffic Denominator Question

The denominator mistake deserves its own section because it is the one I see most often at mid-sized teams that have just gotten access to a full data warehouse.

The impulse is to use the biggest denominator available, because bigger numbers feel more statistically powerful. This is exactly backwards. The correct denominator is the _smallest population your test can actually reach_, because every user outside that population is noise.

Ask yourself a single question: "Can this user experience the variant?"

  • If they never hit the page the test lives on → no, exclude them
  • If they hit the page but are in a segment the test is not targeting → no, exclude them
  • If they hit the page on a device or browser the variant does not render on → no, exclude them
  • If they are in the holdout or a mutually exclusive test → no, exclude them

What is left is your real traffic. That is the number that goes into the power calculator. It is almost always much smaller than the analyst's first instinct, and it is almost always the number that determines whether the test is plannable at all.

What a 60% Eventual Recovery Rate Actually Tells You

Here is a more useful way to read that 60% number — not as a baseline, but as a diagnostic about the shape of your funnel.

If many abandoners eventually return and convert, investigate whether timing, lifecycle contact, cross-device behavior, or in-session friction explains the pattern. The aggregate recovery rate alone does not prove demand or identify the bottleneck.

That diagnostic can motivate several competing hypotheses:

  • Friction reduction — fewer fields, clearer steps, faster load
  • Cognitive load reduction — fewer choices, clearer language, better defaults
  • Reassurance — trust signals, objection handling, payment clarity
  • Performance and latency

Choose the outcome that matches the decision. Time to convert and in-session completion may be useful, but downstream conversion, margin, cancellations, and customer quality may still be required guardrails.

The Checklist I Run Before Every Test Plan Ships

Whenever I audit a test plan, I run through the same four questions before I sign off. If any answer is "no," the plan goes back for revision.

  1. Am I using eligible users only in the denominator? Not total users. Not monthly actives. The users who can actually experience the variant.
  2. Am I treating eventual behavior as immediate behavior? If my baseline number is a 90-day recovery rate, am I implicitly assuming the variant gets credit for 90 days of lifecycle activity?
  3. Does my baseline reflect only the conversions the test can influence? Or is it a blended number that includes sources the test cannot touch?
  4. Am I using a range instead of a single number? Lower, midpoint, upper. Plan on the midpoint, sanity-check with the edges.

A plan that clears all four is not guaranteed to win or even to be valid. It does make the eligibility, time-window, and baseline assumptions easier to inspect before launch.

What You Are Not Seeing

Four more traps I see sitting just behind the baseline mistake, because they share the same root cause — using what is easy to measure instead of what is correct to model.

1. Return behavior might be the real bottleneck. If your "eventual recovery" number is high but the return rate itself is low, the test has limited reach no matter how good the variant is. Measure return rate separately.

2. You may already be near a ceiling. A high eventual conversion rate means fewer users left to capture. Incremental gains get smaller as you approach the ceiling. Lift assumptions that were realistic at 40% baseline are not realistic at 60%.

3. Your lift assumption may be unjustified. Do not borrow a universal UX benchmark. Choose the MDE from eligible traffic and the minimum effect worth implementing, then show whether the proposed mechanism and prior evidence make that scenario credible.

4. Measurement limitations will hide real effects. Without proper sequencing, deduplication, and attribution, the test will report noise as signal and signal as noise. If you cannot trust the measurement, you cannot trust the result — regardless of whether the baseline was right.

Bottom Line

The baseline conversion trap is not a statistics problem. It is a modeling problem disguised as a statistics problem. The analyst is doing correct arithmetic on the wrong numbers.

The fix is three moves:

  1. Anchor in reality. Start with the number you can measure, treat it as an upper bound.
  2. Adjust for causality. Discount for the fraction of that number your test can actually influence.
  3. Work in ranges, not points. Lower, midpoint, upper. Plan on the midpoint.

That is how a reviewable plan separates measured inputs from scenario assumptions before launch.

The surprise is the dashboard. The mistake was three weeks earlier, in the planning doc.

Key Takeaways

  • Eventual recovery is not the same as testable conversion. A 90-day recovery rate includes lifecycle, retargeting, and cross-device returns that no A/B test can influence. Use it as an upper bound, not a baseline.
  • Use the smallest eligible denominator, not the biggest. Traffic is the users who can actually experience the variant — not total site visitors, not monthly actives.
  • Plan on ranges, not point estimates. Lower bound, midpoint, upper bound. Plan on the midpoint, sanity-check with the edges.
  • There is no universal realistic UX lift. Size against eligible traffic and the minimum effect worth implementing; support the scenario with local evidence.
  • An eventual recovery rate is a diagnostic, not automatically the test baseline. It can motivate return-path and friction hypotheses, but it does not identify the cause.
  • The four-question checklist: eligible users only, immediate vs. eventual, testable slice only, ranges not points. Any "no" goes back to the drafting phase.

FAQ

What is the difference between baseline conversion rate and eventual recovery rate?

Baseline conversion rate is the fraction of users who convert through the specific slice of behavior your test can influence. Eventual recovery rate is the fraction of users who eventually convert through any channel over the full measurement window. The two numbers can be very different — eventual recovery is usually much higher because it includes retargeting, lifecycle email, and cross-device returns the test cannot touch.

Why is using total site visitors as the traffic denominator wrong for an A/B test?

Because most A/B tests only run on a specific page, segment, or flow. A user who never reaches that page cannot experience the variant, so they are statistical noise. Including them inflates the apparent sample size by orders of magnitude and makes the power calculation produce runtimes that are not achievable in reality.

How do I estimate a baseline when I cannot measure in-session conversion directly?

Anchor on what you can measure, then use path, identity, source, and treatment-eligibility data to estimate the slice the test can reach. Build a sensitivity range when the path is incomplete, and show how the sample plan changes across that range. Do not apply a portable discount percentage.

What is a realistic lift to expect from a UX-focused A/B test?

There is no portable lift range for a “typical” UX change. Define the smallest effect worth implementing, calculate what the available traffic can detect under the chosen error and power targets, and compare that effect with relevant prior evidence. If the business threshold and detectable effect do not overlap, redesign or deprioritize the test.

How do I size an experiment when traffic is low?

First, confirm the traffic estimate uses the eligible denominator. If traffic is genuinely low, show the detectable effect under the approved design, consider a longer window or a higher-reach outcome, aggregate only where the causal question remains valid, or deprioritize the test. Sequential or Bayesian methods change the decision procedure; they do not manufacture information from an inadequate sample.

What should I review before trusting the dashboard's baseline and metric?

Confirm the population, denominator, outcome window, treatment reach, and decision rule against the written experiment plan. Microsoft Research catalogs related metric-interpretation pitfalls in online controlled experiments, while Google's work on long-term experiment effects explains why a short-window metric may not answer a long-term business question.

Share this article
LinkedIn (opens in new tab) X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified in behavioral economics. Led 100+ in-house experiments at NRG in 2025, with project evidence and limits documented in the case studies.