Outcome-dependent stopping is a common way to invalidate a fixed-horizon test. The size of the error inflation depends on the exact monitoring and stopping rule, so the fix is a precommitted design—not a memorized multiplier.

What You'll Learn

  • What “peeking” is and why outcome-dependent stopping changes the error rate
  • Why the inflation must be calculated or simulated for the actual rule
  • Why short, round-number test durations require more context
  • Two common valid stopping-design families: fixed-horizon and sequential
  • A Monday-morning checklist you can apply to your own tests

Quick Stats Reference

Remember these design rules:

>

- Alpha = 0.05 controls the long-run Type I error at 5% under the null for the specified valid procedure.
- Repeated looks are part of the procedure. Their timing and stopping boundary determine the actual error rate.
- Selection can inflate the reported point estimate, but there is no universal 3–5× correction.
- Duration comes from sample and seasonality, not a universal 14-day minimum.

The Core Definition

Peeking = observing a test's interim results and deciding whether to stop based on what you see. The statistical inference of a test is only valid if you commit to your stopping rule before the test launches. Peeking and stopping at the first favorable observation is a different protocol — one with a much higher false positive rate than the dashboard displays.

Why This Matters for New Analysts

If you are new to CRO, this is an important design concept to learn early. Three reasons:

It can materially increase false positives. Fixing outcome-dependent stopping removes one avoidable source of error; the remaining risk still depends on multiplicity, metric selection, instrumentation, and other parts of the protocol.

It is an operationally simple fix. Pre-committing to a sample target still requires planning, review, and reliable inputs, but it usually does not require changing the treatment itself. Track the added planning time and invalid decisions avoided in your own workflow.

A dashboard label is not a stopping rule. Some A/B testing platforms display a continuously updated "probability variant B beats A" indicator. Do not act on it until the documentation specifies the estimand, monitoring policy, and calibration for the way your team plans to look and stop.


The Mechanism in Plain Language

A two-arm fixed-horizon A/B test with a pre-specified Type I error rate makes a long-run claim under its assumptions: when the null is true, the procedure's false-positive rate is controlled at that chosen level. That calibration belongs to the full procedure, not to a dashboard number viewed in isolation.

That error-rate claim assumes the specified protocol: pre-specify a sample size, run the test to completion, and evaluate under the planned analysis.

The moment you replace that protocol with continuous monitoring — checking the dashboard at days 3, 5, 7, deciding stop-or-continue based on what you see — the inferential properties change.

Quick mental model: Each outcome-dependent look gives a noisy run another opportunity to cross the boundary. Quantify the resulting Type I error by specifying the looks and stopping rule, then use analytical boundaries or simulation.

This distinction is established in sequential statistics, but a dashboard may not surface how its displayed value changes under repeated looks. Read the methodology rather than inferring validity from the interface.


Why Short, Round-Number Durations Need Context

Published DTC case studies often report a round-number duration without the sample target, expected eligible traffic, outcome delay, seasonality, or stopping rule. Duration alone cannot tell you whether a test was valid: a short test can reach a pre-specified sample, and a long test can still be stopped opportunistically.

Ask for the design record. Was the target sample calculated before launch? Were interim looks and boundaries pre-specified? Did the runtime cover the outcome delay and material calendar cycles? Did traffic or eligibility change during the run?

Tip for new analysts: Treat an unexplained duration as missing methodology, not proof of peeking. Downgrade the claim until the sample plan and stopping rule are available; do not invent a universal number of days that separates valid from invalid tests.

The Second Effect: Inflated Lift Magnitude

The peeking problem has a second, less-discussed effect that matters more for the headline numbers in case studies: the lifts you report are systematically biased upward, even when the underlying effect is real.

Here's the intuition. Imagine the true lift is 0% — no real effect. The observed lift varies day to day around zero. Some days it's slightly positive; some days slightly negative. If you apply a peeking rule — "stop the first day the observed lift crosses 15% with apparent significance" — you will eventually get a day where noise pushes the lift above the threshold. You declare a winner. The reported lift is at least 15% by construction.

What just happened: you sampled the upper tail of the distribution of possible observations. The reported "lift" is not an estimate of the true effect. It is a measurement of how far noise traveled before you stopped.

Planning rule: Do not convert a peek-and-stop headline into a “true” effect with a fixed divisor. Treat the estimate as selection-biased and rerun under a valid protocol or use an explicitly modeled shrinkage method.

The cost in production. When the change is rolled out and measured again, a selection-biased headline can shrink because the new noise component is not conditioned on crossing the original stopping boundary. The new estimate can still vary in either direction, so compare estimates and intervals rather than promising a fixed haircut.


The Diagnostic Signal: Failed Replications

When a peek-and-stop finding is tested again under a pre-specified design, the result may shrink, disappear, or reverse. The replication rate and amount of shrinkage depend on the original selection rule, uncertainty, context, and implementation; there is no defensible universal percentage.

Audience differences, seasonality, traffic mix, implementation drift, and selection bias can all explain divergence. Inspect those mechanisms rather than assigning fixed prior odds without a relevant evidence base.

When a result fails to replicate on a comparable surface, update confidence in the original and investigate both selection and context. The failed replication is evidence; it does not identify a single cause by itself.

For new analysts, the takeaway is to read replication failures as a reason to audit the original stopping rule, uncertainty, implementation, and comparability. They are a useful diagnostic, not proof of one specific pipeline failure.


The Two Fixes

Fix 1: Fixed-Horizon Testing (start here)

A practical starting design for teams that do not need interim decisions:

  1. Specify the decision-relevant effect. What is the smallest absolute or relative effect that would change the business choice after implementation cost and risk? "Anything positive" is not a usable threshold.
  2. Calculate required sample size. Use a calculator that matches the metric and analysis. Record the baseline, MDE, allocation, variance assumptions, and chosen Type I and Type II error rates with the sessions required per arm.
  3. Convert sample to time. Use expected eligible traffic by arm, then check whether the resulting period represents the outcome delay and calendar cycles relevant to the decision.
  4. Pre-register. Document the target sample, expected end date, primary metric, MDE, and decision rule. Commit before launch.
  5. Defer evaluation until the end. Observe interim if you must, but do not act. Stop only when the pre-committed sample is reached.
Common objection: "This will slow us down." The duration difference depends on the prior invalid stopping behavior and the valid design's sample requirement. Calculate it rather than promising one extra week.

Fix 2: Sequential Testing (when you really need to monitor)

A more sophisticated alternative that allows interim analyses without inflating false positives. Two principal families:

  • Group sequential designs (e.g., O'Brien-Fleming, Pocock boundaries) — pre-specify a small number of interim checkpoints with adjusted critical values.
  • Always-valid inference (e.g., mixture-SPRT, e-values) — permit continuous monitoring with valid inference at every observation.

Trade-offs depend on the sequential design, look schedule, and alternatives being compared. Calibrated interim looks can permit earlier stopping for strong benefit, harm, or futility while preserving the specified error policy; they do not guarantee a shorter expected run for every effect.

Critical caveat: A Bayesian-looking probability label does not prove that a tool's stopping rule is valid for continuous monitoring. Require documentation of the model, decision threshold, look schedule, and calibration. If validity for your monitoring policy is unclear, use a fixed-horizon design or get statistical review.

A Practical Stopping Rule for DTC

For teams not yet adopting sequential frameworks, document these design questions before launch:

Design questionWhat to specify before launch
Decision-relevant effectMDE or non-inferiority margin tied to the business decision
Required sampleBaseline, allocation, alpha, power, and variance assumptions
SeasonalityWhich cycles materially affect the metric and must be represented
MonitoringFixed horizon or calibrated sequential looks and boundaries
SegmentsWhich cuts are confirmatory and how multiplicity will be handled

Convert the required sample to time using expected eligible traffic, then check whether the period represents relevant seasonality. The result may be shorter or longer than two weeks.

What matters more than the specific durations is that they are documented before launch and not adjusted based on interim observations.


Monday-Morning Checklist

Apply this to your team's next test, today:

  • [ ] Document the MDE before launch (specific number, written down).
  • [ ] Calculate required sample size using a sample-size calculator.
  • [ ] Convert sample to time using eligible traffic; include material outcome delays and calendar cycles.
  • [ ] Add the target end date to your team's testing template.
  • [ ] Set a calendar reminder for the target end date — not before.
  • [ ] If you must look at interim results, do not act on them.
  • [ ] At the end date, evaluate. Then ship or kill, with full reporting.

For your existing test history:

  • [ ] Pull a recent, representative set of consequential "wins."
  • [ ] Note the run time and compare it with the precommitted sample and stopping rule. Duration alone does not prove invalidity.
  • [ ] For any test without a pre-registered sample target, treat it as a candidate finding requiring replication, not as established evidence.

Quick Tips for New Testing Teams

  • The dashboard label is not the procedure. A "94% probability variant B beats A" reading is not a stopping criterion unless the tool documents calibration for your monitoring and decision rule.
  • Failed replications are signal. When the same change tested elsewhere returns flat, reduce confidence in the original and inspect selection, uncertainty, implementation, and comparability. Do not assign one cause without evidence.
  • Pre-registration makes the rule auditable. Document the test plan before launch so the analysis can be checked against the intended procedure.
  • There is no universal 14-day baseline. Size for the effect and seasonality that matter to the decision.

FAQ

What if I don't have enough traffic for the required sample?

Revisit the decision-relevant effect, metric, allocation, and opportunity. You may reduce cadence, choose a higher-traffic surface, accept a larger MDE, extend the window, or decide the question cannot be answered experimentally at present. Sequential methods can improve stopping flexibility but do not create traffic.

My tool reports "94% probability variant B beats A." Can I stop?

Only if the tool documents a stopping procedure calibrated for your monitoring pattern and the result crosses that procedure's pre-specified boundary. Read the methodology. If it is unclear, do not use the displayed percentage as permission to stop.

How should duration relate to purchase cycle?

Purchase cycle and weekly seasonality are different problems. Include enough follow-up to observe the outcome and enough calendar coverage for material traffic cycles. The required duration follows from those inputs and the sample design, not a fixed 14-day rule.

How can I tell if my historical wins were peeking-induced false positives?

Re-run a representative sample of consequential findings under a pre-specified design on a comparable surface. Compare effect estimates and intervals, then audit any divergence against the original stopping rule, implementation, audience, and calendar conditions. Your own paired history—not an uncited universal replication rate—can quantify the program's selection and transport gap.

Is there ever a justification for stopping a test early?

Yes—for benefit, harm, or futility under a calibrated sequential design, or for operational and safety events covered by the protocol. Specify the look schedule, calculation, boundary, and action before launch. Stopping a fixed-horizon test merely because the current result looks favorable is a different procedure from the one originally planned.


Build the Habit Early

For new CRO analysts, a pre-specified stopping design is a useful habit to develop early. It removes one avoidable source of analytic flexibility; its measured effect on false positives in your program depends on what other failure modes are present.

I built GrowthLayer to make pre-registered, fixed-horizon testing the default for new analysts and small teams. The platform sizes the sample for your specified MDE, carries the target horizon into the readout, and reports the result with a confidence interval rather than presenting a continuously changing probability label as a universal stopping rule.

To find experimentation roles where this discipline is the operating standard, explore open positions on Jobsolv.

Or book a consultation for a methodological audit of an existing test history, or for help establishing a disciplined testing program from scratch.

Share this article
LinkedIn (opens in new tab)X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified in behavioral economics. Led 100+ in-house experiments at NRG in 2025, with project evidence and limits documented in the case studies.