Skip to main content
← GlossaryExperimentation Strategy

Holdout Testing

A method for measuring the cumulative impact of all shipped experiments by withholding changes from a small percentage of users.

What Is Holdout Testing?

A persistent holdout can estimate the aggregate incremental impact of shipped changes when assignment, exposure, interference, and measurement remain valid. Holdout size and duration should come from power, operational, and user-equity considerations rather than a fixed percentage.

Individual A/B tests estimate specific changes, while a persistent holdout can estimate a defined bundle. Interaction, population, exposure, and persistence differences mean the aggregate should be measured rather than assumed equal to the sum of individual estimates.

Also Known As

  • Marketing: Global holdback, campaign holdout
  • Sales: Control group, untreated cohort
  • Growth: Growth holdout, program holdback
  • Product: Feature holdback, legacy experience group
  • Engineering: Control bucket, persistent control
  • Data: Cumulative impact measurement, global control

How It Works

Illustrative scenario: a company compares users in a persistent holdout with users eligible for a documented set of shipped changes. The result may differ from the sum of individual estimates because populations, exposure, interactions, and persistence differ. Report the holdout estimate with its interval and design limitations.

Best Practices

  • Size the holdout from the baseline, detectable effect, power, allocation, and planned runtime.
  • Run holdouts in quarterly or half-yearly cycles — long enough to accumulate multiple shipped winners, short enough to release the holdout users to the optimized experience.
  • Track the same primary and guardrail metrics as your individual experiments.
  • Document which tests shipped during the holdout window so you can diagnose disappointing results.
  • Rotate holdout cohorts across cycles to avoid persistent inequity for specific users.

Common Mistakes

  • Contaminating the holdout by accidentally exposing holdout users to new features through marketing emails or in-app announcements.
  • Expecting the sum of individual lifts — teams get defensive when holdouts show smaller aggregate impact, but that's the point of running them.
  • Running holdouts too short to detect significance, producing unreliable program-level conclusions.

Industry Context

SaaS/B2B: Holdouts are especially valuable for measuring the cumulative impact of activation and retention experiments where individual test lifts are small but compound. Low traffic makes percentage holdouts statistically painful, so many B2B teams use account-level rather than user-level holdouts.

Ecommerce/DTC: High traffic may support a smaller allocation, but only a design-specific power calculation can establish adequacy. Do not stack individual checkout lifts as though they were automatically additive.

Lead gen: Holdouts work well for measuring the cumulative impact of landing page optimizations, especially when combined with downstream revenue attribution.

The Behavioral Science Connection

Holdout testing is a structural defense against the illusion of control — the cognitive bias where people overestimate their influence over outcomes. Teams running successful experiments naturally attribute lifts to their work; holdouts force a reality check by showing what would have happened anyway. This honest accounting is culturally expensive but strategically essential.

Key Takeaway

Holdout testing answers the executive question "is our experimentation program actually moving the business?" with evidence no individual test can provide.