This article began as an analysis of a historical working export with more than 200 rows. The current public proof ledger contains 142 source records representing 139 canonical experiments. Those are not the same denominator, and the row-level export is not public enough to reproduce every rate previously reported here. I therefore treat the patterns below as project-specific observations and diagnostic hypotheses—not industry benchmarks.

When I reconciled the working export with the public ledger, I chose the smaller reproducible denominator for public claims and retained the larger export only as historical context.

TL;DR

  • Evidence classes were often mixed. Traditional significance, non-inferiority, holdout evidence, directional signals, and business overrides need different labels.
  • Device aggregation sometimes hid opposite-direction movement. That is a reason to pre-specify device analysis, not a promise that segmentation raises win rates.
  • The historical portfolio skewed toward acquisition. Whether that is wrong depends on the business, available surfaces, and marginal value of retention opportunities.
  • Work in progress mattered. A useful program view includes discovery, blocked work, inconclusive results, and post-launch checks—not only shipped wins.

How the data was collected and anonymized

The historical working export contained more than 200 rows across multiple workflow states. The current public evidence ledger exposes 142 source records representing 139 canonical experiments. Because a workflow row is not always a unique completed experiment, I do not use the larger row count as an audited experiment total.

FieldAnonymization rule
Brand namesStripped — no company, product, or sub-brand identifiers
Test IDsInternal IDs replaced with public Exp-XXX aliases
Specific datesBucketed to quarter; exact run dates removed
Sample sizes per testBucketed to wide ranges (1k-10k, 10k-100k, 100k+)
Lift percentagesReported as ranges, not exact values
Internal funnel codesGenericized (account dashboard, plan selection page, etc.)
Personnel namesStripped except author

The records are anonymized, and the article does not expose enough row-level data to reproduce the old aggregate rates. That limitation matters: the patterns can guide an audit, but they should not be used as planning priors for another program.

Finding 1: Evidence labels were inconsistent

The working export mixed several decision classes: traditional frequentist results, non-inferiority decisions, holdout-informed rollouts, directional ships, and business overrides. Because the public ledger does not expose a reproducible classification for every historical row, I no longer present a single “stat-sig rate.”

The durable finding is operational. A repository where “winner” can mean statistically significant, directionally positive, or simply shipped is not decision-grade. Store the method, stopping rule, uncertainty, and ship rationale separately so a later reader can tell what the result supports.

Finding 2: Device aggregation can hide heterogeneity

Several historical readouts showed mobile and desktop moving in different directions. That made an aggregate result less useful for the decision.

This does not establish that single-device tests outperform combined tests, nor that every program should ship by device. The defensible practice is to pre-specify device analysis when the mechanism or experience differs, maintain multiplicity discipline, and decide whether a conditional rollout is operationally justified.

Finding 3: Methodology should match the decision

The historical ledger relied heavily on standard 50/50 tests, while some decisions were better framed as non-inferiority, staged rollout, or post-launch monitoring. That is a project-specific process observation, not an industry distribution.

The cost is institutional: programs accumulate "we shipped this and the data was noisy" entries in their test repository. A year later, nobody can defend why a particular variant is in production. The methodology was wrong for the conditions, but the documentation never made that distinction.

There is no universal baseline or traffic threshold that selects the “right” regime. Choose from the estimand, decision cost, power, rollout constraints, and the harm of being wrong. Document that choice before the result is visible.

Finding 4: The portfolio skewed toward acquisition

Acquisition surfaces made up most of the historical working export, while retention received less attention. The unreconciled denominator means I do not present the old 89/8 split as an audited portfolio statistic.

The correct allocation is not a universal percentage. Compare the marginal value, feasible sample, decision urgency, and organizational ownership of acquisition and retention opportunities. The ledger should make that trade-off visible.

Finding 5: The ledger must include work in progress

The historical export tracked shipped, reverted, inconclusive, live, discovery, development, and blocked work. Because those workflow rows are not all unique completed experiments, their old counts should not be compared as an audited win distribution.

This is the institutional reality of CRO programs at scale. Most external CRO content presents a clean win → ship → measure narrative. The actual shape of the work is much messier, with substantial discovery, pre-design, legal review, and waiting-on-stakeholder states between hypothesis and shipped variant.

Implication for program leaders: do not judge health from “wins shipped” alone. Track queue age, blocked time, test quality, decision latency, and post-launch follow-through. No single velocity metric proves business value.

Finding 6: The failure-mode taxonomy

The ledger suggested a recurring set of diagnostic patterns. The public evidence does not support audited frequency labels, so use the list as an investigation checklist:

Failure modeDetection signal
Wrong-intent clicksSub-baseline click-to-conversion from specific source pages
Friction injection at destinationModal-mediated CTA has sub-baseline click-to-conv vs equivalent direct-routing
CTA cannibalizationExisting CTAs lose volume in variant; new CTA at sub-baseline conv rate
Confounded-variable trapUpstream + downstream metrics in opposite directions
Form-vs-content mismatchTest optimizes form when user research said content
Above-the-fold competitionMultiple CTAs at same intent/commitment/audience
Aggregate-mask asymmetryAggregate flat or noisy; pre-specified segments move in opposite directions
Trust-badge ambiguityTime-on-page up + FAQ attractiveness up + conversion flat
Time-on-page misreadMetric interpreted without engagement context
Underpowered stat-sig forcingDesign cannot detect the decision-relevant effect with available traffic
Copy-intent mismatchHigh CTR + low click-to-conv on intent-mismatched source pages
Visibility failureLow CTR, high click-to-conv when clicked

Treating them as distinct hypotheses can improve diagnosis. It does not guarantee a higher follow-up win rate.

This site has a dedicated article on each of these failure modes; collectively they form the practitioner's atlas of how CTA tests actually fail.

Finding 7: The gap between CRO advice and CRO reality

A historical ledger makes universal tactics look less useful. “Always use 50/50,” “sticky CTAs work,” “trust badges are free wins,” and “more tests means more wins” each omit the decision context.

The practical gap is evidence labeling. A directional ship is not a statistically significant win; an inconclusive test is not proof of no effect; a secondary metric is not a substitute primary outcome. Publish the definitions and full corpus so readers can evaluate the result.

What this means for practitioners

For CRO managers, analysts, and growth product managers, the data has a few specific implications:

ImplicationAction
Evidence class mattersStore the design, stopping rule, interval, and ship rationale with every result.
Device effects may differPre-specify device analysis when the experiences or mechanisms differ.
Lifecycle allocation deserves reviewCompare acquisition and retention opportunities using marginal value and feasible sample.
Failure modes are diagnostic hypothesesUse the taxonomy to investigate an inconclusive result without promising a fixed remedy.
External win rates are not your benchmarkDefine and audit your own denominator, decision rules, and full outcome distribution.

For founders and VPs of Growth evaluating CRO investment:

QuestionWhat to examine
Should I build in-house or hire an agency?Ownership, access to data, methodology, handoffs, cost, and whether the evidence stays with your team.
How fast should we expect wins?Do not promise a rate. Estimate how quickly the program can answer qualified decisions with the available traffic.
What's a sign the program is working?Decision quality, queue health, implementation follow-through, and recognized business outcomes—not win count alone.
How do I evaluate a CRO consultant or agency?Ask for the full ledger, evidence classes, loss and inconclusive reporting, and examples of forecast reconciliation.

Bottom line

Real CRO programs contain more than a clean sequence of wins. The historical ledger reinforces the need to choose a design for the decision, pre-specify important segments, report the whole portfolio, and separate test-window evidence from modeled business impact.

The operating disciplines I would look for are:

  1. Match methodology to test conditions (50/50, holdout-validated, non-inferiority — pick the right regime)
  2. Segment by device class as a default, not a custom-cut analysis
  3. Run pre-test screens that catch failure modes before experiment budget is committed
  4. Document methodology accurately so future readers know what shipped under what evidence
  5. Audit their test-mix balance between acquisition and retention quarterly

These disciplines are intended to improve defensibility and learning quality. They do not guarantee a particular win rate or revenue result.

This article is the hub for a practitioner's atlas of A/B testing. Each finding above is covered in depth in a dedicated spoke article — sticky CTAs, click-to-conversion ratio, copy intent, modal vs direct routing, holdout-validated shipping, device asymmetry, time-on-page interpretation, trust-badge architecture, confounded-variable recovery, form vs content variable selection, above-the-fold hierarchy, and CTA cannibalization detection. Read whichever pattern matches the test you're about to ship.

Any future quantitative update should publish a reconciled denominator and reproducible classification rules before reporting aggregate rates.

FAQ

Why mention 200+ historical rows when the public ledger is smaller?

The working export contained more than 200 workflow rows, not 200 audited completed experiments. The reproducible public ledger contains 142 source records representing 139 canonical experiments, so public rates should use that reconciled denominator.

What can readers safely learn from the article?

Use the patterns as diagnostic questions for test design and reporting. Do not treat them as universal benchmarks for win rate, device effects, or revenue.

How should a portfolio analysis be reproduced?

Publish inclusion rules, deduplication logic, outcome labels, primary-metric rules, and the analysis query. Microsoft Research's experiment-pitfall catalog and NIST's statistics handbook provide useful methodological checks. Contact me if you want help designing the ledger.

Share this article
LinkedIn (opens in new tab)X / Twitter (opens in new tab)
Atticus Li

Experimentation and growth leader. CXL-certified CRO practitioner, Mindworx-certified in behavioral economics. Led 100+ in-house experiments at NRG in 2025, with project evidence and limits documented in the case studies.