This article began as an analysis of a historical working export with more than 200 rows. The current public proof ledger contains 142 source records representing 139 canonical experiments. Those are not the same denominator, and the row-level export is not public enough to reproduce every rate previously reported here. I therefore treat the patterns below as project-specific observations and diagnostic hypotheses—not industry benchmarks.
When I reconciled the working export with the public ledger, I chose the smaller reproducible denominator for public claims and retained the larger export only as historical context.
TL;DR
- Evidence classes were often mixed. Traditional significance, non-inferiority, holdout evidence, directional signals, and business overrides need different labels.
- Device aggregation sometimes hid opposite-direction movement. That is a reason to pre-specify device analysis, not a promise that segmentation raises win rates.
- The historical portfolio skewed toward acquisition. Whether that is wrong depends on the business, available surfaces, and marginal value of retention opportunities.
- Work in progress mattered. A useful program view includes discovery, blocked work, inconclusive results, and post-launch checks—not only shipped wins.
How the data was collected and anonymized
The historical working export contained more than 200 rows across multiple workflow states. The current public evidence ledger exposes 142 source records representing 139 canonical experiments. Because a workflow row is not always a unique completed experiment, I do not use the larger row count as an audited experiment total.
| Field | Anonymization rule |
|---|---|
| Brand names | Stripped — no company, product, or sub-brand identifiers |
| Test IDs | Internal IDs replaced with public Exp-XXX aliases |
| Specific dates | Bucketed to quarter; exact run dates removed |
| Sample sizes per test | Bucketed to wide ranges (1k-10k, 10k-100k, 100k+) |
| Lift percentages | Reported as ranges, not exact values |
| Internal funnel codes | Genericized (account dashboard, plan selection page, etc.) |
| Personnel names | Stripped except author |
The records are anonymized, and the article does not expose enough row-level data to reproduce the old aggregate rates. That limitation matters: the patterns can guide an audit, but they should not be used as planning priors for another program.
Finding 1: Evidence labels were inconsistent
The working export mixed several decision classes: traditional frequentist results, non-inferiority decisions, holdout-informed rollouts, directional ships, and business overrides. Because the public ledger does not expose a reproducible classification for every historical row, I no longer present a single “stat-sig rate.”
The durable finding is operational. A repository where “winner” can mean statistically significant, directionally positive, or simply shipped is not decision-grade. Store the method, stopping rule, uncertainty, and ship rationale separately so a later reader can tell what the result supports.
Finding 2: Device aggregation can hide heterogeneity
Several historical readouts showed mobile and desktop moving in different directions. That made an aggregate result less useful for the decision.
This does not establish that single-device tests outperform combined tests, nor that every program should ship by device. The defensible practice is to pre-specify device analysis when the mechanism or experience differs, maintain multiplicity discipline, and decide whether a conditional rollout is operationally justified.
Finding 3: Methodology should match the decision
The historical ledger relied heavily on standard 50/50 tests, while some decisions were better framed as non-inferiority, staged rollout, or post-launch monitoring. That is a project-specific process observation, not an industry distribution.
The cost is institutional: programs accumulate "we shipped this and the data was noisy" entries in their test repository. A year later, nobody can defend why a particular variant is in production. The methodology was wrong for the conditions, but the documentation never made that distinction.
There is no universal baseline or traffic threshold that selects the “right” regime. Choose from the estimand, decision cost, power, rollout constraints, and the harm of being wrong. Document that choice before the result is visible.
Finding 4: The portfolio skewed toward acquisition
Acquisition surfaces made up most of the historical working export, while retention received less attention. The unreconciled denominator means I do not present the old 89/8 split as an audited portfolio statistic.
The correct allocation is not a universal percentage. Compare the marginal value, feasible sample, decision urgency, and organizational ownership of acquisition and retention opportunities. The ledger should make that trade-off visible.
Finding 5: The ledger must include work in progress
The historical export tracked shipped, reverted, inconclusive, live, discovery, development, and blocked work. Because those workflow rows are not all unique completed experiments, their old counts should not be compared as an audited win distribution.
This is the institutional reality of CRO programs at scale. Most external CRO content presents a clean win → ship → measure narrative. The actual shape of the work is much messier, with substantial discovery, pre-design, legal review, and waiting-on-stakeholder states between hypothesis and shipped variant.
Implication for program leaders: do not judge health from “wins shipped” alone. Track queue age, blocked time, test quality, decision latency, and post-launch follow-through. No single velocity metric proves business value.
Finding 6: The failure-mode taxonomy
The ledger suggested a recurring set of diagnostic patterns. The public evidence does not support audited frequency labels, so use the list as an investigation checklist:
| Failure mode | Detection signal |
|---|---|
| Wrong-intent clicks | Sub-baseline click-to-conversion from specific source pages |
| Friction injection at destination | Modal-mediated CTA has sub-baseline click-to-conv vs equivalent direct-routing |
| CTA cannibalization | Existing CTAs lose volume in variant; new CTA at sub-baseline conv rate |
| Confounded-variable trap | Upstream + downstream metrics in opposite directions |
| Form-vs-content mismatch | Test optimizes form when user research said content |
| Above-the-fold competition | Multiple CTAs at same intent/commitment/audience |
| Aggregate-mask asymmetry | Aggregate flat or noisy; pre-specified segments move in opposite directions |
| Trust-badge ambiguity | Time-on-page up + FAQ attractiveness up + conversion flat |
| Time-on-page misread | Metric interpreted without engagement context |
| Underpowered stat-sig forcing | Design cannot detect the decision-relevant effect with available traffic |
| Copy-intent mismatch | High CTR + low click-to-conv on intent-mismatched source pages |
| Visibility failure | Low CTR, high click-to-conv when clicked |
Treating them as distinct hypotheses can improve diagnosis. It does not guarantee a higher follow-up win rate.
This site has a dedicated article on each of these failure modes; collectively they form the practitioner's atlas of how CTA tests actually fail.
Finding 7: The gap between CRO advice and CRO reality
A historical ledger makes universal tactics look less useful. “Always use 50/50,” “sticky CTAs work,” “trust badges are free wins,” and “more tests means more wins” each omit the decision context.
The practical gap is evidence labeling. A directional ship is not a statistically significant win; an inconclusive test is not proof of no effect; a secondary metric is not a substitute primary outcome. Publish the definitions and full corpus so readers can evaluate the result.
What this means for practitioners
For CRO managers, analysts, and growth product managers, the data has a few specific implications:
| Implication | Action |
|---|---|
| Evidence class matters | Store the design, stopping rule, interval, and ship rationale with every result. |
| Device effects may differ | Pre-specify device analysis when the experiences or mechanisms differ. |
| Lifecycle allocation deserves review | Compare acquisition and retention opportunities using marginal value and feasible sample. |
| Failure modes are diagnostic hypotheses | Use the taxonomy to investigate an inconclusive result without promising a fixed remedy. |
| External win rates are not your benchmark | Define and audit your own denominator, decision rules, and full outcome distribution. |
For founders and VPs of Growth evaluating CRO investment:
| Question | What to examine |
|---|---|
| Should I build in-house or hire an agency? | Ownership, access to data, methodology, handoffs, cost, and whether the evidence stays with your team. |
| How fast should we expect wins? | Do not promise a rate. Estimate how quickly the program can answer qualified decisions with the available traffic. |
| What's a sign the program is working? | Decision quality, queue health, implementation follow-through, and recognized business outcomes—not win count alone. |
| How do I evaluate a CRO consultant or agency? | Ask for the full ledger, evidence classes, loss and inconclusive reporting, and examples of forecast reconciliation. |
Bottom line
Real CRO programs contain more than a clean sequence of wins. The historical ledger reinforces the need to choose a design for the decision, pre-specify important segments, report the whole portfolio, and separate test-window evidence from modeled business impact.
The operating disciplines I would look for are:
- Match methodology to test conditions (50/50, holdout-validated, non-inferiority — pick the right regime)
- Segment by device class as a default, not a custom-cut analysis
- Run pre-test screens that catch failure modes before experiment budget is committed
- Document methodology accurately so future readers know what shipped under what evidence
- Audit their test-mix balance between acquisition and retention quarterly
These disciplines are intended to improve defensibility and learning quality. They do not guarantee a particular win rate or revenue result.
This article is the hub for a practitioner's atlas of A/B testing. Each finding above is covered in depth in a dedicated spoke article — sticky CTAs, click-to-conversion ratio, copy intent, modal vs direct routing, holdout-validated shipping, device asymmetry, time-on-page interpretation, trust-badge architecture, confounded-variable recovery, form vs content variable selection, above-the-fold hierarchy, and CTA cannibalization detection. Read whichever pattern matches the test you're about to ship.
Any future quantitative update should publish a reconciled denominator and reproducible classification rules before reporting aggregate rates.
FAQ
Why mention 200+ historical rows when the public ledger is smaller?
The working export contained more than 200 workflow rows, not 200 audited completed experiments. The reproducible public ledger contains 142 source records representing 139 canonical experiments, so public rates should use that reconciled denominator.
What can readers safely learn from the article?
Use the patterns as diagnostic questions for test design and reporting. Do not treat them as universal benchmarks for win rate, device effects, or revenue.
How should a portfolio analysis be reproduced?
Publish inclusion rules, deduplication logic, outcome labels, primary-metric rules, and the analysis query. Microsoft Research's experiment-pitfall catalog and NIST's statistics handbook provide useful methodological checks. Contact me if you want help designing the ledger.