Skip to main content
← Experiment Evidence Library
Portfolio synthesis 142 source records · 139 canonical experiments

What 142 Experiment Records Actually Teach Us

A privacy-safe audit of 142 enterprise experiment records: the reported outcome mix, the quality problems hidden inside a win-rate calculation, and the operating metrics CRO leaders should use instead.

The finding

Roughly three in ten source records were labeled winners, but the more important finding was that outcome labels alone were not reliable enough to support a formal meta-analysis or evaluate program quality.

By Atticus Li July 20, 2026 8 min read Aggregate, anonymized evidence

TL;DR

  • The source inventory contained 142 experiment records: 41 labeled winners, 32 labeled losers, and 69 labeled inconclusive.
  • After deduplication, those records represented 139 canonical experiments.
  • The quality audit classified 80 as quantitatively usable, 37 as narrative-only, and 22 as requiring review before quantitative synthesis.
  • The reported win rate is therefore a useful description of the repository, not a reliable benchmark for other teams.
  • The majority of the portfolio did not produce a reported winner. That is where much of the reusable evidence lives.

The easy headline is that roughly three in ten records were labeled winners. The useful headline is that the label was often too blunt to explain the decision.

A team can ship an inconclusive result because the variant passed a non-inferiority threshold and improved a valuable secondary outcome. It can reject a directionally positive result because a deeper metric moved the wrong way. It can also call something a winner even when the written recommendation says not to ship. All three situations appeared in the repository.

That is why this report starts with data quality rather than pretending every row is an independent, equally reliable vote.

The portfolio before the quality audit

The raw outcome fields contained this distribution:

Reported outcomeSource recordsWhat the field says
Winner41The record was stored as a positive outcome
Loser32The record was stored as a negative outcome
Inconclusive69The record did not declare a positive or negative result
Total142Before deduplication and adjudication

Taken literally, 101 of 142 records did not report a winner. That is not evidence that the program failed. It is evidence that experimentation is a search process operating against an existing experience that has already absorbed years of design, compliance, product, and commercial decisions.

If most informed changes reliably improved customer behavior, teams would not need controlled experiments. They would ship the ideas directly.

The repository also spans different surfaces, devices, metrics, test designs, and decision rules. A checkout completion and an upstream engagement event are not interchangeable outcomes. Pooling their lift percentages would create a precise-looking number with no coherent meaning.

What changed after cleaning the evidence

The audit performed four operations before classifying the portfolio:

  1. It grouped duplicate source records into canonical experiments.
  2. It separated records with usable two-arm counts and a named primary metric from narrative-only records.
  3. It flagged known traffic-allocation problems for review.
  4. It compared the reported outcome with the written shipping recommendation and the direction of the stored metric.

The result was a three-level evidence ledger:

ClassificationCanonical experimentsHow it can be used
Quantitative-ready80Aggregate counting or deeper statistical review, subject to metric compatibility
Narrative-ready37Qualitative pattern finding, operating lessons, and hypothesis generation
Review required22Retained in the ledger but not treated as clean quantitative evidence
Total139Three duplicate rows removed

These classifications are mutually exclusive. The quality flags inside them can overlap.

For example, some records lacked the raw counts needed for re-analysis but still contained a strong description of what changed, why the team rejected it, and what the next iteration should isolate. That record is not useless. It simply answers a different kind of question.

Finding one: win rate is a lossy compression of the program

A single outcome field compresses at least four separate judgments:

  • Did the measured effect favor the variant?
  • Was the result statistically or practically credible?
  • Did the variant pass guardrail and downstream checks?
  • Did the team decide to ship, iterate, or stop?

Those judgments can disagree.

Consider a non-inferiority test. The primary metric may move slightly downward while the variant creates incremental value elsewhere. If the decline remains inside a pre-committed tolerance, the team may ship. Calling that result “inconclusive” hides a successful business decision; calling it a “winner” hides the methodology.

The reverse also happens. An upstream metric can improve while completion, revenue, or customer quality deteriorates. Calling that a winner rewards the wrong behavior.

A useful repository therefore stores result, evidence strength, and decision as separate fields.

Finding two: the non-winners are the research advantage

Public case-study libraries are shaped by survivorship bias. Positive results attract attention, strengthen sales narratives, and are easier to explain. Flat and negative results disappear.

An internal portfolio has the opposite advantage: it can show which reasonable ideas repeatedly failed to produce value.

The 101 source records not labeled winners help answer questions such as:

  • Which interventions usually produced noise?
  • Which surfaces were difficult to power?
  • Where did an engagement lift fail to travel downstream?
  • Which ideas worked on one device and failed on another?
  • Which variants bundled too many changes to produce a reusable learning?

This is negative knowledge: evidence about what not to repeat without new context. For CRO managers, it can save more traffic and engineering time than another list of “best practices.”

Finding three: repository quality changes the apparent result

The audit found three duplicate source rows. It also found records requiring review because the stored outcome conflicted with the written decision, the metric direction was ambiguous, or the experiment carried a known traffic-quality flag.

That does not make the source material bad. It shows what happens when a repository grows across tools, teams, taxonomies, and reporting conventions.

The danger appears later, when someone runs a portfolio query and assumes every field has always meant the same thing.

Before using an experiment repository for AI retrieval, benchmarking, or automated hypothesis generation, standardize these fields:

FieldMinimum definition
Result directionPositive, negative, or flat on the pre-committed primary metric
Evidence statusConclusive, directional, underpowered, invalid, or not applicable
Business decisionShip, conditional ship, iterate, stop, or no decision
Metric directionWhether an increase or decrease represents improvement
Quality statusAllocation, instrumentation, novelty, and contamination checks
Reuse statusConfirmed pattern, emerging pattern, contradiction, or isolated observation

Without these distinctions, the repository stores history but cannot reliably create intelligence.

Finding four: a useful program scorecard needs more than win rate

Win rate is still worth monitoring. Sudden changes can expose a weak research pipeline, conservative testing, or outcome reclassification. It should not be the headline KPI.

A more useful scorecard combines:

  1. Verified portfolio impact. Use conservative attribution and, where possible, program-level holdouts.
  2. Revenue or customer value per completed experiment. This rewards magnitude instead of easy wins.
  3. Decision yield. Measure how often a test resolves the decision it was designed to inform.
  4. Save rate. Track harmful or wasteful changes prevented from shipping.
  5. Learning reuse. Track how often prior evidence appears in a later brief or changes prioritization.
  6. Evidence quality. Track the share of completed tests with usable metrics, methodology, and decision notes.

The last metric is often ignored. But if the organization cannot trust or reuse the evidence, the value of an experiment decays as soon as the people who ran it leave.

How to use this report in your own program

Do not compare your win rate directly with the raw ratio in this portfolio. Instead, use the audit structure.

Export your completed experiments and ask:

  1. How many source rows map to unique experiments?
  2. Does “winner” mean statistical result, business decision, or both?
  3. Can the metric be recomputed from stored counts?
  4. Is the direction of improvement explicit?
  5. Are invalid and underpowered tests distinguishable from flat tests?
  6. Can a future analyst understand why the team shipped or rejected the variant?

Once those questions are answered, your repository can support a real cross-experiment synthesis. Before that point, calculating more decimal places creates confidence without clarity.

Methodology and limitations

This is an evidence synthesis, not a formal statistical meta-analysis.

The portfolio is concentrated in related enterprise purchase and enrollment journeys. The experiments use heterogeneous metrics and methods, and several records are not statistically independent. No pooled effect size is reported because doing so would imply comparability the data does not support.

The public analysis uses aggregate counts and anonymized patterns. It does not publish experiment images, client or product names, internal identifiers, exact test dates, per-test sample sizes, or exact per-test lifts.

The reported outcome distribution describes the source repository before adjudication. Records flagged for review remain part of the inventory so the audit does not quietly discard inconvenient evidence, but they are not given the same analytical weight as clean records.

Bottom line

The apparent lesson from 142 experiment records is that roughly three in ten were labeled winners.

The durable lesson is more demanding: an experiment repository becomes valuable only when it can preserve the difference between what the metric did, how credible the evidence was, and what the team decided.

That distinction turns a history of tests into evidence a CRO team can safely reuse.

From evidence to action

Use the pattern. Test the context.

Browse related experiments in GrowthLayer, then adapt the evidence pack to your audience, funnel, and metric hierarchy. If your team needs help building the program around it, I advise growth and experimentation leaders directly.