Evaluate an A/B testing case study by reconstructing the experiment before judging the headline. Find the primary source, identify the sample and allocation, verify the primary metric and duration, locate the stopping rule and SRM check, then separate what worked once from what has actually been replicated.
An A/B testing case study is a narrative report of an experiment; it is not automatically a complete experiment record.
DataForSEO estimates only about 20 US searches per month for “A/B testing case study,” but the query is commercially valuable. People searching it are choosing ideas, vendors, or methods. A trustworthy evaluation framework can prevent them from spending traffic and engineering time on a story that cannot support its headline.
Key takeaways
- Use the primary source, not a chain of roundups.
- A winning outcome does not determine evidence quality.
- Missing sample, duration, stopping rule, or SRM should remain visibly missing.
- Range-bucketed private data supports directional synthesis, not exact meta-analysis.
- Every case should end in an original test plan, not a cloned design.
Why most case studies are incomplete
Case studies are written to communicate a result. Experiment records are written to preserve a decision. Those incentives produce different documents.
A marketing case often emphasizes:
- The dramatic visual change
- A relative lift
- Confidence or significance
- A short explanation of why the winner won
A reusable experiment record also needs:
- Eligibility and assignment unit
- Raw arm sizes and outcomes
- Metric definitions
- Planned allocation
- Duration and business cycles
- Stopping rule
- SRM and instrumentation checks
- Guardrails and downstream effects
- Contradictory segments
- Decision and limitations
In my portfolio audit, records that looked complete at the headline level often became narrative-only after these fields were checked. That is not a reason to discard them. It is a reason to use them for mechanisms and hypotheses instead of pooled effect estimates.
The 12-point Primary Evidence Reconstruction
Use this framework for every case.
| # | Field | Question |
|---|---|---|
| 1 | Primary source | Is this the organization’s report, platform case, paper, or a summary? |
| 2 | Population | Who was eligible, and what traffic sources were included? |
| 3 | Assignment unit | User, session, account, device, region, or something else? |
| 4 | Sample | How many eligible units entered each arm? |
| 5 | Allocation | What split was planned and observed? |
| 6 | Treatment | What changed, and what else changed with it? |
| 7 | Primary metric | Was success defined before the result? |
| 8 | Duration | Did the test cover relevant cycles and novelty? |
| 9 | Stopping rule | Fixed sample, fixed time, or documented sequential method? |
| 10 | SRM and quality | Did assignment and instrumentation behave correctly? |
| 11 | Guardrails | What cost could the primary win hide? |
| 12 | Limitations | What cannot be inferred or transferred? |
Call this the Primary Evidence Reconstruction. A blank cell is a finding, not an invitation to guess.
Worked example: navigation removal
The VWO case report says removing navigation from a registry landing page moved registrations from 3% to 6%. The source report describes the page, traffic sources, treatment, goal, and rates.
It does not report sample size, allocation ratio, duration, stopping rule, SRM, raw counts, or a confidence interval. That earns a C — partially reported.
Calibrated conclusion: the treatment worked once in the reported setting. It suggests that global navigation can compete with a narrow registration goal. It does not prove a 100% expected lift for another landing page.
This sentence is less exciting than the headline and more useful to someone allocating traffic.
Worked example: pricing-page registration
Buttondown’s first-party pricing experiment reports registrations changing from 6.7% to 9.5%. It identifies the downstream registration metric and explains the implementation. It does not publish the sample, duration, allocation, stopping rule, or SRM.
The treatment also changed two elements: it removed a callout and repositioned the signup action. The report supports the business conclusion that the combined page performed better. It cannot isolate which element caused the change.
This is a classic separation:
- Shipping evidence: useful
- Mechanism evidence: limited
- Replication claim: unsupported
The original thesis most summaries miss is that experiment usefulness has two axes: decision value and learning resolution. A case can be high on the first and low on the second.
How to grade evidence
Use five grades:
| Grade | Definition | Safe use |
|---|---|---|
| A — Reproducible | Raw arms, metric, duration, uncertainty, stopping, and quality checks available | Quantitative synthesis when outcomes are comparable |
| B — Verified | First-party or independently reviewed evidence, publicly sanitized | Directional and mechanism synthesis |
| C — Partial | Credible report with important fields missing | Supporting case and hypothesis |
| D — Anecdotal | Claim with little methodological detail | Inspiration only |
| X — Excluded | Duplicate, contradictory, unverifiable, or improperly sourced | Do not count as evidence |
Do not award a better grade because the lift is large or significant. An inconclusive A-grade test is more reusable than a dramatic D-grade winner.
My own public-safe navigation example receives a B. The repository preserves more than 100,000 observations, three approximately equal arms, completed orders, a four-to-eight-week duration, no detected SRM, and a 10%–20% winning range. The stopping rule is missing, and exact data are intentionally not public. That supports a directional first-party contribution but not independent reanalysis.
When can you call the article a meta-analysis?
A formal meta-analysis requires comparable effect estimates and uncertainty. At minimum, you generally need arm-level sample and outcomes or an effect size with a standard error or confidence interval. You also need a defensible reason the populations, treatments, and metrics belong in one model.
Do not pool:
- A checkout order with a landing-page click
- Revenue per session with registration rate without a coherent transformation
- Multiple records from the same underlying experiment as independent tests
- Range-bucketed confidential effects as exact values
- Vendor winners while silently excluding nulls and losses
If those conditions are not met, use systematic evidence review, research synthesis, or case-study comparison. The portfolio evidence audit uses that language because heterogeneous metrics and dependent records make a pooled lift misleading.
How should public and private evidence coexist?
Use two transparent policies.
Public cases
Describe the subject generically when prominence is unnecessary, but keep the publisher citation visible. Quote only what is needed, and evaluate omissions directly.
Private portfolio cases
Anonymize the organization, remove proprietary screenshots, bucket samples and effects, and disclose that public readers cannot reproduce the estimate. Never use range-bucketing to upgrade evidence quality.
This lets the database protect commercial confidentiality without pretending the source is public. It also follows the same separation used in decisions with incomplete data.
Turn every case into an original test plan
An article adds value only if its thesis goes beyond the source. Use this sequence:
- State exactly what the source supports.
- Identify the missing evidence fields.
- Name at least one competing mechanism.
- Add a portfolio pattern, counterexample, calculation, or enterprise constraint.
- Specify where transfer is plausible and where it may fail.
- Define a new control, treatment, primary metric, guardrails, sample plan, and stopping rule.
For navigation, a stronger follow-up compares visible, collapsed, and enclosed states rather than cloning the winning screenshot. See the three-arm navigation test plan and calculate feasibility with the sample-size guide.
Use calibrated claim language
Use these terms consistently:
- Worked once: one credible experiment favored the treatment.
- Suggests: incomplete or heterogeneous evidence points in a direction.
- Supports: multiple credible sources or methods align with the mechanism.
- Replicated: comparable independent experiments reproduce the claim under specified conditions.
“Has been replicated” should be the rarest label in a public case library. Two vendor blog posts that both declare winners are not automatically replications.
Start your evidence review in GrowthLayer
Start your free GrowthLayer workspace to capture the 12 reconstruction fields, assign an evidence grade, link the primary source, and turn the case into an actionable experiment brief.
FAQ
What should an A/B testing case study include?
It should include the population, assignment, sample, allocation, treatment, primary metric, duration, stopping rule, SRM status, guardrails, result, decision, and limitations.
Can I trust a case study without sample size?
Use it as a hypothesis source, not a precise effect estimate. Without sample and uncertainty, you cannot judge how stable the reported rates are.
Is statistical confidence enough to validate a case study?
No. Confidence does not fix poor assignment, an undocumented stopping rule, multiple comparisons, instrumentation errors, weak metrics, or missing guardrails.
Can anonymized internal experiments be included?
Yes, when the author has permission and discloses the sanitization. Range-bucketed private results should support directional synthesis rather than exact public meta-analysis.
When has an A/B test result been replicated?
Use that label when comparable independent experiments reproduce the effect under defined conditions. Similar-looking winners from incomplete reports are only supporting cases.