Multivariate Testing or a Confounded A/B Test: Which Did You Run?
Four bundled changes in one experiment came back inconclusive, and couldn't have told us anything either way. A confounded-test-design lesson.
Articles exploring experiment-design through the lens of behavioral science and experimentation. Practical frameworks for growth leaders who measure in revenue, not vanity metrics.
32 articles
Four bundled changes in one experiment came back inconclusive, and couldn't have told us anything either way. A confounded-test-design lesson.
See why a customer-selector pop-up can fail, how two first-party chooser tests compare, and how to test useful personalization without adding friction.
See why brochure previews may increase downloads, how to grade the evidence, and how to test lead magnet clarity without mistaking images for proof.
See what a product-color matching test really suggests, what its source omits, and how to test color congruence without relying on folklore.
Evaluate any A/B testing case study with a 12-point evidence checklist covering source, sample, metrics, stopping, SRM, limitations, and transfer.
Run a cleaner navigation A/B test with visible, collapsed, and removed treatments, precommitted metrics, guardrails, SRM checks, and decisions.
Should a landing page have navigation? Compare the public evidence, missing methods, intent conditions, guardrails, and a safer A/B test plan.
See four A/B testing examples graded by evidence quality, with missing data, limits, transferable lessons, and safer next-test plans.
Minimum detectable effect (MDE) is the most important input to A/B test design. Learn how to calculate and choose the right MDE for business impact and traffic.
A practitioner's guide to writing A/B test hypotheses — the structure that survives review, the three failure modes that produce inconclusive tests, and how…
Five statistical mistakes that derail new DTC testing teams, with practical checks for sample size, peeking, metrics, and interpretation.
How denominator and time-window mistakes distort experiment baselines, with an illustrative worked example and a framework for sizing under uncertainty.
Standard A/B tests break when users influence each other. Learn how network effects create interference and the experimental designs that handle it.
Master the 'design an A/B test' interview question with a structured framework. Learn the step-by-step approach that impresses hiring managers every time.
Sample ratio mismatch can invalidate an A/B test. Learn how to detect SRM, trace its root cause across the experiment pipeline, and decide whether to rerun.
Pre-registration locks in your experiment plan before seeing results. Learn why it prevents p-hacking, metric shopping, and post-hoc rationalization.
When A/B tests track multiple metrics, statistical complexity increases. Learn frameworks for managing metric conflicts and making sound decisions.
Guardrail metrics prevent A/B tests from causing hidden damage. Learn how to set them up, monitor them, and use them to make better ship decisions.
Your primary metric determines whether an A/B test succeeds or fails. Learn how to select metrics that are sensitive, aligned, and actionable.
Learn how to design rigorous A/B tests from hypothesis to execution. Covers experiment structure, variable isolation, and common design mistakes.
Underpowered tests waste traffic, miss real wins, and erode trust in experimentation. Learn how to diagnose the problem and fix it before it kills your program.
Statistical power determines whether your A/B test can detect real effects. Most experiments run underpowered, wasting traffic and producing misleading results.
Running A/B tests without proper sample size calculation wastes traffic and produces unreliable results. Learn the inputs, formulas, and practical trade-offs.
MDE isn't a calculator input — it's the foundation of your entire experiment design.
Explore how AI and large language models are transforming A/B test hypothesis generation by eliminating confirmation bias, surfacing non-obvious patterns in…
Learn what statistical power means for A/B testing, why 80% is the standard, and how underpowered tests lead to costly false negatives that cause you to…
Master A/B test sample size calculation including the relationship between baseline conversion rate, minimum detectable effect, and statistical power to…
Understand the difference between one-tailed and two-tailed hypothesis tests in A/B testing, when each is appropriate, and the simple conversion rule between them.
A strong hypothesis is the difference between an experiment that teaches you something and one that wastes traffic.
Learn what A/B/n testing is, how traffic splits work with three or more variants, when you need multiple variants, and the tradeoffs compared to simple A/B tests.
Anchoring bias silently distorts A/B test results by making the control variant the psychological reference point against which all alternatives are judged…
How to write experiment briefs that prevent last-minute stakeholder rewrites by building alignment into the document structure.