Regression to the Mean in A/B Tests: A Practical Guide
A plain-English guide to why selected extreme A/B test results often shrink, and why no universal discount can recover the true effect.
Know what to test, when to trust the result, and what to do next. Practical guides for analysts, growth teams, and founders.
A plain-English guide to why selected extreme A/B test results often shrink, and why no universal discount can recover the true effect.
How outcome-dependent stopping changes A/B test error rates, why there is no universal peeking multiplier, and how to precommit a valid stopping rule.
The standard 3-tier pricing-page playbook — anchor + decoy + "Most Popular" badge — works under specific conditions. Here are the ones that break.
Dark patterns optimize for the next quarter. Bright patterns optimize for the next decade. The 12 alternatives, and why your A/B test will misjudge them.
Princeton found dark patterns on 11% of shopping sites. The FTC, EU, and California now regulate them. The 12-pattern taxonomy and what each costs you.
Krug's usability principles still hold — but the cost of misapplying them now compounds in places dashboards cannot see. The complete Krug-2026 reading list.
A practical framework for deciding which above-the-fold CTAs serve distinct intent and which create competition worth testing.
Use click-to-conversion ratios to find wrong-intent clicks, destination friction, and cannibalization hidden by aggregate CTA reports.
Use placement-level conversion and source intent to test whether a directional CTA result is genuinely additive or cannibalized.
Separate CTA visibility from intent match by comparing click-through, destination completion, and source-page context together.
Lessons from a historical enterprise testing ledger, with explicit limits on the denominator and no universal win-rate or device-performance benchmark.
How pre-specified mobile and desktop analysis can reveal heterogeneous CTA effects without claiming that device segmentation universally raises win rates.
Interpret time-on-page with conversion, scroll depth, and interaction data to separate faster decisions from abandonment or confusion.
Use behavioral diagnostics and explanatory content to test whether a benefit badge reduces uncertainty or adds friction.
Most teams discover A/B testing and immediately want to run more experiments. The logic seems sound: more tests mean more data, more data means better
The most expensive misreading in A/B testing is treating 'not statistically significant' as 'no difference.' It actually means 'we didn't collect enough…
There's a pattern worth noticing every time a new category of "automated optimization" software launches: the marketing promises to replace the hard
Every few months, a new platform promises to automate conversion optimization. The pitch is always the same: remove the human bottleneck, run more tests
After 100+ experiments per year, fixed-sample A/B testing's opportunity cost became impossible to ignore.
A/B test repositories don't fail because the schema is wrong. They fail because nobody can find what they need fast enough.
The point of documenting experiments isn't to record what happened. It's to make the next similar hypothesis sharper.
Repeated failed experiments aren't a sign of ambition — they're a sign your team isn't reading its own archive.
The best A/B testing platform isn't a single tool — it's the one that fits your team's scale, statistical needs, integration stack, and cost curve.
A centralized A/B testing database is only as useful as the fraction of experiments you can fully reconstruct.
Know what to test, when to trust the result, and what to do next. Practical decision guides for analysts, growth teams, and founders. Free. Weekly.
Opens Substack to confirm your subscription · Free · Unsubscribe anytime