One A/B calculator.
Before and after launch.
Size the experiment, estimate its runtime, check data quality, and interpret the effect in one defensible fixed-horizon workflow. No account. No black-box score.
What this calculator does—and does not claim.
Conversion planning uses the standard large-sample two-proportion z approximation. Post-test conversion intervals use Newcombe's hybrid Wilson-score method for the difference. Average metrics use allocation-aware planning and Welch's unequal-variance analysis.
This is a fixed-horizon calculator. It does not imitate proprietary sequential or Bayesian
engines, and it does not turn 1 − p into a “confidence” score. If you peeked repeatedly
or changed the stopping rule, the displayed p-value no longer has its nominal fixed-horizon
interpretation.
Sparse conversion data can make large-sample methods unreliable. When expected cells are very small, use an exact method or a statistician rather than forcing a decision from this interface.
Planning a high-stakes test or untangling a disputed result? Work with Atticus on experiment design and CRO strategy.
Read the result without overstating it.
Is this a sample size calculator or a significance calculator?
Both. Plan mode calculates the sample and duration before launch. Analyze mode checks sample ratio mismatch, estimates the effect and its interval, reports the p-value, and compares the result with your planned sample and practical threshold.
What does a p-value mean?
A p-value is the probability of seeing data at least this extreme if the null hypothesis and model assumptions were true. It is not the probability that the result happened by chance, and one minus the p-value is not confidence that the variant is better.
Should I use a one-sided or two-sided test?
Use two-sided by default because a variant can cause meaningful harm as well as improvement. Use one-sided only when the direction was declared before the experiment and a result in the opposite direction would never support the decision.
What is sample ratio mismatch (SRM)?
SRM checks whether observed traffic matches the allocation you planned. This tool uses p < 0.001 as a strong warning threshold. A very small SRM p-value can indicate assignment, instrumentation, or data-loss problems. Investigate that before interpreting a winner.
Why might another calculator give a different answer?
Calculators can use different test families, tail choices, continuity corrections, allocation rules, sequential methods, or Bayesian models. This tool is explicit: fixed-horizon large-sample tests, two-sided by default, a pooled two-proportion z-test for conversions, and Welch inference for averages.