A double-digit-looking loss on a stepper redesign turned out to be neither a loss nor a win — it was noise wearing a costume, and the only professional response was to say so.
TL;DR
- Exp-054 tested a linear vs. radial stepper design on a mobile multi-step form at a large energy retailer; the topline moved sharply negative.
- The reported lift range was -20% to -10% — a number big enough to trigger panic in most stakeholder meetings.
- The traffic behind that number was too thin to reach statistical significance, which is the whole story: a big number and a real result are two different claims.
- The mechanism is statistical power — small samples produce wide, unstable estimates that can look enormous purely by chance.
- The correct response to an underpowered result isn't to ship or kill the variant. It's to hold judgment and rerun with enough traffic to trust the number.
| Detail | Value |
|---|---|
| Citation | Exp-054 |
| Surface | Mobile multi-step form |
| Comparison | Linear stepper vs. radial stepper |
| Duration | ~6 weeks |
| Reported topline lift range | -20% to -10% |
| Statistically significant? | No |
| Verdict | Inconclusive — rerun, don't react |
The number that walks into the room
Every growth leader has sat in a meeting where a dashboard shows a topline number that looks like it's on fire, and the instinct in the room is immediate: kill it, roll it back, write the post-mortem. Exp-054 was built to produce exactly that moment. We were comparing a linear stepper against a radial stepper on a mobile multi-step form at a large energy retailer — a straightforward UX pattern comparison, testing established guidance that linear steppers outperform radial ones for mobile form completion. The hypothesis was reasonable. The published pattern research behind it was reasonable. Nobody walked into this expecting a headline.
Then the topline moved sharply negative — into the -20% to -10% range — and that's where most organizations stop thinking and start reacting. A number that size, on a conversion-adjacent metric, on a live form, reads as an emergency. The pressure to act on it immediately is real, and it isn't irrational: in most companies, a double-digit-looking drop is treated as self-evidently meaningful, and the people asking "wait, is this even a trustworthy number" are the ones who get accused of stalling.
The entire value of running experiments at this level of seniority is knowing the difference between a number that demands action and a number that demands more data. Exp-054 is a clean case of the second kind, and the discipline of checking that distinction before reacting is the actual subject of this case study — not the stepper design itself.
Why a big number can still be nothing
This is where the statistics stop being a formality and start being the whole argument. A/B testing produces an estimate of the true effect of a change, and every estimate comes with a margin of error that shrinks as sample size grows. On a well-powered experiment — enough visitors flowing through both variants for long enough — a large observed lift is trustworthy because the margin of error around it is small. On an underpowered experiment, the same-looking number can be almost entirely noise, because the margin of error is wide enough to swallow it.
This is what statistical power actually measures: the probability that an experiment will detect a real effect if one exists, given the sample size and the size of the effect you're trying to see. Low power doesn't just mean "we might miss a real effect." It also means the opposite failure mode — a result can look dramatic and still be a coin flip, because with too few observations, random variation alone can produce swings that look like a decisive story. Achieving genuine statistical significance in A/B testing requires enough traffic that the pattern you're seeing couldn't plausibly be explained by chance — and Exp-054's traffic, relative to what this specific page would need to trust a topline swing of that size, fell well below that threshold.
The founder-level intuition to build here is simple: a big number and a reliable number are not the same claim, and mistaking one for the other is how organizations make expensive decisions on data that was never solid enough to support them. This is not a subtle statistical footnote — it's the difference between evidence and a story that happens to have a number attached to it.
The discipline: check the math before you trust the story
Here is the part that separates a senior experimentation practice from an ad hoc one. When a topline number moves sharply — in either direction — the reflex in most organizations is to react to the story first and check the math later, if at all. A stakeholder sees a scary chart, asks what happened, and wants an answer in the room. The pressure to have a confident narrative ready is constant, and "it's directionally negative but we can't trust the number yet" is a much harder sentence to say out loud than "radial steppers lost."
The methodology choice that matters here — the one that requires judgment rather than a dashboard — is resisting that pressure and checking sample size and statistical significance in A/B testing before assigning any meaning to the topline at all. That means asking, before anything else: how much traffic actually ran through this comparison, relative to what this page's baseline conversion rate and typical effect sizes would require to detect a real difference reliably? In Exp-054, the traffic was small relative to the retailer's other experiments running in the same window — a small fraction of what a well-powered read on this page would need before a topline swing of this size could be treated as trustworthy rather than noisy.
This isn't caution for its own sake. It's the recognition that acting on an underpowered result is itself a decision with a cost — you can kill a variant that was actually fine, or worse, generalize a false "lesson" (radial steppers are bad on mobile) into future design decisions across other surfaces. The discipline is checking the confidence behind a number before you let the number change your roadmap.
The result: -20% to -10%, and still not significant
Here is the honest accounting. Exp-054's topline lift landed in the -20% to -10% range against the linear stepper baseline — and it was not statistically significant. That combination is the entire point of this case study: the number was large enough to alarm a stakeholder meeting, and simultaneously too thin on traffic to trust as a real effect. Those two facts are not in tension. They're exactly what an underpowered read looks like.
The correct read on Exp-054 was not "ship the radial stepper" and not "kill it and move on." It was neither. The honest conclusion was inconclusive on too little data — a result that calls for a rerun with enough traffic to actually resolve the question, not a verdict in either direction. Treating it as a confirmed loss would have meant writing off a legitimate design pattern based on a number that never earned that authority. Treating it as a false alarm and ignoring it entirely would have thrown away a directionally negative signal that might be real once properly powered. The disciplined move sits between those two impulses: hold judgment, and let more data decide.
The diagnostic catch most teams stop short of
Most teams, shown a double-digit-looking negative number on a stakeholder deck, treat it as a definitive loss and act accordingly — that's the modal behavior, not an edge case. The diagnostic catch in Exp-054 is checking statistical significance against sample size before reacting, and it's a check that gets skipped constantly, not because people don't know it exists, but because organizational pressure rewards a confident story over an honest "we don't know yet."
That pressure is worth naming directly. In most companies, saying "the data was inconclusive" in a meeting where everyone expected a clean answer costs the messenger something socially, even when it's the correct read. The teams that consistently get this right are the ones that have built the habit of running the power check as a reflex, before the topline number gets a chance to set the narrative — not the ones that are simply smarter about statistics. Exp-054 is a case where that reflex prevented a real, if unglamorous, mistake: killing a legitimate mobile UX pattern on the strength of a number that was never strong enough to justify the decision.
FAQ
How do I know if my team is making this mistake?
Ask, the next time a topline number drives a decision in a meeting, whether anyone checked the sample size and significance before the conversation started. If the answer used to justify killing or shipping something is only "look at the number," and nobody can speak to whether that number was statistically significant in the A/B testing sense, that's the mistake in progress. Teams that have internalized this discipline can answer the power question without being asked.
What sample size is "enough"?
There's no universal number — it depends on your baseline conversion rate and the size of the effect you're trying to detect, and smaller effects always require dramatically more traffic to resolve reliably. The practical test isn't a specific count; it's whether the experiment ran long enough, on enough traffic relative to that page's normal volume, that a result in either direction would hold up if you reran it. Exp-054's traffic, relative to that bar, wasn't close.
Why not just trust a number that big? Surely something moved.
Something may well have moved — that's exactly why the honest conclusion here is "inconclusive," not "nothing happened." But a large observed swing on thin traffic is exactly the situation where random variation can produce numbers that look meaningful and aren't. The size of the number doesn't tell you its reliability; the sample size behind it does.
What should we have done differently in that stakeholder meeting?
Reported the number alongside its confidence, not instead of it — "the topline moved sharply negative, and here's why we don't yet trust it enough to act" is a complete, defensible answer. The mistake isn't having an inconclusive result. It's presenting an inconclusive result as if it were a settled one, in either direction.
Does this mean radial steppers are fine for mobile forms?
No — and that's the point. Exp-054 didn't prove radial steppers are fine, and it didn't prove they're a double-digit loser either. It showed that this particular read, on this traffic, couldn't support either conclusion, which is why the right next step is a properly powered rerun rather than a permanent design decision made on an underpowered signal.
Bottom line
A scary topline number is a prompt to check your confidence in it, not a verdict. Exp-054's -20% to -10% swing looked like a clear loss and was, on the traffic available, statistically indistinguishable from noise — which meant the only defensible call was to hold judgment and rerun, not to ship or kill. That distinction, made consistently and under real organizational pressure to have a confident answer, is what separates an experimentation program that compounds good decisions from one that compounds expensive mistakes dressed up as data.
If your team is making calls off toplines without checking whether the underlying number can bear the weight you're putting on it, that's a program design problem, not a one-off judgment call — and it's exactly the kind of thing I help growth teams and founders fix. If you want a second opinion on how your experimentation program handles results like this one, get in touch.
Evidence sources and free next step
VWO's A/B test significance calculator illustrates why lift size alone cannot establish uncertainty. Work through the sample-size guide and case-study evaluation guide before making the decision. Then get started with GrowthLayer free to preserve the planned sample, stopping rule, and interval.