A useful navigation A/B test should compare visibility, accessibility, and removal as separate mechanisms. A simple keep-versus-delete test can reveal a better page, but it cannot tell you whether users needed fewer visible choices, fewer destinations, or merely a clearer primary action.
A navigation A/B test randomly compares how different access to site destinations changes a preselected user or business outcome.
Literal search demand for “navigation A/B test” is small, but the test pattern appears throughout higher-volume questions about A/B testing examples and checkout page design. This article is the implementation spoke: it converts public evidence and portfolio experience into a brief a team can run.
Key takeaways
- Test visible, collapsed, and enclosed navigation when traffic permits.
- Preserve recovery, support, and accessibility in every treatment.
- Use a deep primary metric and precommitted guardrails.
- Check SRM before interpreting any lift.
- Report whether the test worked once, suggests a mechanism, or contributes to replication.
Why the usual two-arm test is ambiguous
The familiar hypothesis is: “If we remove navigation, conversion will increase because users will face fewer distractions.” The treatment changes at least three things:
- The number of visible choices
- The number of available destinations
- The user’s perceived freedom to leave or investigate
If the variation wins, any of those mechanisms could explain the outcome. If it loses, a missing trust or recovery path may have outweighed reduced distraction.
The public VWO case reports registrations moving from 3% to 6% after main navigation was removed from a registry landing page. The primary vendor case report does not disclose sample, allocation, duration, stopping rule, or SRM. It worked once in the reported setting and justifies a hypothesis—not a universal treatment.
My portfolio supplies a useful counterexample. In my experience running a three-arm mobile test with more than 100,000 observations and approximately equal allocation, consolidating two navigation systems produced a 10%–20% improvement in completed orders. The test ran four to eight weeks and had no detected SRM; its stopping rule is missing from the available record. Simplification won without elimination.
The original thesis is therefore:
Navigation salience and navigation availability are different causal variables. Test them separately when the business decision depends on why the treatment works.
There is a practical tradeoff. An additional arm costs traffic and makes the analysis more complex. If leadership only needs to know whether the current header should become enclosed, a two-arm test may be enough. Use the third arm when the answer will shape a reusable design system, when removing access carries material trust risk, or when the team expects to apply the learning across multiple funnel surfaces. Experimental resolution should match the value of the decision. Document that decision explicitly.
The three treatments
!Visible, collapsed, and removed navigation states
_Original conceptual illustration using a fictional interface. It does not reproduce any public or private experiment design._
Control: visible navigation
Use the actual current experience. Do not clean up unrelated elements only in the control or let implementation differences introduce speed changes.
Variation one: collapsed navigation
Keep the same destinations behind an obvious “Menu” control. This tests reduced salience while preserving access. Track menu opens and the destinations selected as diagnostics.
Variation two: enclosed navigation
Remove the global destinations but retain task-specific progress, editing, help, legal information, accessibility, and a deliberate escape hatch. This tests stronger focus without making the page operationally unsafe.
If your question is specifically about complete removal, add it only as a fourth arm when traffic supports the extra comparison. Most teams learn more from a carefully enclosed treatment than from a blank header.
Write the hypothesis before designing
Use this copy-ready hypothesis:
If we reduce the salience of global navigation on this high-intent page while preserving task support and recovery, then qualified completion will increase because fewer users will switch into unrelated objectives before finishing.
Notice what this does not claim. It does not say navigation is bad, every extra choice creates cognitive overload, or the treatment will reproduce a vendor case.
Before launch, document the behavioral evidence:
- Session recordings or journey data showing detours
- Navigation destinations used from the target page
- User research identifying missing information
- Public cases supporting the mechanism
- Contradictory evidence where navigation aids product finding
Baymard’s checkout-flow research supports enclosed checkout design and removal of unrelated distractions. Its homepage and category-navigation research also shows the countercase: users abandon when they cannot find products. Context decides which mechanism dominates.
Choose a primary metric that can settle the decision
Use the deepest outcome the test can power reliably.
| Surface | Preferred primary metric |
|---|---|
| Campaign landing page | Qualified registration or accepted lead |
| Pricing page | Qualified signup, purchase, or revenue per visitor |
| Checkout | Completed order per eligible entrant |
| Logged-in upgrade | Paid upgrade or revenue per eligible session |
| Product discovery | Successful product progression or purchase |
Navigation clicks, CTA clicks, and step progression explain behavior. They should not declare the winner when the business outcome disagrees.
The diagnostic catch is denominator drift. If the control metric uses everyone exposed to the page but the variation metric uses only people who interacted with the new component, the arms no longer answer the same question.
Precommit the guardrails
Use guardrails that represent the cost of excessive focus:
- Revenue and average order value
- Product or plan mix
- Support contacts
- Errors and failed payments
- Backtracking and editing
- Cancellation, refund, or early churn
- Accessibility success
- Page performance
- Return visits when research is deferred
Set action thresholds before looking. “No obvious problem” is not a threshold.
The existing guardrail metrics guide covers how to separate monitoring metrics from decision-changing guardrails.
Calculate sample size and comparisons
Three arms create three pairwise comparisons. Decide which comparison answers the primary question:
- Collapsed versus visible tests salience.
- Enclosed versus visible tests strong focus.
- Enclosed versus collapsed tests availability beyond salience.
Preselect one primary comparison or adjust for multiple comparisons. Do not wait for the dashboard and promote whichever pair looks best.
Use the sample-size calculator guide with the baseline rate, smallest effect worth shipping, power, and significance threshold. If the test would run through major campaigns or seasonal shifts, reduce the number of arms or select a larger minimum detectable effect.
Reconstruct the evidence record
The final report should preserve:
| Field | Required record |
|---|---|
| Sample | Eligible observations and analysis unit per arm |
| Allocation | Planned and observed split |
| Metric | Exact numerator, denominator, and direction of improvement |
| Duration | Start, end, and meaningful business cycles covered |
| Stopping rule | Fixed sample or named sequential method |
| SRM | Test, threshold, and result |
| Limitations | Novelty, contamination, missing data, and generalizability |
| Decision | Ship, conditional ship, iterate, stop, or no decision |
This record lets a future team distinguish result, credibility, and decision—the same distinction in the experiment portfolio audit.
Interpret the result with calibrated language
- Worked once: one valid test favored the treatment.
- Suggests: partial evidence points toward a mechanism.
- Supports: multiple methods or contexts align.
- Replicated: comparable independent tests reproduce the direction under defined conditions.
A win in your checkout works once for your checkout. Adding it to comparable, well-documented cases may support a broader pattern. Calling it replicated requires more than finding another screenshot with the same winner.
Use the diagnostic checklist before declaring the result.
Create the brief in GrowthLayer
Try GrowthLayer free to capture the three treatments, primary comparison, metrics, guardrails, sample plan, stopping rule, evidence grade, and follow-up decision.
FAQ
Should I test navigation removal with two or three arms?
Use three when you need to separate reduced visibility from reduced availability and have sufficient traffic. Use two when one comparison clearly settles the business decision.
What is the best metric for a navigation A/B test?
Use the deepest qualified outcome the page can power, such as completed orders, accepted leads, or paid upgrades. Use navigation engagement only for diagnosis.
What is sample-ratio mismatch?
SRM occurs when observed assignment counts differ unexpectedly from the planned split. It can indicate assignment, eligibility, or tracking problems and should be checked before interpreting results.
Does a winning navigation test prove the mechanism?
Not necessarily. A treatment can support a shipping decision while changing too many things to isolate why it worked.