Atticus Li has led enterprise experimentation programs where engagement metrics, conversion proxies, and financial outcomes had to coexist in one decision system. This guide separates those evidence classes so a proxy is not presented as revenue.
I use a rule that makes some teams uncomfortable: before a test runs, document how its primary metric connects to a business outcome. Measure revenue directly where the sample and data support it; otherwise name the proxy, its validation evidence, and the downstream check.
This rule removes test ideas that cannot support a decision with the available traffic or instrumentation. It also prevents a click or scroll change from silently becoming a revenue claim.
The Vanity Metric Trap
The common trap is to measure the easiest event near the changed interface and stop there.
Click-through rate on a button. Time on page. Scroll depth. Video play rate. "Engagement score" — whatever that means in your organization. Bounce rate reduction. Page load improvements. Form field completion rate.
None of these are bad metrics in isolation. Some of them are genuinely useful diagnostic signals. The problem is when they become the primary success metric for your A/B test. Because when they do, you're optimizing for activity, not outcomes.
Here is an illustrative composite of a recurring failure mode. A team redesigns a landing-page hero, updates the copy, and adds social proof. The variant increases CTA click-through in the high-teens. The team celebrates and ships.
Later, someone pulls the downstream data. Signups and revenue are flat. The click movement only shifted where people clicked; it did not establish more paying customers.
That's not a win. That's a vanity metric masquerading as a business result.
Full-Funnel Accountability
The fix is conceptually simple and operationally hard: measure the whole funnel, not just the step you're optimizing.
When I evaluate a test, I don't just look at the metric closest to the change. I look at what happens downstream. Did the landing page change increase form submissions? Great. Did those form submissions convert to qualified leads at the same rate? Did those leads close at the same rate? Did those customers retain?
The following are illustrative composites of downstream failure modes, not attributable client case studies:
The landing page “win” that hurts lead quality. In this illustrative composite, a simplified form produces a low-twenties completion increase while the resulting leads close at a materially lower rate. More leads do not establish more revenue; the removed fields may have carried qualifying information.
The checkout optimization that increases completions but hurts revenue. In another composite, completed orders rise in the low-teens while average order value falls in the mid-teens and returns increase. Revenue per user declines despite the upstream movement.
The email click win that does not move subscriptions. In a third composite, email click-through rises by roughly one-quarter, but the additional clicks go to content rather than signup. The email is more engaging without establishing more subscriptions.
Every one of these would have been celebrated as a win if the team only measured the metric closest to the change. Full-funnel accountability caught the problem.
The Downstream Check
Here's the framework I use. For every test, before it launches, you have to answer these questions in the experiment brief.
What is the primary metric? This should be as close to revenue as your measurement capability allows. If you can measure revenue directly, that's your primary metric. If you can't, it should be the closest reliable proxy — completed purchases, qualified leads submitted, subscriptions activated.
What is the downstream validation metric? Even if your primary metric is a proxy, you should have a plan to check what happens further down the funnel. If you're measuring form submissions, your downstream check is lead quality and close rate. If you're measuring completed purchases, your downstream check is return rate and lifetime value.
What's the revenue-per-user impact? This is where it gets concrete. You need a documented formula that translates your primary metric into dollars. Revenue per user is the simplest version of this: take the total revenue from the test period, divide by the number of users in each variant, and compare.
Revenue per user can cut through much of the noise when the experiment has enough sample and a reliable revenue join. If it does not move detectably, the test does not support a revenue-lift claim on that evidence; that is different from proving the true effect is exactly zero.
Building a Transparent Dollar Model
Here is a simplified translation model.
Every test gets a projected revenue impact before it launches. The formula is: (projected lift in primary metric) x (conversion rate through remaining funnel steps) x (average revenue per converted user) x (monthly traffic volume) x (12 months).
That gives you an annualized revenue projection. It's an estimate. It will be wrong. But it forces the team to think about the full path from test metric to revenue before they invest a single sprint in the test.
After the test concludes, replace the planning lift with the observed test-window estimate and its uncertainty. The result is still a modeled annual impact, not actual realized revenue. Reconcile it later with recognized post-launch outcomes where measurement allows.
Over time, build a database that keeps planning scenarios, test-window estimates, and recognized outcomes in separate fields. That makes forecast error visible and supports better prioritization.
Why This Is Hard (And Why It Matters Anyway)
I'm not going to pretend this is easy. Measuring full-funnel impact requires instrumentation that most companies don't have on day one. You need to be able to track a user from the experiment through to the revenue event, which might be weeks or months later. You need data pipelines that connect your experimentation platform to your revenue data. You need agreement on what "revenue" means in your context.
Building this instrumentation can require partnerships with data engineering, finance, and IT. The scope and timeline depend on the current identity, revenue, and experimentation data—not a universal implementation schedule.
If you can't measure revenue directly today, start with the best proxy you have and build toward full-funnel measurement. Even moving from CTR to form submissions is a meaningful step. The key is to always be pushing the measurement closer to the thing that actually matters: did the company make more money?
The Culture Shift
The biggest change is not only technical; it is cultural. When you stop classifying proxy movement as business impact, the reported outcome mix may change. Tests that looked positive on CTR or engagement may be null, negative, or unresolved on the downstream metric.
This is where leadership storytelling matters. A project-specific outcome rate on a decision-relevant metric can be more useful than a higher rate produced by permissive proxy definitions, but there is no universal 24%-versus-60% comparison. Define the denominator and evidence class.
When a material test-window effect is reconciled with finance-approved post-launch outcomes, the conversation can shift from activity to business value. Do not substitute an anonymous $2M anecdote for that evidence trail.
Measure what matters. Document the path from proxy to outcome. Report the full portfolio under one denominator and keep observed, modeled, and recognized value separate.
_Need to plan and interpret an experiment without reducing the decision to a single threshold? GrowthLayer's A/B test calculator connects sample planning, data-quality checks, and the final effect._
FAQ
Are clicks and engagement always vanity metrics?
No. They can be useful diagnostics or leading indicators. The problem is treating them as business outcomes without validating the path to revenue.
What if revenue arrives long after the experiment?
Pre-specify a near-term decision metric, retain a downstream holdout or cohort where feasible, and reconcile the estimate later. Google's research on long-term experiment effects explains why the later check matters.
How do I start moving a program toward revenue evidence?
Map one priority funnel from exposure to recognized outcome, document the gaps, and improve the closest broken link first. Contact me if you need help scoping that measurement audit.