Atticus Li has led enterprise experimentation programs where multiple teams needed one defensible record of what a test did and did not show. This guide focuses on the governance that keeps interpretation from changing after results arrive.
I've been in rooms where three different stakeholders presented the same A/B test result and somehow reached three entirely different conclusions. Not because anyone was lying. Because each person was selecting the slice of data that supported their existing position.
This is the dirty secret of enterprise experimentation: the data doesn't speak for itself. People speak for the data. And when you let that happen without guardrails, you end up with an experimentation program that produces ammunition instead of insights.
How Results Actually Get Spun
The examples below are illustrative composites of recurring governance failures, not attributable company readouts.
Selective metric reporting. A test has one primary metric and several secondary metrics. The primary metric is flat, but one secondary metric moves in the low-teens. Guess which number makes it into the executive readout? The stakeholder who championed the change suddenly becomes very interested in an engagement metric they had not prioritized before.
Timeframe manipulation. A multiweek test includes a temporary spike during a seasonal promotion. Someone pulls that interval in isolation and presents it as “the result,” or truncates the planned period before a later regression. The slice may be accurate; the conclusion is manufactured.
Attribution gymnastics. Marketing claims the low-single-digit movement from a checkout test. Product claims the same movement because it redesigned the interface. Both teams then report the same effect upward as separate wins, double-counting one observed result.
Reframing the hypothesis. The original hypothesis was about increasing purchases. The test didn't move purchases. But someone rewrites history: "Actually, we were testing whether the new messaging resonated with users. And the click-through rate on the CTA increased, so it clearly resonated." That was never the hypothesis. But try proving that six weeks later when the original brief has been buried in a Confluence page nobody reads.
The confidence-threshold stretch. A result misses the pre-agreed decision threshold. But the stakeholder presents it as "directionally positive" and recommends shipping. A directional observation can inform a follow-up; it cannot silently replace the rule agreed before launch.
Why This Happens Everywhere
The root cause isn't bad people. It's misaligned incentives layered on top of organizational pressure.
Product managers need to show their roadmap items delivered results. Marketing leaders need to justify campaign spend. Design teams need to prove their redesigns weren't just aesthetic exercises. Engineering leads need to show the sprint was worth the investment.
Every single one of these people has a legitimate reason to want the test to succeed. And when the raw data doesn't cooperate, the temptation to find the angle that supports the narrative is overwhelming. Not because they're dishonest — because their performance reviews, their budgets, and their credibility are on the line.
I've watched senior leaders who I deeply respect subtly reframe test results to protect their teams. It's human nature. Which is exactly why you can't rely on human nature to keep the data clean.
The Center of Excellence Must Be the Single Source of Truth
This is where an experimentation Center of Excellence earns its existence. Not as a consulting team. Not as a support function. As the authoritative source of what happened.
Here's what that means in practice. One team owns the final readout. One team determines whether the test met its primary success metric. One team writes the conclusion. Everyone else can provide context, suggest follow-ups, and challenge the methodology. But they don't get to write an alternative version of reality.
In the programs I lead, every experiment has a standardized report that includes the primary metric result, the pre-registered hypothesis, the sample size, the uncertainty, and the business recommendation. That report is the canonical record. A leadership readout can add context, but it cannot change the recorded result.
This isn't about control. It's about credibility. An experimentation program lives and dies on trust. The moment stakeholders stop trusting the results — because they've seen the same data presented five different ways — you're done. Nobody funds a program they don't trust.
The Standardized Reporting Fix
Here's the specific framework I use to prevent spinning.
One primary metric, declared before the test launches. Not after. Before. In writing. In the experiment brief. If you want to track secondary metrics, great. But the test is evaluated on the primary metric. Period.
Pre-registered hypotheses. You write down what you expect to happen and why before you see any data. This makes it much harder to retroactively claim the test was about something else. Pre-registration isn't just for academia. It's the single most effective defense against organizational spin.
Standardized readout template. Every test gets the same format. Primary metric result. Statistical significance. Confidence interval. Sample size. Duration. Business recommendation. Secondary observations. No creative reframing. No narrative embellishment. The template forces clarity.
Win/loss classification by the experimentation team. Not by the stakeholder. The team that ran the test calls it a win, a loss, or inconclusive. That classification goes into the program database and doesn't change. If a stakeholder disagrees, they can challenge the methodology or request a follow-up test. They don't get to reclassify the result.
Transparent communication of non-wins. Normalize null, negative, invalid, and inconclusive outcomes instead of hiding them. A portfolio that only reports wins has a selection problem. Present the full outcome mix under the same counting rules so positive results have context.
The Story Matters — But It Has to Be Honest
I want to be clear about something. I'm not saying narrative doesn't matter. It absolutely does. Especially early in a program when you're fighting for budget and headcount, the way you frame your results determines whether the program survives.
But there's a massive difference between smart framing and spin.
Smart framing says: "This test didn't move the primary metric, but it gave us enough evidence to stop this idea before committing two more development sprints." That is a decision record. It should not be converted into a cost-savings claim unless the counterfactual spend was documented.
Spin says: "While the primary metric was inconclusive, we saw strong engagement signals that suggest the variant resonated with users." That's not honest. That's someone trying to avoid the word "lost."
The difference is simple. Smart framing presents the actual result and adds business context. Spin presents a different result than what actually happened.
What Happens When You Don't Fix This
I've seen programs collapse under the weight of their own spin. It follows a predictable pattern. Results get spun. Leadership notices the inconsistencies. Trust erodes. Funding questions start. The program gets restructured, downsized, or killed.
The irony is that the spin was meant to protect the program. Stakeholders thought that presenting wins would secure the budget. Instead, the inconsistency between what was presented and what actually shipped destroyed the program's reputation.
Experimentation programs earn trust by reporting the full portfolio and applying the same rules to positive, negative, null, invalid, and inconclusive results. Named company totals belong on the corresponding case study with their evidence limits—not in a generic governance rule.
Build the single source of truth. Enforce the reporting standard. Tell the story honestly. That's how you build a program that lasts.
_If you're building an experimentation program and want one workflow for sample planning, SRM, effect intervals, and the final decision, use GrowthLayer's unified A/B test calculator. It is built for operators, not threshold theater._
FAQ
What belongs in a single source of truth?
Store the original hypothesis, primary metric, guardrails, eligibility, sample plan, stopping rule, result, interpretation, owner, and final decision in one versioned record.
How should exploratory findings be presented?
Label them as exploratory, explain how many cuts were inspected, and require confirmation before using them as causal evidence.
What prevents stakeholders from changing the rule after results?
Pre-commitment and visible governance. Microsoft Research catalogs metric-interpretation pitfalls that a written decision protocol can help expose.