An Experimentation
Operating System.
More experiments. Same evidence bar. PRISM generates testable hypotheses; the operating pipeline sizes each test, ties it to a customer or revenue decision, and closes with a scale, stop, or learn call.
Advisory & program builds available
PRISM: Five Steps. Every Experiment.
Behavioral frameworks do not determine which test will win. They turn observations into falsifiable explanations, while baseline data determines whether the test can detect an effect worth acting on.
The result: every experiment is accountable to a decision, a measurement standard, and a stated evidence limit.
Start with data, not guesses. Quantitative analytics locate the drop-off. Behavioral analytics — heatmaps, session recordings, funnel data — show what users do. Behavioral frameworks then suggest a falsifiable why: trust deficit, cognitive overload, anchoring, or loss aversion. The framework is a hypothesis tool, not proof of the mechanism or a guarantee that the variant will win.
Every finding competes for test bandwidth. Rank hypotheses using eligible traffic, contribution value, detectable effect, confidence, and delivery effort. The test with the strongest decision value runs first — not merely the easiest build or the most interesting idea.
Write a hypothesis that names the proposed behavioral mechanism, expected change, and an impact range based on the baseline. Then build the smallest valid test. Use power calculations, a declared analysis method, and guardrail metrics to catch unintended side effects.
Measure results against the pre-set decision rule and impact range. Scale, stop, or learn — every experiment closes with a recommendation, uncertainty bounds, and revenue context rather than a percentage lift alone.
Ship supported changes, update the baseline, and feed the finding into the next hypothesis cycle. Re-test when audiences, offers, or channels change. At NRG, this operating discipline supported a historical scale-up from roughly 20 to 100+ annual experiments and $30M+ recorded in internal 2025 program reporting.
See PRISM applied to real tests — read the experiment case studies →
From Idea to
Scaled Investment
Once a program scales past a handful of tests, the bottleneck stops being execution and becomes decision quality: which ideas get prioritized, what standards results are held to, and whether winners actually get scaled. This is the pipeline every test moves through — the same operating model that scaled an enterprise program from 20 to 100+ experiments a year.
Idea Intake & Problem Framing
Ideas come from customer behavior, analytics anomalies, market opportunities, channel performance, and stakeholder priorities. Every idea gets reframed as a business problem and the decision it will inform — "what should we invest in?" — before it earns a slot in the pipeline.
Hypothesis Design
IF / THEN / BECAUSE format, with the BECAUSE naming a behavioral mechanism — loss aversion, choice overload, social proof, anchoring. A hypothesis that can’t name its mechanism usually can’t explain its result either.
Prioritization (RICE)
Reach, Impact, Confidence, Effort — scored transparently so the pipeline is driven by expected business value, not by whoever asks loudest. Teams can see exactly why an idea moved forward or didn’t.
Experiment Design
Primary metric, guardrail metrics, audience, minimum detectable effect, power and sample-size requirements, duration, and a pre-agreed decision rule. Then match the method to the decision: onsite changes get A/B tests, lifecycle messaging gets holdouts, brand and market-level questions get geo-lift or incrementality designs.
Feasibility & Effort Sizing
Developer effort, design lift, analytics complexity, and — in regulated industries — legal and compliance review. Cheap to kill an infeasible test here; expensive to kill it after three sprints.
Analytics Setup & Tracking Audit
Instrument the events, QA the tags, and validate the baseline numbers before anything gets built. A surprising result should trigger a tracking and sample-ratio check before it triggers a business decision.
UX/UI Design & Feedback
Wireframes and variant designs reviewed against the hypothesis — does the design actually manipulate the mechanism we named? Usability feedback loops before development, not after launch.
Development & QA
Build the smallest version that tests the hypothesis. Cross-browser, cross-device QA, and a variant-parity check so the test measures the change — not a rendering bug.
Launch & Monitoring
Guardrails watched from day one. No reacting to noisy early results unless the test was explicitly designed for sequential analysis. Repeated unplanned peeking raises false-positive risk.
Analysis & Readout
Bayesian and sequential methods, false-discovery-rate control, segment checks for heterogeneous effects. Every test closes with a verdict — win, save, or learning — in revenue terms.
Business Review & Stakeholder Comms
Executive readouts built on data storytelling: what we believed, what we found, what it’s worth, what we recommend. Stakeholders aligned on the decision criteria before the test see no surprises after it.
Decision & Handoff
Scale, iterate, stop, or hand off to the owning team — marketing, product, lifecycle, or finance. Handoff is where most experimentation programs quietly lose their value, so the owner of a winning idea is named before the test ever runs.
I've Run Programs on ICE,
PIE, and RICE.
RICE won — but not because it's the most sophisticated. Prioritization frameworks sit on a tradeoff between how complex they are and how useful they are. Every hour a team spends scoring ideas is an hour it isn't shipping experiments, and a model elaborate enough to feel "rigorous" is usually a planning tax in disguise.
RICE earns its extra letter: Reach forces the sample-size and feasibility conversation on day one. An idea with beautiful impact scores but no traffic to detect the effect dies in scoring instead of dying six weeks into a doomed test. ICE and PIE let that conversation happen too late.
Just as important: a transparent score defuses loudest-stakeholder bias. Teams can see exactly why an idea moved forward or didn't. But the score structures judgment — it never replaces it.
Fast, lightweight. Good for early-stage programs — but nothing stops a zero-traffic idea from scoring high.
Page-centric and subjective — "potential" and "importance" overlap enough that scores drift toward whoever's in the room.
Reach makes measurement feasibility a first-class input. Slightly more work to score, materially better decisions — the right point on the complexity/usefulness curve.
Measuring What a Program
Is Actually Worth
To make experimentation legible to the C-suite, every verified win runs through a financial impact model:
It's a good executive-communication tool — and taken alone, it flatters the program. A credible experimentation leader tells the CFO where the model is wrong before the CFO finds out:
Winner's curse
Tests that reach significance tend to overestimate their true lift — you selected them because they looked good. Observed lifts get shrinkage applied before anyone annualizes them.
Lift decay & novelty effects
A 46-day lift is not a 12-month lift. Novelty fades, competitors respond, audiences saturate. Naive annualization is the most common way programs inflate their own impact.
Baseline drift & seasonality
Baseline conversion rates move with seasons, pricing, and market conditions. A static baseline in the formula quietly misattributes market movement to the program.
Revenue is not margin
Revenue impact and contribution-margin impact are different numbers, and the CFO cares about the second one. Where margin data exists, the model should use it.
Interaction effects
Ten concurrent winners rarely sum. Overlapping audiences and compounding changes mean the whole is usually less than the sum of the readouts.
Long-Term Holdouts: The Program-Level Audit
One useful program-level check is a long-term holdout. It estimates cumulative differences during the observed window, but it does not isolate each source of bias, repair revenue definitions, or prove durability beyond that window.
Per-test models make the program legible quarter to quarter. Holdouts keep it honest year to year. A mature program runs both.
There Is No
One-Size-Fits-All
Everything above is a starting architecture, not a template. The right version of this system depends on team size, statistical maturity, tooling, company politics, the OKRs and KPIs each team is actually measured on, bandwidth, and who owns what. A five-person growth team and a five-brand enterprise need very different amounts of process — and installing more governance than a team can absorb kills velocity just as surely as having none.
The same failure can happen with an internal build or an outside partner: installing a process without adapting it to how the team makes decisions. The mismatch shows up as friction between teams, and the program gets blamed for what was really a fit problem. The operating system has to serve the team. Never the other way around.
Each experiment should make the next decision better. Findings can compound into institutional knowledge, but they are not permanent truths: audience, offer, and channel changes still require revalidation.
Two Ways to Use
This System
Building a program?
Audits, 90-day sprints, and advisory retainers — I build the experimentation operating system with your team, shaped to how your team actually works.
See CRO consulting → For Hiring TeamsHiring for experimentation or growth?
I lead growth experimentation functions — the operating model, the measurement standards, and the executive decision layer. See the case studies under Work, or reach out directly.
Get in touch →More Experiments. Same Evidence Bar.
PRISM is a hypothesis and decision system, not a promise that every variant wins. Historical in-house work at NRG scaled from roughly 20 to 100+ annual experiments while maintaining revenue measurement standards.
A price is not an ROI forecast. Before a larger engagement, the business case uses your revenue base, eligible traffic, contribution margin, smallest useful effect, and delivery cost. If the evidence cannot justify the work, the right recommendation may be not yet.
Strategy Engagement
$5,000–$15,000 / project
For a capable team that needs the program design, roadmap, and decision rules. Typical scope: four to eight weeks.
- Research-backed experiment portfolio
- Prioritization tied to business value
- Statistical standards and guardrails
- Reporting templates and decision rules
- Cross-functional working sessions
Fractional CRO
Starting at $15,000/mo
I lead the portfolio with your team for a three-month minimum. Your team keeps production ownership; I own program decisions, experiment integrity, and executive interpretation.
- Weekly portfolio leadership
- Experiment design and quality review
- Delivery and stakeholder alignment
- Revenue models and executive reporting
- Cadence matched to traffic and team capacity
- Direct access with no account handoff
The $30M+ NRG result is historical in-house evidence—not a consulting forecast. Your scope starts with what your own baseline and delivery capacity can support.
Ready to Stop Guessing?
Tell me about your growth challenge. No pitch decks, no sales reps — just a direct conversation about whether I can help.
Revenue Frameworks
for Growth Leaders
Every week: one experiment, one framework, one insight to make your marketing more evidence-based and your revenue more predictable.
Opens Substack to confirm · Free · Unsubscribe anytime
Read the archive
A growing archive of experiments, frameworks, and field reports from inside a Fortune 150 growth team.
Open Substack (opens in new tab)