Skip to main content
Research

Practitioner
Research.

I run experiments for a living — and I treat the practice itself as a research subject. My work sits at the intersection of experimentation methodology, causal inference, behavioral economics, and applied AI, grounded in project-specific evidence including 100+ in-house NRG experiments in 2025.

Read Better Decisions →

Free weekly field notes

Focus Areas

Research
Interests

Experimentation Methodology

How organizations should measure the true value of testing programs: win rate as a vanity metric, winner’s-curse correction, sequential testing tradeoffs, and long-term holdouts as program-level audits.

Causal Inference in Marketing

Geo-incrementality, matched-market designs, holdout experiments, and attribution beyond last-click — separating demand that marketing created from demand it merely captured.

Behavioral Economics

When loss aversion, choice architecture, anchoring, and social proof provide useful hypotheses—and what production evidence is needed before treating a mechanism as relevant.

Applied AI for Growth

AI-assisted hypothesis generation, analysis acceleration, and personalization systems — and the governance layer they need: bias detection, guardrails, and transparency standards.

Working Papers

What I'm
Writing

Practitioner research drawn from production experimentation. Papers publish first as citable write-ups on this site and Better Decisions, with PDF versions prepared for scholarly indexing.

In ProgressExperimentation · Organizational Decision-Making

Win Rate Is a Vanity Metric: Toward Decision-Quality Measures for Experimentation Programs

Experimentation programs report win rates as their headline KPI, but win rate is trivially gamed by testing safe, small changes. This paper proposes a decision-quality framework — revenue per experiment, save rate, and learning velocity — drawn from operating a 100+ test/year enterprise program.

In ProgressKnowledge Management · Experimentation

Experiment Repositories as Organizational Memory: Why Testing Programs Repeat Themselves

Experiment knowledge can disappear when ownership changes or readouts are difficult to find. This paper examines repositories as institutional-memory infrastructure and the incentives that determine whether teams contribute to them.

PlannedCausal Inference · Measurement

Long-Term Holdouts and the Systematic Overstatement of Experimentation Program Value

Winner’s curse, novelty decay, and interaction effects can make summed per-test impact differ from a program-level holdout estimate. This planned paper will examine that gap, the assumptions behind each method, and the conditions under which either estimate is decision-useful.

A student considers a path from an abstract signal toward a group of human support figuresOriginal illustration · Human support, not surveillance
Research retrospective · AI validity

The strongest AI decision was rejecting the model.

A four-person hackathon team explored passive facial analysis for student support. Further research exposed an unsupported inference, so I rejected the mechanism and reframed the surviving problem around explicit student input, consent, and human-led navigation.

Mechanism
Rejected
Human problem
Preserved
Startup thesis
On hold

Team research concept · Not built, deployed, or clinically validated

Read the research case study
Reading Notes

What I'm
Reading

Papers and technical writing I'm working through, with notes on how each one holds up against production data.

Better Decisions Newsletter

Make Growth Decisions
You Can Defend

Know what to test, when to trust the result, and what to do next. Practical decision guides for analysts, growth teams, and founders.