Skip to main content
Free evidence-led course · for marketing, product, and growth teams

Behavioral Science for Marketing

Diagnose why a behavior is not happening, choose an intervention with a defensible mechanism, recognize when it may fail, and test commercial and customer outcomes together.

Modules
10
Audit steps
9
Research sources
49
Atticus Li, instructor for the behavioral science marketing course
InstructorAtticus Li · Applied experimentation and behavioral science
A repair, not a repaint

From bias tactics to decision strategy.

This course follows a slide-by-slide audit of the original 52-page class. It keeps the practical Behavioral Audit, corrects blended mechanisms, removes unsupported stories and metrics, and adds current research on generalizability, durability, personalization, and AI-mediated choice.

  • Every empirical claim links to a source and names its boundary.
  • Claims are labeled established mechanism, context-dependent effect, emerging evidence, or testable hypothesis.
  • Every intervention includes customer, commercial, and ethical guardrails.
  • The advanced field chapter audits GrowthLayer archive records behind prior case-study claims, including negative, inconclusive, and integrity-limited results.
Replace bias bingo with disciplined inference. 01

Evidence before effects

Treat each principle as a mechanism-level prior, then earn the right to generalize through diagnosis and measurement.

What you will be able to do

  • Separate a theory, a finding, and a marketing hypothesis.
  • Match claim strength to evidence strength.
  • Name the population, setting, outcome, and time horizon before transfer.

Evidence and boundaries

Behavioral economics integrates psychological evidence into economic analysis of judgment and decision-making.Evidence: established mechanism

Boundary: This does not mean people are uniformly irrational or that every deviation is a mistake.

A comprehensive portfolio of 126 government RCTs produced smaller average take-up effects than the selected published literature used for comparison.Evidence: context-dependent effect

Boundary: This is evidence about two nudge units, not a universal correction factor for marketing.

Choice-architecture effects can be ineffective or counterproductive in some conditions, and current knowledge does not reliably predict every transfer.Evidence: context-dependent effect

Boundary: A useful prior still needs a local counterfactual and outcome measure.

Source-backed example: From anecdote to testable claim

Setting:A launch team sees one countdown campaign outperform its predecessor and credits urgency.

What it shows:The comparison confounds audience, offer, season, creative, and timing. Scarcity is one candidate mechanism, not the finding.

Testable application:Evidence: testable hypothesisRewrite it as: for this audience and offer, a truthful deadline may increase completed purchases versus the same page without a deadline.

Failure mode and guardrails before lift

Failure or backfire
A famous finding is converted into a lift forecast, sending the team toward an untested tactic and away from the actual constraint.
Customer guardrail
Track comprehension, access, welfare, and subgroup outcomes so a business result cannot conceal customer harm or exclusion.
Commercial guardrail
Predefine the primary outcome, uncertainty, unit economics, and scale decision before reading the experiment result.
  • Never convert an anecdote into a causal estimate.
  • Report absolute outcomes and uncertainty, not only relative lift.
  • Keep observed evidence separate from the next test hypothesis.
Classic demonstration → first-party durable outcome · transfer is a hypothesis
Use the strongest relevant evidence—not the most memorable story.
Normative benchmark ≠ observed decision

Map the decision gap before naming a bias.

01 · Benchmark

What would the stated objective imply?

State the customer objective, constraints, information set, and the rule a normative model uses as the comparison point.

02 · Observation

What happens in the actual journey?

Measure choices, completion, downstream outcomes, timing, and variation by journey context or relevant segment.

03 · Competing diagnoses

A difference identifies a question—not a bias. Specify competing explanations and the evidence that would distinguish them.

  1. 01

    First-party durable outcome

    A credible local counterfactual on the target population and business or customer outcome, measured after the intervention ends.

    Watch: Inspect uncertainty, implementation, guardrails, attrition, spillovers, and whether the outcome—not only behavior—persisted.
  2. 02

    First-party causal outcome

    A credible randomized or quasi-experimental estimate on the target population, behavior, and decision-relevant outcome.

    Watch: A short-run effect can decay, displace another outcome, or fail when implementation changes.
  3. 03

    Scaled trials or synthesis

    Multi-site field evidence, a high-quality meta-analysis, or a large coordinated trial with relevant outcomes.

    Watch: Average effects can conceal contexts, segments, nulls, harms, and implementation differences.
  4. 04

    Relevant field evidence

    A well-designed study in a comparable population, decision, channel, and incentive environment.

    Watch: Transfer remains a hypothesis when any material dimension changes.
  5. 05

    Direct replication

    A close repeat tests whether a reported pattern survives a new sample, team, time, or geography.

    Watch: Replication strengthens a pattern but does not make a different marketing application equivalent.
  6. 06

    Classic demonstration

    A controlled laboratory or classroom contrast makes a candidate mechanism visible under simplified conditions.

    Watch: Useful for teaching and treatment generation—not for projecting local lift, revenue, or welfare.
  7. 07

    Mechanism hypothesis

    A reasoned prediction grounded in journey evidence, practitioner experience, or theory.

    Watch: Label it before testing; an anecdote, pattern name, or confident story is not validation.
Evidence in · decision out

Full experiment decision scorecard

Pre-register the decision inputs, then record what the experiment can and cannot support.

01Outcome
Pre-register

Name one primary outcome, its unit and window, and the smallest change that would alter the business decision.

Read after the test

Report the estimate on that outcome and whether the result is practically useful—not only whether a threshold was crossed.

02Uncertainty
Pre-register

Specify the analysis, sample plan, data-quality checks, and which heterogeneous effects are decision-relevant.

Read after the test

Report the uncertainty interval, data-integrity findings, plausible alternatives, and limits on transfer.

03Guardrails
Pre-register

Define safety, autonomy, accessibility, quality, and downstream measures with stop or rollback thresholds.

Read after the test

Check every guardrail, including distribution across users; a primary-outcome gain does not erase a guardrail failure.

04Decision rule
Pre-register

Precommit what evidence leads to scale, revise, rerun, stop, or rollback—and who owns that call.

Read after the test

Apply the rule to the complete evidence packet, record exceptions, and state the next decision the result enables.

Move down the ladder to generate ideas; move up it before forecasting, scaling, or claiming an effect.

Decision lab

Claim ladder

Take one strong claim from your marketing materials and label it observation, mechanism, causal estimate, or transfer hypothesis.

Deliverable:A four-line rewrite with setting, audience, outcome, and uncertainty.

Knowledge check

What is the strongest defensible conclusion after one campaign beats the previous campaign?

  1. The named bias caused the lift
  2. The new campaign was associated with a better result and deserves a controlled test
  3. The result will transfer to every segment
Reveal the answer and reasoning

Answer: The new campaign was associated with a better result and deserves a controlled test

Without a credible counterfactual, the mechanism and causal effect remain uncertain.

The Behavioral Audit is the operating system. 02

Diagnose before you design

Start with a precise behavior and its context, diagnose constraints, then choose an intervention and a counterfactual.

What you will be able to do

  • Define one observable target behavior.
  • Map decision moments and competing explanations.
  • Write a falsifiable intervention hypothesis with guardrails.

Evidence and boundaries

COM-B characterizes behavior as arising through interactions among capability, opportunity, and motivation.Evidence: established mechanism

Boundary: The model organizes diagnosis; it does not identify the causal bottleneck without evidence.

Hands-on application assistance increased aid applications and college outcomes in one randomized financial-aid study, while information alone did not produce the same pattern.Evidence: context-dependent effect

Boundary: The study combines a specific population, process, and assistance model; it is not proof that all service help converts.

The nine-step Behavioral Audit is a practitioner synthesis for translating diagnosis into an ethical experiment.Evidence: testable hypothesis

Boundary: Its value should be judged by decision quality and downstream experiments, not treated as a validated scientific scale.

Source-backed example: Financial-aid help versus more information

Setting:Families received either information or information plus direct help completing an application in a randomized study.

What it shows:The stronger intervention changed practical capability and process friction, not merely message wording.

Testable application:Evidence: testable hypothesisBefore rewriting a form, test whether users lack understanding, documents, time, confidence, or a workable next action.

Failure mode and guardrails before lift

Failure or backfire
The team labels a bias before diagnosing the journey, so an interface treatment misses a structural, capability, or access constraint.
Customer guardrail
Include accessibility, effort, error recovery, and the downstream customer outcome in the audit—not completion alone.
Commercial guardrail
Require a credible comparison, one decision-relevant primary metric, and explicit stop or revise criteria before launch.
  • Diagnose structural constraints before adding persuasion.
  • Do not use a behavioral label as a substitute for customer evidence.
  • Include a no-change or simpler alternative in the test design.
Working canvas · diagnose before design
The Behavioral Audit canvas

Private working area. These fields have no save or submit action and are not sent by this course. Copy your notes before you leave or reload the page.

  1. 01

    What observable action should change, by whom, where, and by when?

  2. 02

    What do behavioral data, interviews, support logs, and accessibility findings show?

  3. 03

    Where do attention, understanding, action, and recovery break down?

  4. 04

    What outcome is the customer hiring this experience for, and how does the business benefit?

  5. 05

    Which physical or psychological capability, physical or social opportunity, and reflective or automatic motivation may constrain behavior?

  6. 06

    Which frictions obstruct the customer, and which protect consent, security, comprehension, or fit?

  7. 07

    What mechanism could change behavior, and in which conditions should it fail or reverse?

  8. 08

    What changes, what stays equivalent, and what counterfactual identifies impact?

  9. 09

    What result is practically valuable, durable, safe, and worth operational cost?

Complete the diagnosis before selecting an intervention.

Decision lab

Run the nine boxes

Choose one stalled SaaS activation behavior and complete every Behavioral Audit step without naming a bias until the mechanism step.

Deliverable:A one-page audit ending in intervention, counterfactual, primary outcome, guardrails, and stop rule.

Knowledge check

Which diagnosis is actionable and testable?

  1. Customers are irrational
  2. Users abandon after identity verification because recovery instructions and document requirements are unclear
  3. People dislike forms
Reveal the answer and reasoning

Answer: Users abandon after identity verification because recovery instructions and document requirements are unclear

It identifies a moment, behavior, and plausible capability or friction constraint that can be investigated.

Remove sludge; preserve protective friction. 03

COM-B and the friction ledger

Map capability, opportunity, and motivation, then classify each friction by whose outcome it protects or obstructs.

What you will be able to do

  • Use the full COM-B taxonomy.
  • Distinguish helpful friction from sludge.
  • Map friction to customer and business outcomes.

Evidence and boundaries

COM-B divides capability into physical and psychological, opportunity into physical and social, and motivation into reflective and automatic processes.Evidence: established mechanism

Boundary: The categories overlap in practice and should organize inquiry rather than force a single label.

Making sales tax salient on grocery shelf labels changed demand in a field experiment.Evidence: context-dependent effect

Boundary: A tax display intervention does not establish that every price disclosure has the same magnitude or direction.

Friction is not inherently harmful: confirmation, comparison, consent, and security steps can protect comprehension and autonomy.Evidence: context-dependent effect

Boundary: Whether a step is protective or obstructive depends on its purpose, burden, placement, and distribution across users.

Source-backed example: A price the buyer can actually compare

Setting:A grocery field experiment displayed tax-inclusive prices before checkout rather than leaving tax less salient until purchase.

What it shows:The interface changed the physical opportunity to evaluate total cost; it did not change product quality or tax rate.

Testable application:Evidence: testable hypothesisInventory every fee, requirement, delay, and recovery path at the decision moment, then ask which ones users can anticipate.

Failure mode and guardrails before lift

Failure or backfire
Removing protective steps increases mistakes or fraud, while adding seller-serving friction suppresses refusal, correction, or exit.
Customer guardrail
Measure time, errors, accessibility, comprehension, reversals, and the effort required to refuse, correct, or leave.
Commercial guardrail
Balance completion against fraud, support cost, rework, refunds, and durable activation rather than optimizing clicks in isolation.
  • Measure errors, reversals, support contacts, and regret alongside completion.
  • Do not optimize friction away from consent or security.
  • Audit who bears the burden, especially in recovery and cancellation.
Diagnostic model · capability + opportunity + motivation
What must be true for the target behavior?
Scientific diagnosis · COM-B

Diagnose capability, opportunity, and motivation as conditions for the behavior before choosing a tactic.

Capability
Can the person understand and perform the action?
Opportunity
Do access, time, tools, and social conditions allow it?
Motivation
Does it feel valuable, safe, salient, and worth doing now?
Target behaviorActivate a SaaS workspace by completing the first data import
Practitioner overlay · friction / fuel

Translate the diagnosis into journey conditions. This applied lens is not a fourth COM-B construct or an effect claim.

Friction
A condition that may make the intended action harder or less likely in this journey. Some friction protects comprehension, consent, security, or fit.
Fuel
A condition that may make the intended action more possible, understandable, or worth doing. Map each candidate back to C, O, or M and test it locally.

Use COM-B to diagnose what must change; it does not prescribe one tactic for every context.

Decision lab

Friction ledger

List every action, wait, decision, information request, and surprise between intent and completion.

Deliverable:A table with COM-B diagnosis, customer cost, protective value, owner, evidence, and testable change.

Knowledge check

Which change is most likely helpful friction?

  1. Hiding cancellation behind repeated menus
  2. Showing a clear irreversible-action confirmation with an undo window
  3. Preselecting data sharing without explanation
Reveal the answer and reasoning

Answer: Showing a clear irreversible-action confirmation with an undo window

A proportionate confirmation can prevent costly errors while keeping the choice understandable and reversible.

A richer model than a loss slogan. 04

Risk, reference points, and framing

Use prospect theory to ask how outcomes are encoded relative to a reference point under risk, without turning parameters into universal copy rules.

What you will be able to do

  • Explain reference dependence, diminishing sensitivity, and decision weighting.
  • Separate risky-choice framing from generic negative wording.
  • Treat loss asymmetry as contextual, not fixed.

Evidence and boundaries

Prospect theory models risky choice using reference-dependent value, diminishing sensitivity, and nonlinear weighting of probabilities.Evidence: established mechanism

Boundary: The original model concerns decisions under risk; not every gain-versus-loss sentence is a prospect-theory test.

A 19-country replication reproduced most original choice patterns but also found attenuation and heterogeneity.Evidence: context-dependent effect

Boundary: The replication supports core patterns, not one universal loss-to-gain ratio.

A certain-gain versus probabilistic-gain contrast tests the reflection or certainty pattern, not loss aversion by itself.Evidence: established mechanism

Boundary: Loss aversion requires comparison around a reference point; probability and certainty also change the choice.

Source-backed example: What the replication actually supports

Setting:Participants across 19 countries completed structured monetary choices adapted from the original prospect-theory items.

What it shows:Most theoretical contrasts replicated, while country and person-level variation remained meaningful.

Testable application:Evidence: testable hypothesisFor a renewal decision, first establish the customer reference point and uncertainty; then test gain and loss frames with identical facts.

Failure mode and guardrails before lift

Failure or backfire
A loss or certainty frame changes comprehension, perceived risk, or trust instead of improving a fact-equivalent decision.
Customer guardrail
Keep outcomes and probabilities identical, disclose material uncertainty, and measure comprehension, perceived pressure, and fit.
Commercial guardrail
Read choice together with retention, refunds, support demand, and trust; a short-term selection shift is not sufficient.
  • Never exaggerate risk or invent a loss.
  • Predefine the reference point rather than inferring it after results.
  • Measure comprehension and trust as well as choice.
Prospect theory · reference dependence and risk
Risk changes with the frame and the probability.

01 Value around a reference point

Schematic prospect-theory value functionA curved line passes through a reference point. Outcomes to the left are framed as losses and outcomes to the right as gains. Both sides flatten farther from the point, showing diminishing sensitivity without specifying a fixed gain-to-loss ratio.Loss frameGain framePerceived valueReference pointDiminishing sensitivityDiminishing sensitivity

02 Probability weighting

Schematic probability-weighting functionA dotted diagonal represents equal objective probability and decision weight. A curved line is above it at low probabilities and below it across much of the moderate-to-high range before meeting it at certainty. This is not a calibrated estimate.Objective probabilityDecision weightLower probabilitiesoften overweightedModerate / higheroften underweightedEqual weighting

03 Fourfold risk pattern

Combine the outcome frame with the probability range before forming a hypothesis.

Higher probability

Gain

Common pattern: risk aversionA sure gain may be preferred to a gamble with a larger possible gain.

Higher probability

Loss

Common pattern: risk seekingA gamble may be preferred to accepting a sure loss.

Lower probability

Gain

Common pattern: risk seekingA small chance of a large gain can support lottery-like choices.

Lower probability

Loss

Common pattern: risk aversionProtection can appeal when it removes a small chance of a large loss.
Classroom probe

Choose first. Then reveal the arithmetic.

There is no scored answer. This compares two equal-expected-value choices; it does not predict what you or a customer will choose elsewhere.

Certain gainWhich gain would you choose?
Certain lossWhich loss would you choose?
Reveal expected values
Gain frame

Option A: 100% × $500 = $500

Option B: (50% × $1,000) + (50% × $0) = $500

Loss frame

Option A: 100% × −$500 = −$500

Option B: (50% × −$1,000) + (50% × $0) = −$500

Equal expected value does not make the options psychologically identical. Treat any contrast you notice as a prompt for diagnosis—not a measured effect or a marketing forecast.

Diagnose
What expectation, status quo, comparison, or goal sets the reference point?
Test
Hold the offer and probabilities constant; vary the frame without hiding information.

These are common prospect-theory patterns, not forecasts. Framing, stakes, experience, ambiguity, and the way probabilities are learned can change a decision.

Decision lab

Reference-point map

For one choice, list the current state, expected state, promised state, and competitor state. Identify which may serve as the reference point.

Deliverable:Two fact-equivalent frames plus a prediction, manipulation check, and trust guardrail.

Knowledge check

What must stay constant in a clean framing test?

  1. Only the button color
  2. The underlying outcomes and probabilities
  3. The emotional intensity
Reveal the answer and reasoning

Answer: The underlying outcomes and probabilities

A framing comparison changes presentation while preserving the substantive outcomes and probabilities.

Design a decision, not a trap. 05

Defaults, choice, and autonomy

Use defaults and assortment design as hypotheses about effort, endorsement, and uncertainty while preserving meaningful choice.

What you will be able to do

  • Explain effort, endorsement, and endowment or reference-point default mechanisms.
  • Diagnose when assortment complexity may create overload.
  • Separate choice uptake from user outcome.

Evidence and boundaries

A meta-analysis found that default effects varied substantially and were partly associated with effort or ease, endorsement, and endowment or reference-point mechanisms.Evidence: context-dependent effect

Boundary: Some studies found null or negative effects, and changing a choice does not guarantee a better outcome.

Automatic enrollment strongly changed retirement-plan participation and contribution patterns in one employer setting.Evidence: context-dependent effect

Boundary: The welfare implications depend on suitability of the default and the path from enrollment to long-term outcomes.

Choice overload is conditional rather than a rule that fewer options always perform better.Evidence: context-dependent effect

Boundary: Complexity, task difficulty, preference uncertainty, and decision goal moderate the effect.

Source-backed example: Default changes choice, then what?

Setting:Automatic enrollment increased participation in an employer retirement plan and clustered contributions at the preset rate and fund.

What it shows:The default reduced action requirements and may have signaled a recommendation, while also shaping the selected outcome.

Testable application:Evidence: testable hypothesisFor onboarding, evaluate whether the preset is safe, representative, visible, editable, and good for the user if left unchanged.

Failure mode and guardrails before lift

Failure or backfire
A default or reduced assortment increases uptake while steering people into poor-fit, hard-to-reverse, or insufficiently considered choices.
Customer guardrail
Keep the default visible, suitable, editable, and reversible, with enough comparison support for preference formation.
Commercial guardrail
Measure activation quality, retention, support, downgrade, and reversal instead of treating initial option share as success.
  • Make defaults visible and easy to change.
  • Evaluate user outcomes after the choice, not uptake alone.
  • Do not hide a dominated or unsuitable option inside complexity.
Choice architecture · defaults and helpful friction
Friction is not automatically bad—and ease is not automatically good.

Consent

Default state
No preselection
Decision friction
One informed choice
Guardrail
Comprehension and reversibility

Deletion

Default state
No destructive action
Decision friction
Confirm and offer undo
Guardrail
Error prevention

Onboarding

Default state
Safe, editable setup
Decision friction
Review consequential choices
Guardrail
Fit after activation
Before shipping
  1. Can people notice the choice?
  2. Can they understand the consequence?
  3. Can they reverse it without disproportionate effort?
Friction ledger

Name who the extra effort protects.

Potentially helpful
  • Comprehension before commitment
  • Security and identity checks
  • Protection from irreversible error
  • Enough information to compare
Potentially harmful
  • Hidden or hard-to-find exit
  • Repeated requests after refusal
  • Surprise fees late in the flow
  • Unequal effort serving only the seller
Supportive choice test

Compare the paired paths, not just the entry path.

Join
Leave
Accept
Refuse
Enable
Disable or correct

Symmetry means proportionate visibility, comprehension, and effort—not necessarily an identical number of clicks.

A default changes what happens without action. Evaluate ease, comprehension, autonomy, and error costs together.

Decision lab

Choice architecture inventory

Document the default, option set, order, labels, comparison aid, exit, and recovery path; contrast a repeat shopper with high preference certainty and a first-time shopper with low preference certainty.

Deliverable:A redesign that reduces task difficulty where needed while keeping alternatives and consequences visible for both shopper states.

Knowledge check

When should a team reduce an option set?

  1. Whenever there are more than three choices
  2. When diagnosis shows complexity or uncertainty is obstructing the target decision
  3. Whenever the premium option needs more sales
Reveal the answer and reasoning

Answer: When diagnosis shows complexity or uncertainty is obstructing the target decision

Assortment size interacts with the task and decision-maker; there is no universal optimal count.

Reference groups, direction, and truth matter. 06

Social influence without social-proof theater

Design norm information around a relevant comparison and a desired direction, then watch for boomerang effects.

What you will be able to do

  • Separate descriptive and injunctive norms.
  • Choose and justify a reference group.
  • Test for positive, null, and counterproductive movement.

Evidence and boundaries

Descriptive norm feedback reduced high household energy use but increased use among some already-efficient households; an injunctive cue countered that boomerang pattern.Evidence: context-dependent effect

Boundary: This was a household energy field experiment, not a license to add approval symbols to every metric.

Beliefs about what others approve can help explain variation in norm-based energy interventions.Evidence: context-dependent effect

Boundary: The relevant group and second-order beliefs need measurement; broad popularity counts may be irrelevant.

Large-scale energy-program evidence finds average conservation effects alongside important heterogeneity.Evidence: context-dependent effect

Boundary: An average program effect does not predict the response of a marketing segment or individual.

Source-backed example: The boomerang hidden in an average

Setting:Households were shown how their energy use compared with nearby homes.

What it shows:A descriptive comparison could motivate high users while signaling permission for low users to consume more; an approval cue helped preserve the desired direction.

Testable application:Evidence: testable hypothesisFor usage, giving, or completion norms, analyze customers above and below the reference separately and predefine unwanted movement.

Failure mode and guardrails before lift

Failure or backfire
A descriptive benchmark licenses regression among people already beyond it, or an irrelevant comparison group reduces credibility.
Customer guardrail
Use a relevant, current, auditable reference group and inspect movement on both sides of the benchmark.
Commercial guardrail
Evaluate segment-level behavior, outcome quality, retention, and trust rather than reporting only an aggregate response.
  • Use real, current, auditable counts.
  • Do not manufacture consensus, ratings, activity, or scarcity.
  • Inspect subgroup movement and downstream quality, not only the mean.
Social influence · descriptive and injunctive norms
A comparison can inform—and it can boomerang.
Reference groupSimilar customers in the same situation

Your use is compared with a relevant, current group.

Informational influence
Other people provide evidence about what may be true, useful, or safe.
Normative influence
Approval, belonging, expectations, or sanctions may change the social payoff.
Descriptive norm
What a relevant group is reported to do.
Injunctive norm
What a relevant group is reported to approve or expect.
Behavior above the reference

The comparison may prompt correction.

Confirm the audience sees the group as relevant and credible.

Behavior already better than the reference

The same comparison may license regression.

Possible correction: Recognize already-desired behavior instead of licensing regression.

Proof credibility audit

Move from a cue to inspectable evidence.

  1. Cue onlyAn unlabeled count, badge, logo, or quote.
  2. AttributedA named source, date, population, and outcome.
  3. RelevantA comparable group facing a comparable decision.
  4. VerifiableThe method, denominator, and limitations can be inspected.

A descriptive norm can move different audiences in opposite directions. Test the reference group and the full message.

Decision lab

Norm-message pre-mortem

Draft one truthful norm message, then identify the reference group, data window, desired direction, and segment that could move the wrong way.

Deliverable:A message, evidence note, subgroup analysis plan, and kill criterion.

Knowledge check

Why can “the average customer uses four features” backfire?

  1. Averages are always false
  2. Customers already above the average may see it as permission to do less
  3. Only percentages can influence behavior
Reveal the answer and reasoning

Answer: Customers already above the average may see it as permission to do less

Descriptive information can pull behavior toward the norm from either side of the comparison.

A constraint is evidence, not decoration. 07

Scarcity, urgency, and credibility

Use scarcity only when a real limit changes the decision, and make the reason and consequence legible.

What you will be able to do

  • Separate quantity scarcity from time scarcity.
  • Identify persuasion-knowledge and trust risks.
  • Design an equivalent-offer scarcity test.

Evidence and boundaries

A marketing meta-analysis found that scarcity tactics can increase purchase intentions, with effects moderated by scarcity and product conditions.Evidence: context-dependent effect

Boundary: Purchase intention is not purchase, retention, margin, or customer welfare.

Online time-scarcity promotions did not consistently outperform equivalent controls in a program of meta-analytic and experimental research.Evidence: context-dependent effect

Boundary: Credible externally grounded reasons sometimes improved responses; results still stopped short of a universal online advantage.

A truthful deadline is best treated as a testable information intervention, not as a guaranteed persuasion device.Evidence: testable hypothesis

Boundary: The hypothesis fails if buyers ignore it, distrust it, rush into poor-fit purchases, or defer until the next promotion.

Source-backed example: A reason for the clock

Setting:Online promotions with time limits were compared with controls and with limits justified by externally grounded events.

What it shows:The online setting can activate persuasion knowledge; a credible reason may matter more than visual intensity.

Testable application:Evidence: testable hypothesisState what ends, when, why, and what remains available. Test against the same offer without time pressure.

Failure mode and guardrails before lift

Failure or backfire
Pressure activates suspicion, delays the decision, or accelerates a poor-fit purchase that later becomes regret, cancellation, or distrust.
Customer guardrail
Show the real constraint, source, expiry behavior, consequence, and remaining alternatives without inventing total loss.
Commercial guardrail
Measure net completed value after cancellations, refunds, support, and trust—not clicks or checkout starts alone.
  • No resettable clocks or invented inventory.
  • Keep the underlying offer equivalent across test cells.
  • Track cancellations, regret, refunds, and trust after purchase.
Scarcity audit · information before intensity
A credible constraint leaves an operational trail.
01 · Constraint

[Verified capacity, inventory, or deadline]

Quantity, time, capacity, eligibility, or another specific limit.

02 · Truth source

[Named system or operational owner]

A current record and a person accountable for accuracy.

03 · Expiry behavior

[Exact time and what changes afterward]

The interface, offer, and saved work behave as described.

04 · Customer consequence

[What remains available to the customer]

The message explains what ends without inventing total loss.

Reject: a timer, stock count, or deadline that resets, cannot be sourced, or leaves the real consequence ambiguous.

Mechanism map

Separate the constraint from the customer’s interpretation.

Observable source · Supply

A verifiable limit in inventory, capacity, access, or production.

Observable source · Demand

Current demand changes remaining availability; report the denominator and update rule.

Observable source · Time

A real deadline changes price, eligibility, fulfillment, or another stated condition.

Interpretation · Inference

People may treat limited availability as information about popularity, quality, or fit.

Interpretation · Suspicion

Pressure, vague provenance, or a resetting limit may instead reduce credibility.

Measure the interpretation. Do not assume that a real constraint creates urgency or trust.

Illustrative audit—not a live scarcity claim. Replace every bracketed field with verifiable operational facts.

Decision lab

Constraint proof

Choose one capacity, inventory, or deadline claim and document its operational source and customer consequence.

Deliverable:A fact-equivalent control, truth owner, expiry behavior, trust measure, and post-purchase quality guardrail.

Knowledge check

Which scarcity message is most defensible?

  1. Hurry while supplies last, with no inventory signal
  2. Applications close Friday because reviewers start Saturday; saved drafts remain available
  3. A resettable ten-minute timer
Reveal the answer and reasoning

Answer: Applications close Friday because reviewers start Saturday; saved drafts remain available

The constraint, timing, reason, and consequence are specific and verifiable.

Make value comparable and total cost visible. 08

Pricing context, decoys, and free offers

Treat anchors, option context, and zero price as conditional mechanisms while preserving transparent comparison.

What you will be able to do

  • Distinguish anchor, compromise, and attraction effects.
  • Interpret field magnitudes without folklore.
  • Include monetary and nonmonetary costs in offer design.

Evidence and boundaries

Anchoring estimates can hide opposing responses across people and experimental conditions.Evidence: context-dependent effect

Boundary: An arbitrary high number is not a dependable or ethically sufficient pricing strategy.

An archival analysis of 3.6 million grocery purchases detected a modest, history-dependent attraction effect, while a field experiment across 146,851 flight searches found no general effect from a recommended decoy.Evidence: context-dependent effect

Boundary: Different products, choice sets, outcomes, and study designs prevent either result from becoming a universal pricing rule.

Zero-price experiments found a discontinuous response to free options, but free samples show heterogeneous acceleration, expansion, and cannibalization effects.Evidence: context-dependent effect

Boundary: A zero sticker price can coexist with shipping, time, data, switching, or opportunity costs.

Partitioned and drip prices can change evaluation and search by changing total-price visibility and when fees appear.Evidence: context-dependent effect

Boundary: Effects vary by presentation, surcharge, category, and regulation; transparent all-in comparison is the safety baseline.

Source-backed example: Observed shift, experimental null

Setting:Researchers found a modest association across millions of wine purchases, while a randomized decoy test across more than 140,000 flight searches found no general attraction effect.

What it shows:The contrast makes implementation, setting, study design, and customer history part of the mechanism—not footnotes after the result.

Testable application:Evidence: testable hypothesisValidate the dominance structure, then test whether a third option improves comprehension and fit rather than assuming it will move target-plan share.

Failure mode and guardrails before lift

Failure or backfire
The option set or free offer creates confusion, hidden cost, cannibalization, or low-fit selection rather than better comparison.
Customer guardrail
Show mandatory total and nonmonetary costs before commitment, with truthful attributes and meaningful alternatives.
Commercial guardrail
Evaluate contribution margin, cost to serve, expansion, retention, refunds, and support—not target-plan share alone.
  • Show mandatory total cost before commitment.
  • Do not use an option designed only to confuse or misdirect.
  • Measure plan fit, downgrade, refund, and support outcomes.
Context effects · mechanisms are not interchangeable
Map the choice set before explaining the choice.
Choice options plotted by cost and benefitEach option is positioned on cost and benefit axes. Circles identify primary options; a dashed square identifies the proposed decoy. This configuration isconsistent with asymmetric dominance on the two plotted dimensions.More monetary cost →More benefit →Alternative: alternativeAlternativeTarget: targetTargetIllustrative decoy: decoyIllustrative decoy
  • AlternativealternativeLower cost · Core outcome
  • TargettargetHigher cost · Expanded outcome
  • Illustrative decoydecoyHighest cost · Lower than target

Mechanism map

Start with the diagnostic condition; do not infer a mechanism from layout alone.

Starting valueAnchor
An initial number becomes an input to a later estimate. It need not be a credible prior price.
Comparison baselineReference price
An expected, remembered, or displayed price becomes the baseline for judging the current price.
PositionCompromise
An option becomes intermediate on meaningful attributes, not merely centered on screen.
AttentionSalience
A feature becomes easier to notice or compare. Attention is not dominance.
RelationAsymmetrically dominated decoy
The target dominates the decoy; the competitor does not. Impact still requires a test.

Decoy configuration check

For this diagram only: lower cost and higher benefit are preferred.

Result · Valid asymmetric-dominance structure
  • PassTarget dominates decoy. No worse on either dimension; better on at least one.
  • PassCompetitor does not dominate decoy. The alternative and decoy retain a trade-off.
  • PassPrimary options retain a trade-off. Neither primary option beats the other on both dimensions.
Invalid configurations
  • Trade-off option: it beats the target on one dimension.
  • Universally dominated: the competitor also dominates it.
  • Trivial set: the target already dominates the competitor.

Schematic choice set. The plotted logic uses lower cost and higher benefit; real decisions may include omitted attributes. Any change in choice remains an empirical question.

Zero price, sampling, and price visibility
Put every cost in the same decision frame.
Decision receiptBefore commitment
Advertised priceShown at selection
$72
Mandatory service feeMust also be visible now
$8
Time, data, or switching costNonmonetary but material
Name it
Comparable total
$80 + nonmonetary costs
01 · Zero price

Which costs remain?

Include shipping, time, data, switching, attention, and opportunity cost.

02 · Sampling

What changed later?

Separate trial, accelerated purchase, category expansion, and cannibalization.

03 · Total price

When was it knowable?

Test comprehension, comparison, and trust when mandatory fees appear before the choice.

Free-to-paid economics

Choose the model by the learning and risk it creates.

Sample
Customer job
Experience a bounded unit before buying more.
Business risk
Cannibalization or attracting use without learning.
Measure next
Incremental category purchase and repeat behavior.
Time-limited trial
Customer job
Test the full workflow long enough to judge fit.
Business risk
Deadline-driven activation without durable value.
Measure next
Activation quality, informed conversion, and retention.
Freemium
Customer job
Use a durable core before needs expand.
Business risk
Service cost without a credible expansion path.
Measure next
Value realization, upgrade trigger, and cost to serve.
Paid pilot
Customer job
Reduce implementation risk with a scoped commitment.
Business risk
Custom work that does not generalize or expand.
Measure next
Decision quality, implementation cost, and expansion.

This is a strategy-selection map, not evidence that one model wins across contexts.

Illustrative receipt. The numbers are not observed performance; the lesson is total-cost visibility.

Decision lab

Total-cost option map

Map each option on price, mandatory fees, time, data, reversibility, and customer outcome. Mark every anchor and dominated relationship.

Deliverable:A transparent comparison plus hypotheses for choice, comprehension, margin, and regret.

Knowledge check

What does the recent large purchase analysis justify?

  1. Always include a decoy plan
  2. Context effects can appear in real choices but may be modest and person-dependent
  3. Place the target second from the right
Reveal the answer and reasoning

Answer: Context effects can appear in real choices but may be modest and person-dependent

The evidence supports detectable conditional context effects, not a universal merchandising recipe.

Acquisition is the first time horizon, not the last. 09

Motivation over time and durable effects

Design progress and commitment around a real goal, then measure whether behavior and outcomes survive the intervention.

What you will be able to do

  • Connect present bias to a specific timing conflict.
  • Design honest progress and reversible commitment.
  • Measure persistence and mechanism after treatment ends.

Evidence and boundaries

Loyalty-program studies found effort accelerated as people approached a reward, and an endowed head start increased completion when requirements were equivalent.Evidence: context-dependent effect

Boundary: Clear progress toward one reward is not evidence that every progress bar improves durable outcomes.

A 2026 meta-analysis found present-bias estimates differed between money, effort, and other nonmonetary rewards.Evidence: context-dependent effect

Boundary: A population average cannot select a message, incentive, or reminder schedule for a specific journey.

A review found demand for commitment devices and their effectiveness vary, while commitment can also impose costs when people misforecast or need flexibility.Evidence: context-dependent effect

Boundary: The device, population, time horizon, exit terms, and alternative supports determine whether commitment helps.

Pooled evidence from 38 home-energy field experiments found that part of the reduction persisted after treatment and was consistent with technology adoption.Evidence: context-dependent effect

Boundary: This mechanism is context-specific and does not imply that all behavioral interventions form habits.

Large multi-arm field studies can compare many plausible messages under one shared outcome and population.Evidence: context-dependent effect

Boundary: A winner ranking in vaccination reminders is not a reusable ranking for commercial messages.

Source-backed example: Persistence through changed infrastructure

Setting:Home-energy studies followed outcomes beyond active norm feedback and after resident moves.

What it shows:Part of the effect remained with the home, consistent with durable technology adoption rather than a message-only habit story.

Testable application:Evidence: testable hypothesisFor onboarding, distinguish temporary prompts from lasting changes in setup, skill, or environment and measure both after prompts stop.

Failure mode and guardrails before lift

Failure or backfire
A progress cue or commitment increases exposure-period activity while creating pressure, lock-in, low-quality completion, or no durable outcome.
Customer guardrail
Show only earned progress, make commitment informed and voluntary, and preserve a proportionate exit or revision path.
Commercial guardrail
Measure durable activation, retention, incremental cost, and post-treatment value rather than prompt-period clicks.
  • Do not imply progress that has not occurred.
  • Make commitment optional, comprehensible, and easy to exit.
  • Measure post-treatment outcomes and unintended lock-in.
Present bias, progress, and persistence
Measure beyond the moment the intervention is visible.
  1. T0
    During treatment

    Immediate target behavior

    Did exposure change action?
  2. T1
    After treatment

    Behavior without the prompt

    Did the effect persist?
  3. T2
    Later outcome

    Customer and business result

    Was persistence useful and safe?
Competing explanation

Structural adoption: Track setup, skill, environment, or technology changes that could carry the behavior forward.

Temporary response: behavior fades when reminders, rewards, or social feedback stop.

A persistent outcome does not prove habit. Measure competing mechanisms and keep a credible comparison over time.

Decision lab

Three-horizon scorecard

Define immediate, intermediate, and post-treatment outcomes for one lifecycle intervention.

Deliverable:A measurement plan with mechanism indicators, holdout or withdrawal design, and a date for the keep-revise-stop decision.

Knowledge check

What is the cleanest evidence of durability?

  1. A click during the intervention
  2. The target outcome remains after the intervention ends, with a design that addresses alternative causes
  3. Users say the message was memorable
Reveal the answer and reasoning

Answer: The target outcome remains after the intervention ends, with a design that addresses alternative causes

Durability concerns downstream behavior after treatment, not exposure-period response alone.

Advanced research lab 10

Advanced lab: AI, personalization, and agency

Audit what the system knows, chooses, explains, and lets the user change before optimizing recommendation acceptance.

What you will be able to do

  • Map data provenance and inference risk.
  • Design choice and control autonomy while anticipating reactance.
  • Test acceptance alongside fit, correction, and trust.

Evidence and boundaries

Personalized advertising can increase relevance while covert data collection increases felt vulnerability and can reduce response.Evidence: context-dependent effect

Boundary: The studies predate generative AI and do not quantify every modern personalization context.

Three online vacation-choice experiments found greater recommendation acceptance when users received multiple recommendations and some control over recommendation format.Evidence: emerging evidence

Boundary: Multiple options can add overload in harder tasks, and acceptance is not the same as decision quality.

Psychological reactance research describes how a perceived threat to freedom can motivate counterarguing, resistance, or efforts to restore the threatened choice.Evidence: established mechanism

Boundary: Not every refusal, correction, or low response is reactance; perceived freedom, threat, context, and individual differences matter.

Three controlled song experiments found that randomized or perturbed recommendation signals shifted willingness-to-pay judgments beyond measured preferences.Evidence: context-dependent effect

Boundary: Song valuations in experiments are not completed purchases, durable satisfaction, or evidence that recommendation influence improves welfare.

Large-scale interface audits and experiments document manipulative design patterns that can impair autonomous choice.Evidence: context-dependent effect

Boundary: Pattern labels do not replace design-specific causal evidence or legal advice.

Digital choice interfaces can face explicit restrictions on manipulative design in applicable jurisdictions.Evidence: established mechanism

Boundary: Scope depends on service, conduct, and jurisdiction; obtain qualified legal review.

Source-backed example: One answer versus a choice with control

Setting:Online experiments compared a single algorithmic recommendation with multiple recommendations and user control over the format.

What it shows:Choice and control autonomy increased acceptance in the simulated vacation decisions.

Testable application:Evidence: testable hypothesisFor an AI plan selector, expose assumptions, offer alternatives, let users correct inputs, and compare acceptance with fit and reversal.

Failure mode and guardrails before lift

Failure or backfire
A persuasive recommendation is accepted despite poor fit, opaque inference, reduced autonomy, or unequal errors across groups.
Customer guardrail
Expose material inputs and rationale, offer meaningful alternatives, allow correction, and preserve a no-personalization path.
Commercial guardrail
Read acceptance with fit, willingness to pay, reversals, support, retention, trust, and subgroup outcomes.
  • Disclose material data use and inferred attributes.
  • Provide meaningful alternatives and correction controls.
  • Never optimize acceptance without fit, reversal, trust, and subgroup checks.
Advanced lab · AI and recommendation autonomy
Audit what the system knows, chooses, explains, and lets people change.
Single recommendation

Low comparison effort, but the ranking carries more of the decision. Expose why and an exit.

Multiple recommendations

Preserves comparison. Audit set size, meaningful differences, ordering, and omitted options.

User-controlled recommendations

Let the person set goals, filters, or constraints and revise the resulting choice set.

01 · Inputs

Known or inferred signals

  • Declared goals
  • Observed behavior
  • Uncertain inference
02 · Choice set

Recommendations shown

  1. Primary recommendation
  2. Meaningful alternative
  3. No personalization
03 · Agency

User controls

  • View the rationale
  • Correct inputs
  • Choose alternatives
  • Turn personalization off
Review loop

Expose the rationale → let the user correct inputs → preserve alternatives → learn from the outcome.

Acceptance is incomplete evidence. Also measure fit, correction, reversal, trust, and subgroup outcomes.

Decision lab

Recommendation-system audit

Trace one AI recommendation from collected data through inference, ranking, explanation, user correction, choice, and downstream outcome.

Deliverable:A test matrix for single versus multiple recommendations, visible rationale, edit controls, and a no-personalization baseline.

Knowledge check

Which metric set best protects agency?

  1. Recommendation acceptance alone
  2. Acceptance, fit, corrections, reversals, trust, and subgroup outcomes
  3. Time on page alone
Reveal the answer and reasoning

Answer: Acceptance, fit, corrections, reversals, trust, and subgroup outcomes

Acceptance can rise even when the recommendation is poor or manipulative; downstream quality and control complete the picture.

Advanced field evidence chapter

What the archive can support—and what it cannot.

This is a public-safe audit of historical experiments stored in the GrowthLayer evidence system and interpreted alongside current behavioral-science research. GrowthLayer stores and audits the evidence; it did not independently run every historical test.

The practical standard is intentionally strict: treatment and outcome can be historical facts, while the proposed psychological mechanism remains a hypothesis unless the design measured or isolated it.

Historical portfolio · evidence reconciliation
Reconcile the archive before reading outcomes.

The source-row labels and canonical evidence classifications use different units. Keep the denominators visible so a record label cannot quietly become a portfolio win rate.

  1. 01 · Ingest142source records
  2. 3duplicate rows
  3. 02 · Adjudicate139canonical experiments
Before adjudication · 142 source rows

Source-row outcome labels

Positive
41
Negative
32
Inconclusive
69

These are inherited record labels—not independently recomputed outcomes and not canonical experiment-level winners or losers.

After adjudication · 139 canonical experiments

Canonical evidence classifications

Quantitative-ready
80
Narrative-ready
37
Review-required
22

Readiness describes what each canonical record can responsibly support. It does not encode whether a treatment helped, harmed, or produced an uncertain result.

No pooling. Do not average lifts, combine p-values, or infer a cross-experiment effect without compatible outcomes, units, assignment checks, and a prespecified synthesis method.

Eight corrected field notes

Wins, losses, non-detections, and records that must stop the story.

Public ranges protect prior organizations and clients. The interpretation beside each result is narrower than the headline because accuracy is more valuable than a cleaner anecdote.

01Source-reported resultA pricing architecture package moved completed transactionsReported +15% to +20%; p < .01
What changed
The variant combined product-first ordering, visibility of all price points, and a plan callout.
Primary readout
Completed transactions
Public-safe scope
30,000–35,000 sessions · Four to five weeks
Archive read
The historical record reports a positive result for the package on the transaction outcome.
What the record supports

This bundled pricing-page package outperformed its control in this recorded experiment.

What it cannot establish

The archive cannot attribute the result to price order, a decoy, the callout, or any single behavioral mechanism.

Evidence boundary: Treat the package as one intervention. A follow-up component test would be required to isolate which element mattered.

02Source-reported resultMaking the cheapest price prominent moved the primary outcome downReported -10% to -5%; no public p-value
What changed
The variant showed all price points and made the cheapest option the strongest visual price anchor.
Primary readout
Funnel starts
Public-safe scope
25,000–30,000 sessions · Four to five weeks
Archive read
The historical record labels the variant negative; the public projection does not include the arm-level table needed to check the estimate.
What the record supports

This implementation did not support shipping stronger cheapest-price prominence on this page.

What it cannot establish

It does not show that anchoring caused the result or that prominent low prices generally reduce conversion.

Evidence boundary: Visual prominence, copy, page hierarchy, and the specific choice set travel together in this treatment.

03Directional evidenceA checkout countdown produced directional negative evidenceReported -5% to 0%; p ≈ .08
What changed
The variant added a rate-lock countdown during a considered checkout.
Primary readout
Checkout completion
Public-safe scope
5,000–10,000 sessions · About four weeks
Archive read
The estimate leaned negative but did not clear the conventional two-sided .05 threshold in the archived summary.
What the record supports

The result was sufficient evidence not to roll out this countdown unchanged in this context.

What it cannot establish

It does not establish reactance as the cause or show that accurate urgency cues fail in other journeys.

Evidence boundary: Decision policy and mechanism attribution are separate questions: a cautious product decision can coexist with uncertainty about why.

04Directional evidenceA progress indicator produced an exploratory positive readReported +5% to +10%; p ≈ .09
What changed
A three-arm test compared progress treatments during product selection.
Primary readout
Completed transactions
Public-safe scope
30,000–35,000 sessions · Seven to nine weeks
Archive read
The archive labels a positive variant, while its rounded p-value supports an exploratory rather than confirmatory interpretation.
What the record supports

The treatment generated a follow-up-worthy signal in this funnel stage.

What it cannot establish

It is not a verified replication and does not show that progress indicators reliably increase completion.

Evidence boundary: A three-arm analysis also requires a prespecified comparison and multiplicity policy, neither of which is available in the public projection.

05Non-detectionA linear-versus-radial stepper test showed why magnitude without precision misleadsReported -20% to -10%; p ≈ .38
What changed
The variant changed the progress treatment from a radial display to a linear one.
Primary readout
Completed transactions
Public-safe scope
2,000–3,000 sessions · About six weeks
Archive read
A large negative point estimate coexisted with weak evidence against the null in a small sample.
What the record supports

The experiment did not resolve whether the linear-versus-radial change affected the primary outcome.

What it cannot establish

The point estimate alone cannot justify a stable negative effect or a claim that the treatment had no effect.

Evidence boundary: The missing per-arm counts and interval prevent independent precision, power, or sample-ratio checks.

06Non-detectionReducing fourteen plans to nine did not produce a clear detectionReported 0% to +5%; p ≈ .14
What changed
The variant reduced a product set from fourteen plans to nine.
Primary readout
Completed transactions
Public-safe scope
250,000–300,000 sessions · About eight weeks
Archive read
The large traffic count did not turn the observed estimate into a clear detection under the archived read.
What the record supports

This experiment did not provide clear evidence that this reduction improved the primary outcome.

What it cannot establish

It neither establishes no effect nor overturns the conditional choice-overload literature.

Evidence boundary: Traffic volume is not power by itself; baseline rate, allocation, variance, minimum detectable effect, and stopping policy are also required.

07Integrity-limitedConflicting navigation metadata blocks the attractive storyMetadata conflict: result withheld
What changed
The variant consolidated competing mobile navigation systems.
Primary readout
Orders
Public-safe scope
200,000–250,000 sessions · About five weeks
Archive read
The recorded winner field conflicts with the variation-level conversion ordering and the archived p-value does not support the headline.
What the record supports

The record requires reconciliation against the experiment platform before any outcome claim is published.

What it cannot establish

No lift, winner, Hick-law effect, or causal navigation claim is supportable from the conflicting projection.

Evidence boundary: A large sample cannot rescue internally inconsistent metadata. Integrity checks precede interpretation.

08Integrity-limitedA four-change landing page failed the allocation checkSRM flag: causal read blocked
What changed
One variant bundled four landing-page changes.
Primary readout
Completed transactions
Public-safe scope
65,000–70,000 sessions · Four to five weeks
Archive read
The archive flags a sample-ratio mismatch, so the allocation does not match the expected randomization split.
What the record supports

The correct action is to diagnose instrumentation or allocation before interpreting the outcome.

What it cannot establish

No causal effect, component attribution, or reusable lift is supportable while the mismatch remains unresolved.

Evidence boundary: Even without the mismatch, four simultaneous changes would identify only the package, not which component mattered.

2024–2026 research frontier

What changed in the evidence—not just what got published.

Recent work is less useful as a stream of new tactics than as a correction to how marketers transfer evidence: replication criteria matter, hypothetical magnitude is noisy, interfaces change welfare as well as clicks, and AI autonomy depends on the job and context.

  1. 01
    274 replication attempts across 164 papers

    Published effects often shrink under direct replication

    Replication success depends on the criterion, and the project reported a lower median effect size in replications than in original studies.

    Boundary: The project sampled selected positive claims from earlier publications; its aggregate rates do not forecast a specific marketing intervention.

  2. 02
    20 preregistered hypothetical experiments; n = 16,114

    Hypothetical tests may help with direction, not dependable magnitude

    In this study set, reconstructed scenarios often recovered the direction of field effects while magnitude estimates varied substantially.

    Boundary: The comparison covered five field interventions; stated or imagined choices remain weaker than observed behavior in the target setting.

  3. 03
    Three vaccination field RCTs; n = 314,824

    Transfer can occur for one component and fail for another

    Selected reminder and ownership-language elements transferred, while several additions chosen from intentions or predictions had no detectable benefit.

    Boundary: The results concern booster uptake in one health-system setting and should inform a transfer test, not an automatic rollout.

  4. 04
    Retail-platform field experiment involving 1.6 million consumers

    Choice overload can follow a conditional, non-linear pattern

    Purchase probability followed an inverted-U pattern as recommendation-set size increased, with search initiation explaining much of the later decline.

    Boundary: One recommender system does not determine the best assortment size for another product, audience, or decision journey.

  5. 05
    Emerging working paper; 23,685 cookie choices

    Small consent frictions can materially alter recorded choices

    Hiding rejection behind an extra step reduced rejection by 17.8 percentage points in survey visits and 9.4 points during organic browsing; subtler ordering and visual treatments moved behavior less.

    Boundary: The study is a working paper in one browsing environment; behavioral movement does not establish informed consent or consumer welfare.

  6. 06
    67 completed subscriptions across 34 news sites in four countries

    Cancellation asymmetry can be audited as product behavior

    The audit documented cancellation journeys requiring many steps and, in some cases, a phone call.

    Boundary: An observational interface audit does not estimate causal retention, customer welfare, or the legal status of every flow.

  7. 07
    Marketplace field experiment plus an emerging randomized study, n = 1,608

    Price presentation can change quantity, quality, and the value of decision support

    Upfront mandatory fees changed purchase quantity and tier selection in a ticket marketplace; in a newer randomized study, a digital shopping assistant reduced the reported incidence of overpayment by 63%.

    Boundary: The newer result is a working paper using endowment-funded gift-card purchases, and neither marketplace estimate transfers automatically to subscriptions or ordinary retail.

  8. 08
    Randomized recommendation studies plus one ad field study and four experiments

    AI nudges depend on transparency, preference match, autonomy, and context

    Transparency badges and preference match shaped choice confidence in recommendation studies, while scarcity increased perceived benefits and acceptance of higher-autonomy assistance in a separate study set.

    Boundary: Student shopping tasks, clicks, and hypothetical acceptance do not establish objective decision quality or lasting adoption, and the findings do not justify fabricated scarcity.

The reusable operating system

From principle to decision in seven records.

A mature behavioral strategy is a chain of auditable records. If one link is missing, the claim should get weaker—not more eloquent.

  1. 01
    Decision

    Name who must decide what, by when, and what happens if evidence stays uncertain.

  2. 02
    Observed behavior

    Define the consequential action, denominator, time window, and current journey.

  3. 03
    Competing diagnosis

    Write at least two explanations that predict different observable patterns.

  4. 04
    Treatment contract

    State exactly what changes, what stays fixed, and whether the treatment is a bundle.

  5. 05
    Evidence plan

    Precommit the primary metric, guardrails, allocation, sample logic, and stopping rule.

  6. 06
    Integrity read

    Check instrumentation, allocation, contamination, missingness, and decision-rule fit before lift.

  7. 07
    Decision and reuse

    Store result, evidence strength, business decision, mechanism status, and next test separately.

Public source registry

Inspect the evidence, not just the lesson.

The registry distinguishes study design from claim strength and keeps each limitation next to the citation. Links go to original papers, publisher records, or official sources.

Browse all 49 sources and their limitations
  1. foundational paper · Foundational experiments and model

    Prospect Theory: An Analysis of Decision under Risk

    Daniel Kahneman and Amos Tversky · 1979

    Use: Risk and reference points

    Limit: Original stylized risky-choice studies; parameters and effects are not universal constants.

  2. multi-country replication · Preregistered 19-country replication

    Replicating patterns of prospect theory for decision under risk

    Kai Ruggeri et al. · 2020

    Use: Risk and reference points, Evidence boundaries

    Limit: Replicated structured monetary choices with attenuation and cross-country heterogeneity; not a marketing lift forecast.

  3. program evaluation · 126 RCTs covering 23 million people

    RCTs to Scale: Comprehensive Evidence From Two Nudge Units

    Stefano DellaVigna and Elizabeth Linos · 2022

    Use: Foundations, Experimentation

    Limit: Government nudge-unit portfolio; channel mix and outcomes differ from commercial marketing.

  4. review · Current generalizability review

    Generalizability of choice architecture interventions

    Barnabas Szaszi, Daniel G. Goldstein, Dilip Soman, et al. · 2025

    Use: Foundations, Experimentation

    Limit: Synthesizes heterogeneity and transfer problems; it does not identify one best intervention.

  5. meta-analysis · 99 observations across 7,202 participants

    Choice overload: A conceptual review and meta-analysis

    Alexander Chernev, Ulf Bockenholt, and Joseph Goodman · 2015

    Use: Choice architecture

    Limit: Overload is conditional on complexity, task difficulty, preference uncertainty, and decision goal.

  6. field experiment · Field experiment and observational evidence

    Salience and Taxation: Theory and Evidence

    Raj Chetty, Adam Looney, and Kory Kroft · 2009

    Use: COM-B and friction, Pricing

    Limit: Grocery tax salience and alcohol-tax evidence; transfer to other prices or interfaces must be tested.

  7. field experiment · Randomized field experiment

    The Role of Application Assistance and Information in College Decisions

    Eric P. Bettinger, Bridget Terry Long, Philip Oreopoulos, and Lisa Sanbonmatsu · 2012

    Use: Behavioral Audit, COM-B and friction

    Limit: Financial-aid application context with hands-on assistance; not evidence that copy simplification alone is sufficient.

  8. program evaluation · Large-scale program evaluation

    Social norms and energy conservation

    Hunt Allcott · 2011

    Use: Social influence

    Limit: Residential energy program; average effects conceal household and program heterogeneity.

  9. meta-analysis · Theory-guided anchoring meta-analysis

    Recovering the Anchoring of Economic Valuations

    Liang Guo · 2024

    Use: Pricing and context

    Limit: Shows opposing and moderated treatment effects; an arbitrary number is not a dependable pricing tactic.

  10. foundational paper · Controlled zero-price experiments

    Zero as a Special Price: The True Value of Free Products

    Kristina Shampanier, Nina Mazar, and Dan Ariely · 2007

    Use: Pricing and free offers

    Limit: Specific low-cost product choices; free does not erase time, shipping, privacy, or switching costs.

  11. controlled experiment · Controlled market experiment

    Drip pricing and its regulation: Experimental evidence

    Alexander Rasch, Miriam Thone, and Tobias Wenzel · 2020

    Use: Pricing and free offers

    Limit: Experimental market; legal obligations and buyer expectations vary by jurisdiction and category.

  12. meta-analysis · 86-study meta-analysis

    A Meta-Analysis of Quasi-Hyperbolic Discounting

    Stephen L. Cheung, Agnieszka Tymula, and Xueting Wang · 2026

    Use: Motivation over time

    Limit: Discounting estimates vary by reward domain and selective reporting; they do not prescribe a single reminder cadence.

  13. review · Economics review

    Commitment Devices

    Gharad Bryan, Dean Karlan, and Scott Nelson · 2010

    Use: Motivation over time

    Limit: Commitment can help sophisticated users but can also impose costs; easy exit and informed consent matter.

  14. review · Psychological reactance review

    Understanding Psychological Reactance: New Developments and Findings

    Christina Steindl, Eva Jonas, Sandra Sittenthaler, Eva Traut-Mattausch, and Jeff Greenberg · 2015

    Use: AI and personalization, Choice architecture

    Limit: Synthesizes varied settings and outcomes; a freedom threat does not imply that every refusal, correction, or low conversion is reactance.

  15. controlled experiment · Three controlled recommendation experiments

    Effects of Online Recommendations on Consumers’ Willingness to Pay

    Gediminas Adomavicius, Jesse C. Bockstedt, Shawn P. Curley, and Jingjing Zhang · 2018

    Use: AI and personalization

    Limit: Digital-song willingness-to-pay judgments in experiments; stated valuations are not completed purchases or a welfare measure.

  16. online experiment · Two experimental studies

    Shining a Light on Dark Patterns

    Jamie Luguri and Lior Jacob Strahilevitz · 2021

    Use: Ethics, AI and personalization

    Limit: Online experiments on dark-pattern exposure; effects and legal standards vary by design and population.

  17. official guidance · International practice review

    Fixing Frictions: Sludge Audits Around the World

    OECD · 2024

    Use: COM-B and friction, Ethics

    Limit: Public-policy cases and audit guidance; commercial implementation needs local customer and legal review.

  18. regulation · Binding EU regulation

    Regulation (EU) 2022/2065: Digital Services Act

    European Parliament and Council of the European Union · 2022

    Use: Ethics, AI and personalization

    Limit: Legal scope is service-, conduct-, and jurisdiction-specific; this course is not legal advice.

  19. multi-country replication · 274 attempted replications; 55.1% were significant in the original direction

    Investigating the replicability of the social and behavioural sciences

    Andrew H. Tyner et al. · 2026

    Use: Evidence boundaries, Experimentation

    Limit: The project sampled positive claims published from 2009 to 2018; median Pearson’s r fell from 0.25 to 0.10 and success ranged from 28.6% to 74.8% across criteria, so no single rate forecasts a specific intervention.

  20. online experiment · 20 preregistered hypothetical experiments, n = 16,114

    Hypothetical nudges provide directional but noisy estimates of real behavior change

    Linnea Gandhi, Anoushka Kiyawat, Colin Camerer, and Duncan J. Watts · 2025

    Use: Evidence boundaries, Experimentation

    Limit: The scenarios reconstructed five field experiments and usually recovered direction in this set, but magnitude estimates varied widely; hypothetical choices are not dependable field-effect forecasts.

  21. field experiment · Three preregistered field RCTs, n = 314,824; selected reminders added 1.13–1.93 percentage points

    Field testing the transferability of behavioural science knowledge on promoting vaccinations

    Silvia Saccardo et al. · 2024

    Use: Evidence boundaries, Experimentation, Motivation over time

    Limit: COVID-19 booster messages in one health-system setting; reminders and ownership language transferred, while several intention- or prediction-selected additions had no detectable benefit.

  22. field experiment · Platform field experiment involving 1.6 million consumers

    The Choice Overload Effect in Online Recommender Systems

    Xiaoyang Long et al. · 2024

    Use: Choice architecture, Evidence boundaries

    Limit: One retail platform and retargeting recommender implementation; up to 64% of the later purchase decline was attributed to fewer consumers starting search, and patterns varied by category, price, and timing.

  23. observational study · 67 completed subscriptions across 34 news sites in four countries

    Staying at the Roach Motel: Cross-Country Analysis of Manipulative Subscription and Cancellation Flows

    Ashley Sheil, Gunes Acar, Hanna Schraffenberger, Raphael Gellert, and David Malone · 2024

    Use: COM-B and friction, Ethics

    Limit: Some cancellation flows required up to ten steps and four required a call; this observational audit does not estimate causal conversion, retention, welfare, or the legal status of every design.

  24. controlled experiment · Emerging working paper; randomized marketplace experiment, n = 1,608

    Measuring and Mitigating Drip Pricing Overcharge: Evidence from an Online Marketplace Experiment with a Digital Shopping Assistant

    Benjamin Lu, Daniel Markovits, Andrew Miller, and Rory Van Loo · 2025

    Use: Pricing and free offers, AI and personalization, Ethics

    Limit: A working paper using endowment-funded gift-card purchases; drip pricing raised average prices paid by up to about 10% and an assistant reduced overpayment incidence by 63%, but effects may differ in other markets.

  25. field experiment · Large-scale ticket-marketplace field experiment

    Price Salience and Product Choice

    Tom Blake, Sarah Moshary, Kane Sweeney, and Steve Tadelis · 2021

    Use: Pricing and free offers, Ethics

    Limit: One ticket marketplace with product-quality and seller responses; quality substitution accounted for at least 28% of the reported revenue decline, which is not a universal estimate for other categories.

  26. field experiment · One Google Ads field study and four online experiments with about 2,000 people

    Consumer Acceptance of High-Autonomy AI Assistants Is Driven by Perceived Benefits in Online Shopping Settings Characterized by Scarcity

    Darius-Aurel Frank, Michal Folwarczny, and Tobias Otterbring · 2026

    Use: Scarcity and credibility, AI and personalization

    Limit: Scarcity increased perceived benefits and acceptance of high-autonomy assistance in this study set, but ad clicks and hypothetical choices do not establish adoption or justify manufactured scarcity.

Capstone

Build an evidence-led Behavioral Audit

Choose one consequential marketing decision and produce an intervention portfolio that improves customer and business outcomes without hiding cost, constraining agency, or overstating evidence.

  1. Behavior and baselineTarget behavior, population, journey moment, baseline rate, and measurement window.
  2. Evidence packetBehavioral data, customer evidence, source ladder, and documented unknowns.
  3. DiagnosisCOM-B map, friction ledger, reference points, and at least two competing explanations.
  4. Intervention portfolioA structural fix, a communication treatment, and a no-change or simpler comparator.
  5. ExperimentUnit, allocation, primary outcome, minimum useful effect, guardrails, and analysis plan.
  6. Ethics and operationsTruth owner, consent and autonomy review, accessibility check, failure recovery, and rollback.
  7. Decision memoPrecommitted scale, revise, and stop rules plus the next decision each result enables.

Precommit the stop rules

  • Stop when the primary outcome misses the predeclared practical threshold.
  • Stop when trust, error, refund, complaint, or subgroup guardrails breach their limits.
  • Stop when the mechanism check fails even if a vanity metric improves.
  • Do not scale beyond the tested population or context without a transfer test.
Want to teach this as a team capability?

Turn the audit into a live working session.

Review the training approach, or bring a real decision to a private conversation.

Continue with the behavioral economics experimentation guide or the behavioral economics in marketing article.