The best study on whether nudges work found they work about one sixth as well
Stefano DellaVigna and Elizabeth Linos assembled 126 randomised controlled trials covering 23 million individuals — every trial run by two of the largest nudge units in the United States. They compared them against nudge trials published in academic journals, drawn from two meta-analyses.
In academic papers, the average nudge produced an 8.7 percentage point increase in take-up — a 33.5% lift over a control base of 26%.
In the nudge units, the average effect was 1.4 percentage points — an 8.1% lift over a control base of 17.2%. Highly statistically significant, and roughly one sixth the size.
They tested five explanations for the gap. The finding that matters: publication bias in the academic journals, made worse by low statistical power, accounts for the full difference. Academic involvement explained none of it.
Most forecasters, academics and practitioners alike, over-predicted the effect.
Read that as calibration, not as debunking
The temptation is to treat this as evidence that behavioural science is oversold. That reading is wrong, and the numbers show why.
A 1.4 percentage point improvement across a population of millions, achieved by rewording a letter, is enormous in absolute terms and costs almost nothing. That is a genuinely good return on a copy change.
What the finding kills is the promise of transformation. Behavioural interventions produce modest effects, reliably, at very low cost. Anyone offering large effects from a nudge is quoting the published literature, which the largest available study says is inflated by selective publication.
That distinction is the whole practical value. It tells you what to budget, what to expect, and what to be suspicious of.
Which frameworks survived
Behavioural science took heavy damage in the replication crisis. Several findings that dominated marketing decks for a decade — ego depletion, social priming, power posing — failed large replication attempts, in some cases with original authors withdrawing support.
Two frameworks came through with their methodology intact, because both were built for applied trials rather than for laboratory demonstrations.
COM-B
Developed by Susan Michie and colleagues and published in Implementation Science in 2011, as the core of the Behaviour Change Wheel.
Behaviour requires three things simultaneously:
Capability. Can the person do it — physically and psychologically. Do they have the knowledge, skill, and capacity.
Opportunity. Does the environment allow it. Is it available, affordable, permitted, and socially acceptable.
Motivation. Do they want to — including both reflective motivation, meaning intentions and beliefs, and automatic motivation, meaning habits and impulses.
The value is diagnostic. If a behaviour is not happening, COM-B forces the question of which of the three is missing before anyone designs an intervention.
That ordering matters commercially. Most marketing addresses motivation by default. When the actual barrier is capability or opportunity, motivational work produces people who want to act and still cannot.
EAST
The Behavioural Insights Team's applied framework: make it Easy, Attractive, Social, Timely.
Easy. Reduce friction. Default settings, fewer steps, simpler language.
Attractive. Draw attention. Personalisation, salience, meaningful incentives.
Social. Use accurate social information. What comparable people actually do.
Timely. Intervene when the person is receptive, usually near a decision or a moment of change.
EAST is a design checklist rather than a theory, and it comes from the organisation whose trials the DellaVigna and Linos data partly covers. It is the practical companion to COM-B's diagnosis.
Diagnose before designing
The two frameworks operate in sequence, and running them in the wrong order is the most common failure.
COM-B first. Which of capability, opportunity, or motivation is the binding barrier for this specific behaviour, for this specific audience.
EAST second. Given the barrier, which design levers address it.
A checkout abandonment problem illustrates it. If the barrier is opportunity — the payment method the customer uses is not offered — then making the page more attractive changes nothing. If the barrier is capability — the customer cannot tell what the delivery cost will be — then a discount does not help either.
Most conversion work applies EAST-style tactics without the COM-B step, which is why so much of it produces effects indistinguishable from noise.
What to expect, in numbers
Set expectations from the at-scale data rather than from case studies.
A well-designed intervention on an existing behaviour: low single-digit percentage point improvement. That is the honest central estimate.
Effects are larger where friction is genuinely high. Removing a real barrier outperforms adding persuasion.
Effects are smaller where the population is already motivated. There is little headroom among people who were going to act anyway.
Detecting an effect of this size requires real sample. A 1.4 percentage point improvement is undetectable in a test with a few hundred participants, which means most small businesses running behavioural tests will conclude nothing either way. That is a measurement limitation, not evidence of no effect.
Where behavioural science is oversold
Lists of cognitive biases. The catalogue format implies each entry is equally supported. Many are not, and several prominent ones failed replication. A framework with published trials behind it is different from a list with a hundred entries.
"Neuromarketing" claims. Brain imaging in commercial settings rarely supports the inferences drawn from it.
Persuasion framing for structural problems. If the product is hard to find, hard to pay for, or hard to understand, no amount of message optimisation resolves it.
Single-study findings presented as laws. The replication crisis was specifically about compelling single studies that did not hold. Anything resting on one memorable experiment deserves the same scepticism this series applies elsewhere.
The ethical line, stated plainly
Behavioural techniques work on people who did not consent to being worked on. That is not a reason to avoid them, but it is a reason to hold a position.
The workable test: would the intervention still be acceptable if the person could see it? Reducing friction, clarifying options, and timing a message well all pass. Manufactured scarcity, dark patterns in cancellation flows, and social proof that misrepresents what others do all fail.
The commercial argument aligns with the ethical one. Interventions that fail this test tend to produce short-term conversion and long-term damage to the thing the Drive stage is actually building, which is being thought of favourably when the buyer next has the need.
Diagnostic: is this behavioural design or decoration?
Six tests.
The specific behaviour being targeted is written down, in observable terms.
The binding barrier has been identified as capability, opportunity, or motivation — not assumed to be motivation.
The intervention addresses that barrier specifically.
The expected effect size is stated in advance, and it is modest.
The test has enough sample to detect an effect of that size.
The intervention would survive the customer seeing exactly what was done and why.
Test two is the one that separates behavioural design from copywriting with citations attached.
What this produces
Small, reliable improvements to behaviours that already nearly happen.
That is the honest scope, and it is worth having. Across enough interactions, single percentage points compound into material revenue, and the interventions are cheap enough that the return holds even at modest effect sizes.
What behavioural science does not do is create demand where none exists, rescue a product nobody wants, or substitute for being known when the need arises. Those are different problems addressed by different parts of this stage.
The businesses that get value from this work are the ones that use it to remove friction from a path people were already trying to walk. The ones that get nothing are usually applying persuasion to a structural barrier.
Frequently asked questions
Do nudges actually work?
Yes, at about one sixth the size published research suggests. Across 126 trials covering 23 million people, average take-up rose 1.4 percentage points, against 8.7 points in academic journal papers. Publication bias, worsened by low statistical power, accounts for the entire gap.
What is COM-B?
A model holding that behaviour requires capability, opportunity, and motivation simultaneously. Its value is diagnostic — establishing which of the three is the binding barrier before designing an intervention.
What is the EAST framework?
The Behavioural Insights Team's design checklist: make the behaviour Easy, Attractive, Social, and Timely. It is applied after diagnosis, not instead of it.
How much improvement should we expect?
Low single-digit percentage points on an existing behaviour is the realistic central estimate. Larger effects are possible where genuine friction is being removed, and case studies reporting large gains are usually selected examples.
Why do our behavioural tests show nothing?
Frequently because the sample is too small. Detecting a 1.4 percentage point effect requires substantial traffic. Inconclusive results at small scale are a measurement limitation rather than evidence the intervention failed.
Are cognitive bias lists useful?
Less than they appear. The format implies uniform evidential support, and several prominent entries failed large replication attempts. Prefer frameworks with published trials behind them.
Where is the ethical line?
A workable test is whether the intervention would still be acceptable if the customer could see exactly what was done and why. Reducing friction and improving timing pass. Manufactured scarcity and obstructive cancellation flows do not.
Sources
DellaVigna, S. and Linos, E., RCTs to Scale: Comprehensive Evidence From Two Nudge Units, Econometrica 90(1), 2022 — 126 RCTs across 23 million individuals
Michie, S., van Stralen, M. and West, R., The Behaviour Change Wheel, Implementation Science (2011) — COM-B
Behavioural Insights Team, EAST: Four Simple Ways to Apply Behavioural Insights
Published replication failures in behavioural research, including ego depletion and social priming
Structure your next phase
Zerologic designs behavioural campaigns from diagnosis rather than from tactics — identifying the actual barrier before deciding what to change.
Talk to us: partners@zerologic.io · zerologic.io



