eBay turned off its search ads and almost nothing happened
In March 2012, eBay stopped buying brand-keyword paid search on MSN and measured what happened to traffic.
Around 99.5% of the forgone paid clicks were immediately picked up by organic search. Substitution was nearly complete. The ads had been buying clicks the company would have received for free, and every attribution system in the stack had been recording those clicks as advertising-driven conversions.
Blake, Nosko, and Tadelis published the result in Econometrica in 2015, using randomisation across 210 media markets. The non-branded search finding is the one that should stop a marketing meeting: the same spend produced an estimated return of over 4,100% by regression and minus 63% by experiment.
Not a different number. A different sign.
Their heterogeneity finding explains the mechanism. Non-branded ads did have a positive effect on new and infrequent users. On average returns were negative because most impressions were served to frequent customers who were coming anyway. The targeting was working exactly as designed, and the design was the problem.
It is not just search, and not just eBay
Facebook, 15 randomised controlled trials. Gordon, Zettelmeyer, Bhargava, and Chapsky compared experimental results against observational methods across 15 US campaigns — around 500 million user-experiment observations and 1.6 billion impressions — published in Marketing Science in 2019.
In half the experiments, observational estimates of purchase lift were wrong by a factor of three or more. In one case, a true lift of 2.4% was estimated at 1,306%.
The critical detail is that the errors do not run in one direction. Observational methods sometimes overstate enormously and sometimes understate. That rules out the obvious workaround: you cannot apply a discount factor to platform-reported numbers and arrive at truth. There is no fudge factor, because there is no consistent bias to correct.
And eBay is an extreme case, which matters. Coviello and colleagues replicated the experiment at Edmunds.com and found the same substitution effect at far smaller magnitude — less than half of paid search traffic returned organically. eBay was exceptionally well known, so its brand searches were unusually substitutable. Your business is probably somewhere between the two, and the only way to find out where is to test.
One disclosure worth noting: both landmark studies were run with the cooperation of the platforms concerned, and several authors held positions at those companies while conducting the work. It is unusual for such results to be published at all.
What incrementality actually measures
Attribution asks which touchpoints preceded a conversion. Incrementality asks whether the conversion would have happened without the touchpoint.
Only the second supports a spending decision, because the first has no way of distinguishing advertising that caused a purchase from advertising that was present when a purchase was going to happen.
Four practical methods.
Geo holdout tests. Suspend or increase spend in a set of matched geographies while holding others constant. The most accessible method for most businesses, and the one that works without platform cooperation.
Ghost ads and PSA holdouts. The control group is served an unrelated ad, or the ad slot is logged without delivery. Requires platform support and is the cleanest design available.
Full-channel blackouts. Turning a channel off entirely for a defined period. Crude, expensive, and unambiguous — this is what eBay did.
Marketing mix modelling. Statistical modelling of spend against outcomes across channels and time. Not experimental, so it produces estimates rather than causal proof, but it covers channels that cannot be experimentally isolated and has returned to prominence as user-level tracking has degraded.
How to run one properly
The pilot design discipline applies without modification.
Decide the decision first. What spend change follows from what result. Written down before the test starts.
Set a minimum detectable effect on commercial grounds. The smallest true incremental return that would change the budget. That threshold determines the size and duration of the test.
Accept that the test costs money. A holdout means deliberately not spending in some markets, or spending in markets where you would not have. That cost is the price of the answer, and it is almost always smaller than a year of misallocated budget.
Run for at least one full purchase cycle, and preferably more. Suppressing advertising has delayed effects that a two-week test will miss entirely.
Expect a lower number than the platform reports. If the incremental figure comes back close to the platform's, check the test design before celebrating.
What businesses find when they test
Three patterns recur often enough to anticipate.
Branded search is the most commonly overstated line item. People searching your brand name have already decided to visit you. Some share of that spend is buying traffic that was free, and the share is measurable — eBay's near-total substitution is the extreme, not the norm, but the direction is consistent.
Retargeting is frequently second. By construction it reaches people who already visited, and a meaningful proportion of them were returning regardless.
Upper-funnel activity is often understated. Because its effects are delayed and diffuse, attribution systems under-credit exactly the activity that creates future demand. That is the measurement asymmetry that pushes budgets toward activation, seen from the other side.
The pattern is consistent: attribution over-credits the spend closest to the purchase and under-credits the spend that produced the intent.
Where incrementality testing goes wrong
The test is too small. A lift of a few percent needs substantial volume to detect. Underpowered incrementality tests conclude "no effect" for channels that work.
Geographies are not matched. Control and treatment markets must be comparable on the outcome before the test. Unmatched markets measure the difference between the markets.
The window is too short. Advertising effects persist after exposure ends. A short blackout will understate the loss, making the channel look less valuable than it is.
One test is treated as permanent. Incrementality changes with creative, competition, audience saturation, and season. It is a repeating measurement.
The result is ignored. The most common outcome in practice. A test showing a channel is not incremental threatens budgets, headcount, and agency relationships, and organisations frequently commission the test and then decline to act on it.
Diagnostic: is measurement causal?
Six tests.
At least one incrementality test has been run in the past year, on a material channel.
The decision rule was written before the test.
Control and treatment groups were matched on the outcome metric beforehand.
The test ran for at least one full purchase cycle.
Reported channel returns are not taken solely from the platforms selling the media.
A budget decision has actually changed as a result of an incrementality finding.
Test six is the one that matters. Measurement that never changes an allocation is expensive reporting.
What this produces
A defensible answer to the only question that matters about media spend: what did it cause.
That answer is usually less flattering than the dashboard and more useful than any amount of attribution modelling. The eBay result is worth carrying as a permanent reminder — the same spend, measured two ways, produced returns of positive 4,100% and negative 63%. One of those numbers was in a report someone acted on.
The businesses that test do not necessarily spend less. Frequently they spend the same amount differently, having discovered that some of what looked efficient was accounting for demand it did not create, and that some of what looked wasteful was creating demand it never got credit for.
Frequently asked questions
What is incrementality testing?
Measuring the difference in outcomes between a group exposed to advertising and a comparable group that was not. It establishes what the advertising caused, rather than what happened near it.
How different are incremental and attributed results?
Potentially very. In eBay's experiment, non-branded search returned over 4,100% by regression and minus 63% by experiment. Across 15 Facebook RCTs, half the observational estimates were off by a factor of three or more, including one case of 2.4% true lift estimated at 1,306%.
Can we just discount platform-reported numbers?
No. The errors run in both directions — sometimes overstating enormously, sometimes understating — so there is no consistent bias to correct with a multiplier.
Is branded search always a waste?
No, and eBay is an extreme case. A replication at Edmunds.com found less than half of paid search traffic returned organically, against near-total substitution at eBay. Substitution depends on how well known the brand is, which is exactly why it needs measuring rather than assuming.
What is the easiest incrementality test to run?
A geo holdout. Suspend or increase spend in matched markets while holding others constant. It requires no platform cooperation and works for most businesses with sufficient volume.
Does marketing mix modelling count as incrementality?
Not strictly. It is statistical rather than experimental, so it produces estimates rather than causal proof. It remains useful for channels that cannot be experimentally isolated, and its value has risen as user-level tracking has degraded.
How often should incrementality be tested?
Periodically, on material channels. Incrementality shifts with creative, competition, saturation, and season, so a single test is a snapshot rather than a standing fact.
Sources
Blake, T., Nosko, C. and Tadelis, S., Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment, Econometrica 83(1), 2015 — conducted with eBay Research Labs
Gordon, B., Zettelmeyer, F., Bhargava, N. and Chapsky, D., A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook, Marketing Science 38(2), 2019 — authors held research positions at Facebook to access the data
Coviello, L., Gneezy, U. and Goette, L., replication at Edmunds.com (2017)
Structure your next phase
Zerologic designs measurement that separates what advertising caused from what happened near it — geo tests, holdouts, and modelling, with the decision rule agreed in advance.
Talk to us: partners@zerologic.io · zerologic.io



