One third, and that is the good number
Microsoft evaluated its own experiments — well-designed, well-executed, each built to improve a specific metric. Roughly one third succeeded at improving the metric they were designed to improve.
The fuller breakdown from the same programme: about a third of ideas were positive and statistically significant, a third were flat, and a third were negative and statistically significant. Ideas that made things measurably worse were as common as ideas that worked.
Failure rates elsewhere run higher. Published figures put Microsoft broadly around 66%, Bing near 85%, Google Ads and Netflix around 90%, and Airbnb around 92%. More mature and better-optimised the product, higher the failure rate, because the easy gains are gone.
Microsoft's comparatively strong number came at a cost worth noting: significant upfront work went into scoping and refining ideas before they entered an experiment at all.
If two thirds of carefully considered ideas from well-resourced product organisations fail, the working assumption for any business planning a pilot should be that the idea is probably wrong. A pilot exists to find that out cheaply.
The number most teams never calculate
Here is the part that changes how results should be read.
At the median organisation, roughly 10% of experiments move the metric they were designed to move. Run the arithmetic on false positive risk at that success rate, with a standard 5% significance threshold and 80% power, and the probability that a statistically significant result is actually a false positive comes out around 22%.
Roughly one in five wins is not a win.
Most teams read p < 0.05 as a 5% chance of being wrong. It is not. The false positive risk depends on how often the ideas being tested are true in the first place, and in most organisations that base rate is low. Weak prior, significant result, high chance of nothing.
This is the mathematical case for the upfront work Microsoft did. Better ideas going in raise the base rate, which lowers the false positive risk on everything that comes out.
A pilot produces a decision, not a result
The distinction is the whole discipline.
A result is a number. A decision is what the business does differently because of the number. The gap between them is filled by something that has to exist before the pilot runs: a rule stating what result leads to what action.
Without that rule, the result gets interpreted after the fact, by people who have already formed views and now have a number to support them. This is not dishonesty. It is the default behaviour of every team, and the only reliable defence is deciding in advance.
The single most useful question to ask before any pilot: what result would change our mind? If nobody can answer it, the pilot is not going to settle anything.
Six elements of a pilot worth running
1. A hypothesis with a mechanism. Not "we think this will improve conversion," but "we think conversion improves because this removes the pricing uncertainty buyers raise on calls." The mechanism is what makes the result interpretable, and what makes a failure informative rather than merely disappointing.
2. One primary metric, chosen first. Multiple primary metrics guarantee a winner, because something will move. Secondary metrics can be tracked; only one decides.
3. A minimum detectable effect, set before the sample. Decide the smallest improvement that would actually justify the change — commercially, not statistically. That number determines the sample size required. Running the pilot the other way round, gathering whatever sample is available and seeing what emerges, produces a test that cannot detect anything smaller than a large effect while appearing to.
4. A decision rule, written down. If the result is above X, we do this. If below, we do that. If between, we do this specific third thing. Write it before the pilot starts and circulate it.
5. A fixed duration. Set the end date at the start, based on the sample calculation and at least one full business cycle. Weekday and weekend behaviour differ; a pilot that runs Tuesday to Friday measures Tuesday-to-Friday customers.
6. Guardrail metrics. What must not get worse. A pilot that improves conversion while degrading retention or margin has not succeeded, and without guardrails nobody finds out until later.
The peeking problem
The most common way a pilot lies is being checked daily and stopped when the number looks good.
Statistical significance thresholds assume a single test at a predetermined sample size. Checking repeatedly and stopping at the first favourable reading inflates the false positive rate substantially — the test becomes a search for a moment when random variation happened to look like signal.
The observed effect at that moment is also biased upward. Teams stop early on an inflated number, ship, project revenue from it, and then watch the real effect settle at a fraction of what was reported. The credibility damage from that sequence is usually larger than the value of the change itself.
The fix costs nothing: set the duration in advance, and do not look at the primary metric until it ends. Watch guardrails for genuine harm, not the metric that decides.
Pilots that are not A/B tests
Most of the experimentation literature is about digital product testing at scale. Most businesses reading this do not have that traffic, and the underlying logic still applies.
Concept and message testing. Structured comparison of positioning or messaging directions with a defined audience, before commitment. Cheap, fast, and useful precisely because positioning determines pricing and channel decisions downstream.
Geographic pilots. Run in one market, hold another comparable market as control. The control is what makes it a pilot rather than a launch. Without one, a seasonal effect is indistinguishable from a treatment effect.
Smoke tests. A landing page and a paid campaign against a proposition that does not exist yet, measuring intent to proceed. Answers whether demand exists before anything is built. The ethical version stops at expressed interest and tells people honestly where things stand.
Structured focus groups. Useful for understanding why, not for measuring how many. Small qualitative samples generate hypotheses. They do not validate them, and treating six enthusiastic participants as evidence of demand is one of the more expensive errors available.
Controlled digital pilots. Small paid campaigns across defined audience segments, comparing response. Available to almost any business and frequently the fastest route to a directional answer.
Where pilots mislead
Novelty effects. Existing users respond to change because it is different. The effect decays. Short pilots on established audiences systematically overstate.
The wrong metric. Click-through improves while purchases fall. Sign-ups rise while activated users do not. Choose the metric closest to the commercial outcome, even when it is slower and noisier.
No counterfactual. Running a campaign and observing that sales rose is not a pilot. Without a control, you have measured a period, not a treatment.
Sample that is not the market. Testing on the most engaged existing customers answers a question about the most engaged existing customers. This is the same bias that inflates product-market fit scores.
Stopping at the number. The result tells you what happened. The follow-up work tells you why, and why is what transfers to the next decision.
Diagnostic: is this pilot going to settle anything?
Seven tests, run before it starts.
The hypothesis includes a mechanism, not only a predicted direction.
One primary metric is named, and it is the one closest to commercial outcome.
A minimum detectable effect was set on commercial grounds, and the sample was sized from it.
A decision rule is written down and circulated before launch.
The end date is fixed, and covers at least one full business cycle.
Guardrail metrics are defined.
Someone can state the result that would cause the business to abandon the idea.
Failing test seven means this is a demonstration, not a pilot.
What this produces
Permission to spend, or permission to stop.
That is the output, and both directions are valuable. Two thirds of good ideas from strong teams do not work. A pilot programme that never kills anything is not producing evidence — it is producing justification, at a cost that shows up later as capital committed to an assumption nobody tested.
The reason to do this properly is not methodological neatness. It is that the next allocation depends on it. Evidence closes the question, and a closed question is what releases the budget for the phase that follows.
Frequently asked questions
What proportion of experiments actually succeed?
At Microsoft, roughly one third of well-designed experiments improved the metric they targeted, with about a third flat and a third actively negative. Failure rates elsewhere run higher — reported figures reach 85% at Bing and above 90% at Google Ads, Netflix, and Airbnb.
How reliable is a statistically significant result?
Less than most teams assume. At a 10% base rate of ideas that work — typical of the median organisation — with a 5% significance threshold and 80% power, roughly 22% of statistically significant results are false positives. Improving the quality of ideas entering testing is what improves that number.
Why is stopping a test early a problem?
Significance thresholds assume one test at a predetermined sample size. Checking daily and stopping at the first favourable reading inflates false positives and biases the measured effect upward. Set the duration in advance and hold to it.
How long should a pilot run?
Long enough to reach the sample size implied by the minimum detectable effect, and covering at least one full business cycle. Behaviour varies across days of the week and across the month, and a short window measures a slice rather than the market.
Can a small business run meaningful pilots?
Yes. Geographic pilots with a control market, controlled digital campaigns across defined segments, concept and message tests, and smoke tests all work at small scale. What cannot be scaled down is the requirement for a control and a decision rule set in advance.
Are focus groups a form of validation?
They generate hypotheses and explain reasoning. They do not measure demand. A small enthusiastic group is not evidence that a market exists, and treating it as such is a common and expensive error.
What is the single most important thing to fix?
Write down what result would change your mind, before the pilot starts. Almost every other failure mode is downstream of not having done that.
Sources
Kohavi, R., Crook, T. and Longbotham, R., Online Experimentation at Microsoft (2009); Kohavi, R., Trustworthy Online Controlled Experiments
Kohavi, R. and Thomke, S., The Surprising Power of Online Experiments, Harvard Business Review (2017)
Published experiment failure rates across Microsoft, Bing, Google Ads, Netflix, and Airbnb
False positive risk calculation at typical organisational success rates, standard significance and power thresholds
Structure your next phase
Zerologic designs and runs validation programmes — pilots structured to close a question, with the decision rule agreed before anything goes live.
Talk to us: partners@zerologic.io · zerologic.io



