Start with the statistic that props up the field
Almost every article about product innovation opens the same way. Thirty thousand new products launch each year, and 95% of them fail, attributed to Clayton Christensen of Harvard Business School.
The figure is repeated by MIT's professional programmes, by consultancies, by product blogs, and in pitch decks. It is very difficult to trace to a primary study. Christensen's own published work does not present it as a research finding with a stated method.
A 2022 paper in the Journal of Marketing and Consumer Behaviour in Emerging Markets examined the claim directly and concluded that the widespread belief in a 90% new-product failure rate is not supported by empirical evidence. Category-level figures that do have method behind them run considerably lower and vary enormously by sector — grocery and consumer packaged goods sit high, other categories much lower.
This matters beyond pedantry. A framework introduced by an unverifiable statistic invites the reader to accept everything that follows on the same terms. Jobs to be Done is genuinely useful. It does not need the 95%.
Two frameworks, one name
The larger problem is that Jobs to be Done refers to two different things, and most teams use one while expecting results from the other.
The narrative branch is Christensen's. Customers "hire" products to make progress in a situation. The milkshake study is its teaching device: commuters buying milkshakes in the morning were hiring them to make a boring drive more tolerable, which reframes the competition as bagels and boredom rather than other milkshakes.
The measurable branch is Anthony Ulwick's Outcome-Driven Innovation. A job is broken into steps, each step generates desired outcome statements, and customers rate every outcome on importance and current satisfaction. The gap produces a score that ranks where to invest.
On sequence, since it is usually reported backwards: Ulwick named ODI in 1999, published it in Harvard Business Review in 2002, and expanded it in What Customers Want in 2005. Christensen introduced Jobs Theory in The Innovator's Solution in 2003, citing Strategyn's work on job and outcome-based thinking.
The narrative branch became far more popular. The measurable branch is the one that produces a decision.
The narrative branch: what it does and cannot do
It does one thing extremely well. It breaks a team out of describing customers by attribute — age, income, industry, company size — and moves them to describing what the customer is trying to accomplish. That shift reliably changes what a team considers a competitor, which changes positioning and roadmap.
Its limitation is structural, and it is worth stating plainly.
Christensen's version does not specify what would falsify a job statement. If a hypothesised job fails to predict purchase behaviour, the analyst re-narrates. The job was not commute sustenance, it was emotional reassurance during a boring drive. The theory refits the data every time.
A framework that cannot be wrong cannot be tested. It can still be valuable as a way of seeing, and it is. But an untestable framework should not be the last step before a funding decision, because it will confirm whatever the team already believed with a fresh vocabulary attached.
The measurable branch: how ODI actually works
Five steps. The detail is the point, because the detail is what the narrative version lacks.
1. Define the job. State what the customer is trying to accomplish, independent of any solution. Not "I need a car" but "arrive at work on time and unstressed."
2. Map the job into steps. The full process from start to finish. For a commute: decide when to leave, choose a route, travel, park, walk in.
3. Generate desired outcome statements. For each step, the metrics the customer uses to judge success. These are written to a strict form — a direction of improvement, a unit of measure, and an object of control. "Minimise the time spent finding parking" qualifies. "Better parking" does not.
4. Survey a statistically valid sample. Customers rate every outcome twice: how important it is, and how satisfied they currently are. Ulwick recommends a sample in the range of 150 to 600, which is substantially larger than most qualitative research a startup runs.
5. Calculate opportunity scores. The standard formula is importance plus the gap between importance and satisfaction, where the gap is not allowed to go below zero. Outcomes that are important and poorly satisfied score highest. Those are underserved. Outcomes that are well satisfied relative to importance are overserved, and building there produces features nobody notices.
The output is a ranked list of opportunities with numbers attached. That is the entire reason to prefer this branch. It can be wrong, and the survey will show it.
The 86% claim
ODI is widely marketed with a success rate of 86% against an industry average of 17%.
That figure originates from Strategyn's own track-record study of its engagements. It has not been independently replicated. The comparison baseline comes from a different body of work with different definitions of success.
Treat it as a vendor claim. Under the evidence grading used across this series, ODI is a documented practitioner method — developed in the field, internally consistent, refined across many engagements, with evidence produced by the people who sell it.
That grade is not a reason to avoid it. It is the correct grade for most of the useful strategy toolkit. It is a reason not to cite 86% to a board.
The case for ODI rests on something simpler: it produces a falsifiable output. The narrative branch does not.
Which to use, and when
Use the narrative branch to reframe. Early exploration, when the team is stuck describing customers demographically, or when the competitive set feels obviously wrong. It is fast, cheap, and it changes how a room thinks in an afternoon.
Use the measurable branch to decide. When capital is about to be committed to a roadmap, a feature set, or a segment. When the question is not "what might customers want" but "which of these seven things should we build first."
Most teams run the narrative version and then make measurable-branch decisions with it. That is the error this article exists to name. A milkshake insight is a hypothesis. Treating it as a finding is how a team funds a confident guess.
Where ODI strains
It costs real money and time. A properly executed ODI study means interviews to generate outcome statements, then a survey of 150 to 600 respondents. For a business with a small addressable market or a short runway, that is a significant commitment.
It requires an existing job to study. ODI works best in categories where customers already accomplish the job somehow, even badly. For genuinely novel behaviour with no current workaround, there is no satisfaction level to measure.
Outcome statements are harder to write than they look. Most teams produce solution statements wearing outcome clothing. "Reduce the time to generate a report" is an outcome. "Add a one-click export" is a solution. Getting this wrong invalidates the survey before it runs.
It answers what, not whether. ODI ranks opportunities inside a job. It does not tell you whether the job is worth serving, whether the category can support your margins, or whether you should be in this market at all. Those are different questions with different frameworks.
Diagnostic: which version is the team actually using?
Six tests.
There is a written job statement containing no product or solution.
The job has been mapped into steps, and the steps are the customer's, not the company's process.
Outcome statements are written with a direction, a measure, and an object — and would survive a stranger reading them.
Importance and satisfaction have both been measured, on a sample large enough to mean something.
At least one opportunity was identified as overserved, and something was deprioritised because of it.
Someone can state what result would have proved the hypothesis wrong.
Passing one to three and failing four to six means the team is using the narrative branch. That is fine, provided nobody is treating the output as evidence.
What this produces
A ranked, quantified list of underserved customer outcomes, with a stated sample and a method that could have produced a different answer.
That is what makes it usable downstream. A strategic choice about where to play and how to win needs to rest on something more specific than a compelling story about a commuter. Positioning needs to name a value the customer already measures. Both of those requirements are satisfied by outcome data and not by narrative alone.
The story version teaches a team to see differently. The measurable version tells them what to do about it. Knowing which one is in the room is the whole discipline.
Frequently asked questions
What is the difference between Christensen's and Ulwick's Jobs to be Done?
Christensen's version is narrative: customers hire products to make progress in a situation. Ulwick's Outcome-Driven Innovation is quantitative: jobs are broken into steps, outcomes are rated by importance and satisfaction, and the gap produces a ranked opportunity score. Ulwick's version came first, in 1999.
Is the 95% product failure statistic accurate?
It is very difficult to trace to a primary study, and a 2022 academic paper found no empirical support for the widely repeated 90% figure. Category-level failure rates that do have stated methods vary considerably and are generally lower.
How large a sample does ODI need?
Ulwick recommends 150 to 600 respondents for the outcome survey. Smaller samples produce rankings that cannot be distinguished from noise, which defeats the purpose of using the quantitative branch at all.
How do you write a good outcome statement?
Three components: a direction of improvement, a unit of measure, and an object of control. "Minimise the time required to reconcile invoices" works. "Easier reconciliation" does not, because it cannot be rated meaningfully.
Can a small startup use ODI?
Yes, though the survey requirement is a real constraint. A workable adaptation is to run the job mapping and outcome statement work properly, then survey a smaller sample and treat the ranking as directional rather than decisive. What should not be adapted away is measuring satisfaction alongside importance.
Does Jobs to be Done replace segmentation?
It replaces attribute-based segmentation with outcome-based segmentation. Groups are formed by which outcomes they find important and underserved, rather than by demographics or firmographics. In practice this often cuts across conventional segments.
Where does Jobs to be Done sit in the Define sequence?
After structural analysis of the category and before the strategic choice. It establishes which customer outcomes are underserved, which is the raw material for deciding where to play and on what basis to win.
Sources
Ulwick, A., What Customers Want (2005) and Jobs to be Done: Theory to Practice (2016); Outcome-Driven Innovation, Harvard Business Review (2002); Strategyn track-record study (self-reported)
Christensen, C., The Innovator's Solution (2003), and subsequent Jobs Theory writing
Journal of Marketing and Consumer Behaviour in Emerging Markets (2022), on the empirical basis of new-product failure rate claims
Published critique of falsifiability in the narrative branch of Jobs to be Done
Structure your next phase
Zerologic runs consumer and behavioural research that produces ranked, measurable customer outcomes — the input a roadmap and a positioning decision can be built on.
Talk to us: partners@zerologic.io · zerologic.io



