Run the arithmetic on your own score

A team surveys 40 users. Sixteen say they would be very disappointed to lose the product. That is 40%, the threshold, and the result gets reported to the board as product-market fit achieved.

The 95% confidence interval on 16 out of 40 runs from roughly 26% to 55%.

The true figure could be well under the threshold or comfortably over it. The survey cannot distinguish between those cases. And that interval assumes a clean random sample, which almost no startup has, because the users who churned are not in the list.

The number was reported to three significant figures of confidence it does not possess. This is the most common failure in product-market fit measurement, and it is arithmetic rather than judgement.

Where the 40% came from

Sean Ellis introduced the test in 2009. One question, asked of active users: how would you feel if you could no longer use this product? Four options, of which one matters — very disappointed.

Ellis derived the 40% threshold from patterns he observed across roughly 100 startups, including work with Dropbox, LogMeIn, Eventbrite, and Lookout. Companies clearing 40% tended to grow. Companies below it tended to struggle.

Under the evidence grading used across this series, this is a documented practitioner method. It has not been independently validated. The specific number is a heuristic derived from pattern observation by one practitioner, not a finding with a published method and a control group.

That does not make it useless. It makes it a signal to be triangulated rather than a gate that authorises spending.

What the question actually measures

It measures dependence, which is why it works better than the alternatives most teams reach for.

Net Promoter Score asks about willingness to recommend — a social question, answered with reference to reputation and context. The Ellis question asks what the user would lose. That is closer to the thing that predicts retention, because it probes whether the product has become load-bearing in the user's workflow.

But it is still attitudinal. It captures what people say, not what they do. Stated preference and revealed preference diverge routinely, and the divergence is not random — it skews optimistic, because respondents are answering a question from a company they have a relationship with.

This is why the survey is a leading indicator that needs behavioural confirmation, never a substitute for it.

Four ways the number lies

Surveying only engaged users. The single most common error. Buffer ran the survey to a highly engaged segment and recorded 78% very disappointed. The broader user base did not retain at anything like that level. Survey your happiest cohort and the score reports on your happiest cohort.

The users who would answer "not disappointed" have mostly already left. They are structurally absent from the sample. Every PMF survey has this bias; the question is how much.

Too few responses. Ellis suggests 30 responses makes the result directionally useful and that confidence improves substantially at 100 or more. Rahul Vohra puts the directional threshold around 40. Below those levels the interval is wide enough that the threshold carries no information, as the opening arithmetic shows.

Surveying signups rather than users. The question only means something to someone who has experienced the core value. A workable filter is users who have used the product at least twice in the past fortnight. Survey everyone who ever registered and the score measures onboarding, not fit.

Treating the score as permanent. Fit is a relationship between a product and a market condition. Pricing changes, competitor launches, and audience expansion all move it. A score from eighteen months ago describes a market that may no longer exist.

The useful part is what Vohra did with it

Rahul Vohra's contribution at Superhuman is more valuable than the threshold itself, because it converts a score into a roadmap.

Segment the respondents into three groups, then treat them differently.

Very disappointed. These users have the fit. Ask what the main benefit is, in their words. Find what they have in common — role, company size, use case, how they arrived. That profile is the segment to build toward, and usually it is narrower than the one the business currently describes.

Somewhat disappointed. The productive group. Among them, isolate the ones who name the same main benefit as the very disappointed group, and ask what stops them depending on it. Those blockers are the roadmap.

Not disappointed. Ignore them. Building for users who do not want the product moves the score down, not up, because it dilutes the thing the first group values.

Superhuman ran this and moved from 22% to 58% by narrowing rather than broadening. The mechanism is worth stating plainly: the score went up because the audience got smaller and the product got more specific.

Confirm with behaviour

The survey should agree with the data. When it does not, the data wins.

Retention curve shape. Plot cohort retention over time. A curve that declines and then flattens indicates a group for whom the product became habitual. A curve that continues toward zero has no fit regardless of survey sentiment.

Cohort consistency. Newer cohorts should retain at least as well as older ones. If they do not, the fit existed for an early audience and is not travelling.

Organic share. What proportion of new users arrive without paid acquisition. Pull the business did not buy is among the more honest available signals, though the meaningful level varies enormously by category.

A strong survey score alongside a collapsing retention curve is not a contradiction to be resolved in the survey's favour. It usually means the sample was biased, or that users are attached to a promise the product does not yet keep.

What to do below 40%

Not: broaden the market. That is the instinct and it makes things worse.

The Superhuman sequence applies. Find the users who do have fit, describe them precisely, and narrow toward them. Fit is almost always found in a segment before it is found in a market, and the path runs through specificity.

If no segment shows fit, the problem is upstream of the product. That is a question about which customer outcomes are underserved and whether the strategic choice was right — not a question a survey can answer.

Diagnostic: is the score trustworthy?

Six tests.

  1. The respondent count is reported next to the percentage, every time it is quoted.

  2. Respondents had used the product enough to have formed a view, with a stated activity filter.

  3. The sample was not restricted to the most engaged cohort.

  4. The score has been checked against the retention curve for the same period.

  5. The very disappointed group has been profiled, and the profile is narrower than the current target market.

  6. Something on the roadmap changed because of the somewhat-disappointed responses.

Failing tests one to three means the number is not measuring what it claims. Failing four to six means it was measured and not used.

What this produces

A defensible answer to one question: should acquisition spend increase now, or should the product keep changing.

That is the decision the survey exists to inform, and it is worth getting right, because scaling acquisition ahead of fit is the most reliably expensive mistake available to an early-stage business. The score is one input to that decision. The retention curve is another. The segment profile is the third, and the most actionable.

Reported with its sample size, checked against behaviour, and used to narrow rather than to celebrate, the survey earns its place. Reported as a single number that cleared a threshold, it is a way of feeling certain.

Frequently asked questions

Is the 40% product-market fit threshold reliable?

It is a practitioner heuristic derived by Sean Ellis from patterns across roughly 100 startups. It has not been independently validated. Use it as a directional signal triangulated against retention data, not as a gate that authorises spending on its own.

How many responses does the PMF survey need?

Ellis suggests 30 for a directional read and 100 or more for confidence. At 40 responses, a 40% score carries a confidence interval running roughly from 26% to 55% — wide enough that the threshold tells you very little. Always report the respondent count beside the percentage.

Who should be surveyed?

Users who have experienced the core value, filtered by recent activity — for example, used the product at least twice in the past fortnight. Surveying signups measures onboarding. Surveying only the most engaged cohort inflates the score.

What is the difference between the PMF survey and NPS?

NPS asks about willingness to recommend, which is a social question. The PMF survey asks what the user would lose, which probes dependence. Dependence predicts retention more directly than recommendation intent does.

What should we do if the score is below 40%?

Narrow rather than broaden. Profile the users who said very disappointed, isolate the somewhat-disappointed users who name the same core benefit, and build toward removing their blockers. Superhuman moved from 22% to 58% by making the audience smaller and the product more specific.

Can the survey be run before launch?

No. The question requires a user who has experienced the product. Before launch, customer discovery interviews and structured pilots are the appropriate methods.

How often should product-market fit be re-measured?

When pricing changes, when the audience expands, or when a competitor launches. Fit describes a relationship between product and market condition, so it moves when either side moves.

Sources

  • Ellis, S., product-market fit survey (2009), derived from work across approximately 100 startups including Dropbox, LogMeIn, Eventbrite, and Lookout

  • Vohra, R., Superhuman product-market fit engine — segmentation and roadmap method

  • Buffer, published account of PMF survey results across engaged and broader user cohorts

  • Standard confidence interval calculation for a sample proportion

Structure your next phase

Zerologic runs validation and pilot programmes that produce evidence a leadership team can act on — before acquisition spend scales against an assumption.

Talk to us: partners@zerologic.io · zerologic.io