The denominator is the problem
RICE scores an initiative as reach multiplied by impact multiplied by confidence, divided by effort.
Look at what is in the denominator. Effort is an estimate of how long work will take — and estimation is the single most thoroughly discredited input available to a business.
The evidence is not ambiguous. Asked to predict thesis completion, students averaged 33.9 days. Asked for a worst case — everything going as badly as possible — they said 48.6. The actual average was 55.5 days. Reality was worse than the worst case they could imagine, and being reminded of past prediction failures did not improve subsequent estimates.
So a RICE score divides three judgements by the number we know to be least reliable, and produces a figure to one decimal place.
That figure then gets sorted, and the sort gets called a roadmap.
The arithmetic does not survive its inputs
Three of the four RICE inputs are judgements on invented scales. Impact is commonly scored 0.25, 0.5, 1, 2, or 3. Confidence is a percentage picked by the person filling the form.
Multiplying and dividing values like these produces a number that looks like a measurement. It is not. Ordinal judgements combined arithmetically yield something with the appearance of a ratio scale and none of its properties. The difference between 12.4 and 11.8 is not a difference in anything.
The test that settles it takes five minutes. Take the current scored list. Change one input on the second-ranked item by 20% — the kind of adjustment nobody would argue about in a meeting. Re-sort.
If the ranking changes, the framework did not determine the priority. A judgement did, and the arithmetic distributed responsibility for it across a spreadsheet.
This is what false precision means practically: not that the numbers are wrong, but that they are reported with a confidence the method cannot support, and that confidence transfers to the decision.
Then there is the finding that complicates all of it
Here is where the easy argument breaks down, and where this article gets more interesting.
Robyn Dawes published The Robust Beauty of Improper Linear Models in Decision Making in American Psychologist in 1979. Building on Paul Meehl's work from 1954, he reviewed evidence that simple linear models outperform expert intuition at predicting numerical outcomes.
The striking part is what happens with bad weights. Dawes showed that improper linear models — weights set by intuition, weights set equal, even weights set at random — still outperformed clinical judgement. Random linear models performed roughly as well as models built to mimic the judges themselves.
And a search of the literature failed to turn up studies where clinical judgement beat statistical prediction on the same coded inputs.
The explanation is the sentence worth carrying out of this article. People, especially experts, are much better at selecting and coding information than at integrating it. The linear model cannot decide what to look for. That is exactly the expertise humans have. What humans do badly is combine several considerations consistently, and that is precisely what a model does well.
So both things are true
The precision is fake. RICE cannot support one decimal place, and treating small score differences as meaningful is an error.
The consistency is real. A scoring framework applied the same way across every item removes a category of noise that intuitive prioritisation does not. That effect is well evidenced and has been for decades.
The failure mode is not using scoring frameworks. It is using them as answer generators rather than as consistency devices — and then treating the output as the reason a decision was made.
A team that scores ten items and takes the top three has done something defensible. A team that scores ten items and defends the ordering of items four through seven has confused the tool with a measurement.
Five rules
1. Use the framework to choose factors, not to produce answers. Dawes's finding puts human expertise in selecting what matters. Deciding that reach, impact, confidence, and effort are the right considerations is the valuable judgement. The arithmetic afterwards is bookkeeping.
2. Run the sensitivity test every time. Move one input 20%. If the top of the list reorders, say so out loud before proceeding. The framework has not resolved the question and pretending otherwise is where the damage happens.
3. Score into bands, not ranks. Three tiers — clearly worth doing, genuinely unclear, clearly not — is what the input quality supports. A ranked list of eighteen items claims discrimination the method cannot deliver.
4. Keep effort out of the denominator where possible. Given how unreliable estimation is, a fixed-time approach that asks what an outcome is worth avoids the problem rather than encoding it. Appetite is a decision; effort is a forecast.
5. Record the disagreement. When two people score the same item differently, that gap is information about a genuine uncertainty. Averaging it away destroys the most useful thing the exercise produced.
The pattern across this cluster
Nine articles on execution frameworks, and the same failure recurred in almost every one.
Theory of Constraints works when the constraint is identified and subordinated to. It fails when organisations skip to buying capacity, because that step feels like action.
Team Topologies rests on Conway's Law, which is well evidenced, and on Dunbar's numbers, which have been challenged to the point where the confidence intervals run from 4 to 520 people.
Working Backwards works when a document can fail. If no PR-FAQ has ever killed an idea, the process produces justification.
The Opportunity Solution Tree works when opportunities come from customers. Populated from internal preference, it formalises the politics it was meant to interrupt.
Shape Up works when the circuit breaker is used. The first extension ends the method.
Service blueprinting works when it maps what happens rather than what is supposed to happen.
Design systems fail on adoption and ownership, not on design — and 42% of surveyed teams felt the original build created debt.
Delivery metrics work in pairs. Throughput alone is the flattering half, and DORA's own 2025 report moved the four keys to a footnote.
One pattern. Every framework in this cluster fails the same way: it stops being a test and becomes a record of a decision already made.
The Define cluster reached the same conclusion from the other end, where the evidence problem was inflated statistics rather than inflated precision. Both are versions of the same thing — a method producing more confidence than its inputs justify.
What actually improves decisions
Not better frameworks. The research on this is consistent and slightly deflating: across 1,048 major strategic decisions, process mattered more than analysis by a factor of six.
Which means the interventions that work are structural and cheap.
State the disconfirming result before the analysis. What would we see if this is wrong.
Name the evidence grade when a number enters the room. Empirically grounded, practitioner method, or thinly evidenced.
Require an exclusion. A prioritisation that rules nothing out has not prioritised.
Give dissent somewhere to go. Analysis that is never challenged does not improve.
Cap the cost of being wrong. Fixed cycles and stage gates do more for outcomes than better estimates ever will, because they change what a mistake costs rather than trying to prevent it.
Diagnostic: is this prioritisation or is it justification?
Seven tests.
The sensitivity test has been run, and the result is known.
Output is bands rather than a precise ranking.
Something has been excluded by name, not merely ranked low.
Scoring disagreements are recorded rather than averaged.
The factors being scored were chosen deliberately, not inherited from a template.
Someone can name an item the framework demoted that the team then chose to build anyway, and why.
The list is different from what the team would have chosen without it.
Test seven is the one worth sitting with. A framework that reproduces the team's prior intuition has added consistency, which is worth something, and has not added information.
Closing the Build stage
Across two clusters and twenty articles, the frameworks were rarely the problem. The problem was the confidence attached to them.
The Build stage closes when a business can produce what its strategy requires, repeatedly, and can see where it is slow. Frameworks help reach that state efficiently. None of them substitute for the process around the decision — and the evidence says that process is where most of the improvement actually lives.
A decision specific enough to act on, cheap enough to be wrong about, and honest enough about its own inputs to be revised. That is the whole of it.
Frequently asked questions
Is RICE a good prioritisation framework?
It is useful as a consistency device and unreliable as a source of precise rankings. Its denominator is an effort estimate, and estimation is among the least reliable inputs available. Use it to form bands, not to order a list.
What is false precision in prioritisation?
Reporting a number with more confidence than the method's inputs can support. Multiplying ordinal judgements produces a figure that looks like a measurement without having the properties of one.
How do I test whether my scoring model is meaningful?
Change one input on a highly ranked item by 20% and re-sort. If the ordering changes, the framework did not determine the priority — a judgement did, and the arithmetic obscured whose.
Should we abandon scoring frameworks entirely?
No. Dawes showed in 1979 that simple linear models outperform expert intuition, and that even models with intuitive, equal, or random weights beat clinical judgement. Consistency is the benefit. Precision is not.
Where does human expertise actually help?
In selecting and coding what matters, not in integrating it. Dawes's conclusion was that people, especially experts, are far better at deciding what to look for than at combining several considerations consistently. The model handles the combining.
What is the single biggest improvement to a prioritisation process?
Requiring an explicit exclusion. A prioritisation that ranks everything and rules out nothing has produced an ordering, not a decision.
If frameworks are limited, what actually improves decisions?
Process. Research across 1,048 major strategic decisions found process mattered six times more than analysis. Stating the disconfirming result in advance, grading evidence out loud, and capping the cost of being wrong all do more than a better scoring model.
Sources
Dawes, R.M., The Robust Beauty of Improper Linear Models in Decision Making, American Psychologist 34(7), 1979; Meehl, P., Clinical Versus Statistical Prediction (1954)
Buehler, R., Griffin, D. and Ross, M., Exploring the "Planning Fallacy", Journal of Personality and Social Psychology 67(3), 1994
Lovallo, D. and Sibony, O., The Case for Behavioral Strategy, McKinsey Quarterly (2010)
Evidence qualifications drawn from the preceding nine articles in this series, with primary sources cited in each
Structure your next phase
Zerologic works with leadership teams to build the operating system a strategy requires — and to keep the confidence attached to a decision proportionate to what supports it.
Talk to us: partners@zerologic.io · zerologic.io



