The framework's authors have moved on, and most of the industry has not
For roughly a decade, four metrics have been the standard language for software delivery performance: deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time.
In the 2025 DORA report, they appear in a footnote — described as one of several ways to measure software development. DORA changed both its measures and its clustering methodology that year.
Gene Kim, one of the original authors, went further, saying he had started calling the previous year's report and its findings "the DORA 2024 anomaly."
The framework has not been withdrawn, and the underlying research remains the strongest evidence base in this cluster. But anyone presenting the four metrics as current DORA doctrine is describing a position its own authors have visibly moved away from. That is worth knowing before building a measurement programme on it.
What the four metrics are, and why they worked
Throughput
Deployment frequency. How often the organisation releases to production.
Lead time for changes. From code committed to code running in production.
Stability
Change failure rate. The proportion of changes causing degraded service or requiring remediation.
Failed deployment recovery time. How long it takes to restore service after a failure.
The design is the insight. Two measures of speed and two of stability, read together. Each pair constrains the other, so gaming one shows up in the other.
The finding that mattered most
Speed and stability move together. They do not trade off.
This contradicts what most leadership teams believe walking in — that shipping faster necessarily means breaking more. The research consistently found the opposite: teams deploying more frequently also failed less and recovered faster, because the practices producing one produce the other. Small batches are easier to test, easier to diagnose, and easier to reverse.
The organisational finding attached to it: teams performing well on these measures were around twice as likely to meet or exceed their organisational goals, including profitability, market share, and productivity.
The evidence, with its limits
Empirically grounded, with caveats worth stating.
The scale is real. Annual State of DevOps surveys have run since 2014, synthesised in Accelerate in 2018 by Nicole Forsgren, Jez Humble, and Gene Kim, with cumulative respondents in the tens of thousands — the 2024 report drew on more than 39,000 professionals, and the 2025 report on nearly 5,000 new responses.
Four limits.
It is self-reported survey data. Respondents describe their own deployment frequency and failure rates. Nobody instrumented their pipelines.
Respondents are self-selected. People who answer a DevOps survey are not a random sample of software organisations.
The relationship is correlational. DORA applies statistical methods intended to support causal inference, but the design is observational. Teams that measure well may be well-run for reasons the survey does not capture.
DORA acknowledges priming. Its own 2025 methodological note observes that respondents may have been primed to some extent by the survey instrument.
None of this makes the research unusable. It makes it the best available evidence rather than settled fact, which is a different and more accurate description than it usually receives.
What happens when metrics become targets
Goodhart's law: when a measure becomes a target, it ceases to be a good measure.
Deployment frequency is unusually easy to game. Mandate more deploys and teams will split changes into smaller commits, push trivial updates, and deploy configuration changes that alter nothing. The number rises. Nothing improves.
This is why the pairing matters. Frequency rising while change failure rate also rises is not progress — it is a warning. A dashboard showing only throughput will report that warning as a success.
Two rules follow.
Never measure throughput without stability. Either half alone produces predictable distortion. Stability alone produces teams that avoid shipping.
Never use these to evaluate individuals. They are system-level and team-level indicators. Applied to people they measure the situation someone was placed in, and they guarantee gaming.
The 2025 findings are a live example
The most recent DORA research studied AI-assisted development, and the result illustrates exactly why paired measurement matters.
Around 90% of technology professionals reported using AI at work, up sharply on the previous year, and over 80% believed it increased their productivity. Roughly 30% reported little or no trust in AI-generated output.
The measured effect split. AI adoption correlated with higher software delivery throughput and, simultaneously, higher instability — more change failures, more rework, longer resolution cycles. Individual effectiveness, code quality, and team performance improved. Burnout and friction were unchanged.
DORA's framing is that AI acts as an amplifier rather than a solution. Teams with mature pipelines and automated testing gained on both dimensions. Teams without them generated code faster than their review and deployment systems could absorb.
An organisation measuring only throughput would have read this as unambiguous success. The instability is only visible because the framework insists on measuring both.
Note also what this describes structurally: code generation stopped being the constraint, and the constraint moved to review and testing. Adding capacity upstream of a bottleneck produces work in progress, not output.
Applying this outside software
The specific metrics are software-shaped. The structure transfers to any repeatable delivery process.
Throughput equivalents. How often does the business ship something to customers, and how long from decision to live. For a content operation, a campaign team, or a fulfilment process, both are measurable.
Stability equivalents. What proportion of things shipped require rework, correction, or apology, and how long does correction take.
The discipline is identical: measure both pairs, never one, and read them together. A marketing team publishing twice as often with a rising correction rate has the same problem as an engineering team deploying twice as often with a rising change failure rate.
Where measurement programmes go wrong
Only throughput is reported. The most common failure, because throughput is the flattering half.
Metrics become individual targets. Guarantees gaming and destroys the signal.
Benchmarks are applied without context. Elite benchmarks describe organisations with particular architectures and risk profiles. A regulated business with a monthly release window is not failing because it does not deploy daily.
The measurement replaces the diagnosis. Knowing lead time is nine days does not say where the nine days went. Metrics identify that a problem exists; finding it requires looking at the flow.
Improvement is pursued everywhere at once. Delivery performance has a binding constraint like any other system. Improving stages with spare capacity moves nothing.
Diagnostic: is measurement helping?
Six tests.
Both throughput and stability are reported, always together, in the same view.
No delivery metric is used in individual performance evaluation.
Targets are set relative to the organisation's own trend, not to published elite benchmarks.
When a metric moves, someone investigates why rather than reporting the movement.
The team can name the current bottleneck in the delivery flow.
At least one metric has moved in an unwelcome direction and been reported anyway.
Test six is the culture test. A measurement programme that only ever produces good news has stopped measuring.
What this produces
Visibility into whether the delivery system is improving or only appearing to.
That is the honest scope. These metrics do not tell a business what to build, whether customers want it, or whether the economics work. They tell it whether the machine that produces things is getting better or worse, which is a narrower question than it is often asked to answer.
Used well, that is enough. A leadership team that knows its lead time, its failure rate, and the direction both are moving can tell whether an investment in the delivery system paid off. Most cannot, which is why most such investments are justified afterwards rather than evaluated.
Frequently asked questions
What are the four DORA metrics?
Deployment frequency and lead time for changes, which measure throughput; change failure rate and failed deployment recovery time, which measure stability. They are designed to be read in pairs so that gaming one shows up in the other.
Are the DORA metrics still current?
Partially. The 2025 DORA report references them only in a footnote as one of several measurement approaches, and the team changed its measures and clustering methodology that year. The underlying research remains strong; presenting the four keys as current doctrine overstates it.
Do speed and stability trade off?
The research consistently found they do not. Teams deploying more frequently also failed less and recovered faster, because small batches are easier to test, diagnose, and reverse.
How reliable is the DORA research?
It is the strongest evidence base in software delivery measurement and it has real limits: self-reported, self-selected respondents, observational design, and DORA's own note that respondents may have been primed. Best available evidence rather than settled fact.
Should DORA metrics be used in performance reviews?
No. They are system-level and team-level indicators. Applied to individuals they measure the situation a person was placed in and reliably produce gaming.
Why did our change failure rate rise after adopting AI tools?
DORA's 2025 research found AI adoption correlating with both higher throughput and higher instability. Code is generated faster than review and deployment systems can absorb, which moves the constraint downstream rather than removing it.
Can these metrics work outside software?
The structure transfers. Measure how often you ship and how long from decision to live, alongside what proportion requires rework and how long correction takes. Always both pairs, never one.
Sources
Forsgren, N., Humble, J. and Kim, G., Accelerate (2018); DORA State of DevOps research, annual since 2014
DORA, State of AI-assisted Software Development (2025), including methodological notes on priming and the revised measurement approach
Published commentary on the 2025 report's treatment of the four key metrics
Goodhart, C., on measures becoming targets
Structure your next phase
Zerologic builds performance measurement frameworks that show whether a delivery system is improving — paired metrics, named owners, and diagnosis rather than dashboards.
Talk to us: partners@zerologic.io · zerologic.io



