You cannot improve delivery you do not measure, and for a long time software delivery had no agreed measure. DORA metrics changed that. They are four numbers, validated by years of research into thousands of engineering organizations, that together describe how well a team ships software. They matter because they are the rare engineering metrics that predict business outcomes, and because they give an engineering leader a way to talk about delivery in terms a CFO recognizes: speed and risk, tracked over time.

DORA stands for DevOps Research and Assessment, the research program the metrics come from. There are four of them, and the discipline is in using all four together, because each one on its own can be gamed or can hide a problem the others reveal.

The four metrics

The four split cleanly into two that measure speed and two that measure stability, which is the point: a good delivery process is fast and reliable at once, and measuring only one half hides the cost paid on the other.

Deployment frequency measures how often you release to production. It is the clearest signal of throughput, and higher is better, because frequent releases are small releases, which are easier to verify and safer to reverse.

Lead time for changes measures how long it takes a commit to reach production. It captures the whole path from written code to running code, so it exposes the delays that deployment frequency alone can miss, such as a fast pipeline sitting behind a slow manual approval.

Change failure rate measures the share of deployments that cause a failure needing remediation. It is the first of the two stability metrics, and it is what keeps deployment frequency honest, because shipping often is only an achievement if the releases work.

Mean time to recovery measures how long it takes to restore service after a failed change. It accepts that failures happen and asks the more useful question: when one does, how fast do you recover? A team that recovers in minutes can afford to move faster than one that recovers in days.

What good looks like

The DORA research groups teams into performance levels, from low to elite, by where they land on the four metrics. Elite performers deploy on demand, many times a day, carry a lead time measured in hours, keep change failure low, and recover from an incident in under an hour. Lower performers deploy weekly or monthly, measure lead time in weeks, and recover in days. The gap between the levels is large, and the research consistently finds that the elite pattern correlates with better organizational performance, not just happier engineers.

The useful part is not the label. It is that the four metrics locate a team honestly. A team that deploys daily but takes two days to recover is fast and fragile. A team that rarely fails but ships monthly is stable and slow. Only looking at all four shows which problem you actually have, which is why a real improvement program baselines every one of them before changing anything.

The mistakes that make the metrics useless

The metrics are easy to measure and easy to misuse. The most common error is tracking one and ignoring the rest, usually deployment frequency, which produces teams that ship constantly and break constantly. The second is turning the metrics into individual targets, which invites gaming: deployment frequency rises because releases are split into meaningless slices, and the number improves while delivery does not. The third is treating them as a scorecard rather than a diagnostic. The metrics are most valuable not as a grade but as a way to find the specific bottleneck slowing a team down, so the next improvement is aimed rather than guessed.

Used well, they are a feedback loop. Baseline the four, make a change to the pipeline or the process, and watch which metric moves. That is how delivery improvement stops being a matter of opinion.

How we use them

We treat DORA metrics as the frame for the whole engagement. Before any pipeline work, we baseline the four so there is a starting point, which is also what makes the work accountable to a number rather than a feeling. As we implement, the same metrics become the operating dashboard: deployment automation shows up as a lower change failure rate and faster recovery, and CI/CD pipeline work shows up as higher deployment frequency and shorter lead time. For Solarman Engineering Projects, the four moved together, with release frequency tripling while the change failure rate held below 5% and incident detection became 60% faster. Measuring them is what turned "delivery feels better" into a result an engineering leader could show.

Measurement is what makes delivery improvement real

DORA metrics are the difference between a DevOps program that can prove its value and one that just asserts it. They connect the pipeline work to an outcome leadership cares about, and they keep the work honest by making both speed and stability visible at once. Measurement is the fourth capability in a DevOps engagement, and it ties the others together. The full picture is in our guide to enterprise DevOps services.

Ready to measure your delivery?

If you are improving delivery without a baseline, you are guessing at whether it is working. We start every DevOps engagement by baselining your four DORA metrics and use them to aim and prove every change that follows. Reach out to talk through how your delivery measures up today.

Frequently asked questions

What are DORA metrics?

DORA metrics are four measures of software delivery performance from the DevOps Research and Assessment program: deployment frequency, lead time for changes, change failure rate, and mean time to recovery. Two measure speed and two measure stability, and together they are research-validated predictors of both engineering and organizational performance.

What are the four DORA metrics?

Deployment frequency (how often you release to production), lead time for changes (how long a commit takes to reach production), change failure rate (the share of deployments that cause a failure), and mean time to recovery (how long it takes to restore service after a failed change). The first two measure speed; the second two measure stability.

What is a good deployment frequency?

It depends on the team, but the DORA research finds that the highest performers deploy on demand, often many times a day, while lower performers deploy weekly or monthly. Frequent deployment is not the goal in itself; it is valuable because small, frequent releases are easier to verify and safer to reverse than large, infrequent ones.

How do you measure DORA metrics?

By instrumenting the delivery pipeline so deployment frequency, lead time, change failure rate, and recovery time are recorded as the pipeline runs, rather than gathered by hand. The important step is to baseline all four before making changes, so each improvement can be measured against a real starting point.

Are DORA metrics still relevant?

Yes. The four keys remain the standard measure of delivery performance, and the research now also emphasizes reliability alongside them, so many teams track an operational reliability measure as a companion. The core discipline is unchanged: measure speed and stability together, and use the numbers to find and fix bottlenecks rather than as a scoreboard.