Skip to content
compiler.dev

DORA Metrics: What to Measure and How

DORA metrics come from the DevOps Research and Assessment program, now part of Google Cloud, which has surveyed tens of thousands of professionals since 2014. They measure how fast and how stably a team ships software. Used well they point at bottlenecks; used badly they become targets that get gamed. This guide defines each metric, shows where to get the data, and explains what to do with it.

The metrics

The long-standing "four keys" are two measures of speed and two of stability:

MetricQuestionCategory
Deployment frequencyHow often do you deploy to production?Throughput
Lead time for changesHow long from commit to running in production?Throughput
Change failure rateWhat share of deployments cause a failure needing remediation?Stability
Time to restore serviceHow long to recover from a failed deployment or incident?Stability

The 2024 DORA report refined the stability side. It renamed time to restore as "failed deployment recovery time", and added "deployment rework rate", the share of deployments that are unplanned, made in response to a production issue. The research groups the metrics into throughput (lead time, deployment frequency, recovery time) and stability (change failure rate, rework rate). Check the current DORA site, dora.dev, for the latest definitions, since they have changed before.

How to measure each one from GitHub

You can compute all of them from data you already have.

Deployment frequency. Count successful production deployments per day or week. Source: the GitHub Deployments API, a workflow that deploys to a production environment, or release tags.

gh api "repos/OWNER/REPO/deployments?environment=production&per_page=100" \
  --jq '[.[] | .created_at] | length'

For a fair figure, count deployments to production only, and count each service separately if you deploy services independently.

Lead time for changes. For each change, time from the first commit (or PR opened, if you need something simpler) to the moment it is deployed. The common approximation:

lead time = deployment time - first commit time of the PR

Report the median and the 90th percentile, not the mean; a few long-lived branches skew averages. Splitting it into phases shows where time goes: coding, waiting for review, CI, waiting to deploy. For most teams, review wait and CI time are the biggest two. See speeding up GitHub Actions.

Change failure rate. Failed deployments divided by total deployments. You need a definition of failure that you apply consistently: a rollback, a hotfix deployed within 24 hours, or an incident linked to a deployment. Tag incidents with the deploy that caused them, or label hotfix PRs and count them.

Failed deployment recovery time. For each failure, the time from deployment to resolution (the fix deployed or rollback done). Take it from your incident tracker. Use the median.

Deployment rework rate. Unplanned deployments divided by all deployments. Label hotfix and rollback deployments to count them.

Our DORA metrics calculator takes those four inputs and shows how you compare with the report's performance bands. The error budget calculator is a useful companion for the reliability side.

What the benchmarks say

DORA's 2024 report clusters teams into performance levels. For the top ("elite") cluster the report describes lead time of less than a day, on-demand deployment (multiple times a day), a change failure rate around 5 percent and recovery in under an hour. Lower clusters are slower on each axis. The report is based on survey responses, not instrumented data, and its own authors caution against treating the bands as targets. Use them for orientation only.

An important finding of the research: speed and stability are not a trade-off. Teams that deploy frequently tend also to have lower failure rates, because small changes are easier to review, test and roll back.

Make the measurement trustworthy

  • Define your terms once and write them down: what counts as a production deployment, what counts as a failure, what the start of lead time is.
  • Measure per team or service, not across the company; company averages hide everything.
  • Use medians and distributions. A 4-day median with a 30-day 90th percentile tells you more than a 7-day mean.
  • Automate collection. Hand-built spreadsheets go stale. Pull from the API weekly into a dashboard.
  • Trend, not snapshots. Look at 8 to 12 weeks.

Misuse to avoid

DORA metrics describe a team's delivery system. They are not individual productivity measures, and ranking engineers or teams by them invites gaming. Splitting deployments into trivial ones inflates frequency; narrowing the definition of failure improves the failure rate on paper. Goodhart's law applies: when the measure is the target, it stops being a good measure.

Use the metrics to ask questions: why did lead time double in March? Where do PRs wait? What did we change before the failure rate rose?

Improve each metric

To improveLook at
Lead timePR size, review latency (CODEOWNERS and review rules), CI duration, manual approvals, batch releases
Deployment frequencyAutomated deploys, feature flags, trunk-based development, merge queues
Change failure rateTest quality, flaky tests, staged rollouts, better review of risky paths
Recovery timeRollback automation, observability, alerts, runbooks, on-call rotation (rotation generator)

CI is the lever teams most often forget. If CI takes 25 minutes, no one pushes small changes. Use the cost of slow CI calculator to put a price on the waiting.

Other frameworks

DORA covers delivery. The SPACE framework adds satisfaction, collaboration and flow; developer experience surveys capture what metrics cannot. Combine a few quantitative measures with periodic surveys.

FAQ

How often should I review DORA metrics?

Monthly for the team, with weekly dashboards available. Quarterly reviews hide problems for too long.

Is a lower deployment frequency bad?

Not always. Regulated software or mobile apps with store review cycles will deploy less often. Compare with your own history and your own constraints.

Can I compute DORA metrics without a tool?

Yes, with the GitHub API, a deployments workflow and an incident log. Dedicated tools save effort when you have many repositories.

Where does CI speed show up?

Mostly in lead time, and indirectly in deployment frequency. Slow checks also encourage large, riskier batches.

Where compiler.dev fits

CI time is the part of lead time you can shorten by changing hardware rather than habits. compiler.dev's comparison mode runs your workflow on faster runners so you can measure the effect on your own pipeline before changing anything else.

Made by compiler.dev. Free tools · Pricing