DORA Metrics: What to Measure and How
DORA metrics come from the DevOps Research and Assessment program, now part of Google Cloud, which has surveyed tens of thousands of professionals since 2014. They measure how fast and how stably a team ships software. Used well they point at bottlenecks; used badly they become targets that get gamed. This guide defines each metric, shows where to get the data, and explains what to do with it.
The metrics
The long-standing "four keys" are two measures of speed and two of stability:
| Metric | Question | Category |
|---|---|---|
| Deployment frequency | How often do you deploy to production? | Throughput |
| Lead time for changes | How long from commit to running in production? | Throughput |
| Change failure rate | What share of deployments cause a failure needing remediation? | Stability |
| Time to restore service | How long to recover from a failed deployment or incident? | Stability |
The 2024 DORA report refined the stability side. It renamed time to restore as "failed deployment recovery time", and added "deployment rework rate", the share of deployments that are unplanned, made in response to a production issue. The research groups the metrics into throughput (lead time, deployment frequency, recovery time) and stability (change failure rate, rework rate). Check the current DORA site, dora.dev, for the latest definitions, since they have changed before.
How to measure each one from GitHub
You can compute all of them from data you already have.
Deployment frequency. Count successful production deployments per day or week. Source: the GitHub Deployments API, a workflow that deploys to a production environment, or release tags.
gh api "repos/OWNER/REPO/deployments?environment=production&per_page=100" \
--jq '[.[] | .created_at] | length'
For a fair figure, count deployments to production only, and count each service separately if you deploy services independently.
Lead time for changes. For each change, time from the first commit (or PR opened, if you need something simpler) to the moment it is deployed. The common approximation:
lead time = deployment time - first commit time of the PR
Report the median and the 90th percentile, not the mean; a few long-lived branches skew averages. Splitting it into phases shows where time goes: coding, waiting for review, CI, waiting to deploy. For most teams, review wait and CI time are the biggest two. See speeding up GitHub Actions.
Change failure rate. Failed deployments divided by total deployments. You need a definition of failure that you apply consistently: a rollback, a hotfix deployed within 24 hours, or an incident linked to a deployment. Tag incidents with the deploy that caused them, or label hotfix PRs and count them.
Failed deployment recovery time. For each failure, the time from deployment to resolution (the fix deployed or rollback done). Take it from your incident tracker. Use the median.
Deployment rework rate. Unplanned deployments divided by all deployments. Label hotfix and rollback deployments to count them.
Our DORA metrics calculator takes those four inputs and shows how you compare with the report's performance bands. The error budget calculator is a useful companion for the reliability side.
What the benchmarks say
DORA's 2024 report clusters teams into performance levels. For the top ("elite") cluster the report describes lead time of less than a day, on-demand deployment (multiple times a day), a change failure rate around 5 percent and recovery in under an hour. Lower clusters are slower on each axis. The report is based on survey responses, not instrumented data, and its own authors caution against treating the bands as targets. Use them for orientation only.
An important finding of the research: speed and stability are not a trade-off. Teams that deploy frequently tend also to have lower failure rates, because small changes are easier to review, test and roll back.
Make the measurement trustworthy
- Define your terms once and write them down: what counts as a production deployment, what counts as a failure, what the start of lead time is.
- Measure per team or service, not across the company; company averages hide everything.
- Use medians and distributions. A 4-day median with a 30-day 90th percentile tells you more than a 7-day mean.
- Automate collection. Hand-built spreadsheets go stale. Pull from the API weekly into a dashboard.
- Trend, not snapshots. Look at 8 to 12 weeks.
Misuse to avoid
DORA metrics describe a team's delivery system. They are not individual productivity measures, and ranking engineers or teams by them invites gaming. Splitting deployments into trivial ones inflates frequency; narrowing the definition of failure improves the failure rate on paper. Goodhart's law applies: when the measure is the target, it stops being a good measure.
Use the metrics to ask questions: why did lead time double in March? Where do PRs wait? What did we change before the failure rate rose?
Improve each metric
| To improve | Look at |
|---|---|
| Lead time | PR size, review latency (CODEOWNERS and review rules), CI duration, manual approvals, batch releases |
| Deployment frequency | Automated deploys, feature flags, trunk-based development, merge queues |
| Change failure rate | Test quality, flaky tests, staged rollouts, better review of risky paths |
| Recovery time | Rollback automation, observability, alerts, runbooks, on-call rotation (rotation generator) |
CI is the lever teams most often forget. If CI takes 25 minutes, no one pushes small changes. Use the cost of slow CI calculator to put a price on the waiting.
Other frameworks
DORA covers delivery. The SPACE framework adds satisfaction, collaboration and flow; developer experience surveys capture what metrics cannot. Combine a few quantitative measures with periodic surveys.
FAQ
How often should I review DORA metrics?
Monthly for the team, with weekly dashboards available. Quarterly reviews hide problems for too long.
Is a lower deployment frequency bad?
Not always. Regulated software or mobile apps with store review cycles will deploy less often. Compare with your own history and your own constraints.
Can I compute DORA metrics without a tool?
Yes, with the GitHub API, a deployments workflow and an incident log. Dedicated tools save effort when you have many repositories.
Where does CI speed show up?
Mostly in lead time, and indirectly in deployment frequency. Slow checks also encourage large, riskier batches.
Where compiler.dev fits
CI time is the part of lead time you can shorten by changing hardware rather than habits. compiler.dev's comparison mode runs your workflow on faster runners so you can measure the effect on your own pipeline before changing anything else.
Made by compiler.dev. Free tools · Pricing