DORA Metrics and SPACE: Measuring Engineering Productivity Without Gaming It

DORA Metrics and SPACE: Measuring Engineering Productivity Without Gaming It

DORA Metrics and SPACE: Measuring Engineering Productivity Without Gaming It

Every engineering leader eventually gets the same request: show me that the team is productive. The tempting answer is a dashboard of numbers pulled from Git, and that is precisely where most measurement programs go wrong. Within a quarter, developers learn which number is watched, and the number improves while the outcomes it was meant to represent do not. Goodhart’s law, in its common phrasing, says that when a measure becomes a target it ceases to be a good measure.

DORA metrics are the best-researched antidote available, and the SPACE framework is the best-argued reason not to rely on them alone. This post explains what the five current DORA metrics actually measure, how to compute them from GitHub or GitLab events and CI/CD logs with real SQL and Python, how SPACE and DevEx widen the lens, where gaming creeps in, and what changes when AI coding assistants inflate the volume of code your pipeline has to absorb. You will leave with definitions that survive an argument, a minimal data model, and a decision matrix that connects each metric movement to an action rather than a performance review.

What this covers: the current DORA definitions, the SPACE and DevEx frameworks, a working instrumentation pipeline, gaming and anti-patterns, AI-assistant effects, and a metric-to-action matrix.

Context and Background

The DORA research program began as the DevOps Research and Assessment team, was led in its early years by Nicole Forsgren, and is now run within Google Cloud. Its core contribution was statistical rather than rhetorical: it showed, through multi-year surveys of technology professionals, that a small set of delivery measures correlate with organizational performance. The original set was known as the four keys. The current model, documented on the official DORA metrics guide, has five metrics split into two groups, and that restructuring matters because it changed how teams should read the numbers.

The five metrics are grouped as follows. Throughput contains change lead time, deployment frequency, and failed deployment recovery time. Instability contains change fail rate and deployment rework rate. DORA’s own definitions are short and worth quoting in spirit: change lead time is the time for a change to go from committed to version control to deployed in production; deployment frequency is the number of deployments over a period or the time between them; failed deployment recovery time is the time to recover from a deployment that fails and requires immediate intervention; change fail rate is the ratio of deployments that require immediate intervention; and deployment rework rate is the ratio of deployments that are unplanned but happen as a result of an incident in production.

Two terminology shifts deserve attention. First, the older “mean time to restore” (MTTR) became “failed deployment recovery time”, narrowing the clock to recoveries from a deployment you shipped rather than any incident from any cause. The official page names the change but points to a separate history article for the reasoning, so treat claims about exactly why it happened with caution. Second, rework rate is the newest addition to the family, and DORA’s page does not explain its origin in detail. What the definition does make clear is that it counts deployments, not lines of code, which separates it from the code-churn metrics that vendors often sell under a similar name.

SPACE arrived from a different direction. Published in ACM Queue in 2021 by Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Brian Houck, and Jenna Butler, The SPACE of Developer Productivity argued that productivity cannot be reduced to one dimension or one number. It names five dimensions: satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow. The paper is explicit that DORA-style delivery measures live inside this larger map, mostly under performance, activity, and efficiency.

The practical consequence is a division of labor. DORA tells you how well a team’s delivery system converts changes into running software and how often that goes badly. SPACE tells you whether you are measuring enough of the picture to trust that reading. Platform teams, who own the paved road that other teams ride, need both: DORA to evaluate the road, SPACE to ask whether the people on it are healthy. If you are building the underlying delivery capability, our guide to AI-generated code guardrails in the CI/CD pipeline covers the controls that sit upstream of these metrics.

A note on evidence. DORA’s research is survey-based and correlational, which is a strength for breadth and a limit for causation. When a vendor says a given tool “improves DORA metrics by X percent”, ask which population, which period, and whether the comparison was against a matched control. This post cites only what the primary sources state and labels the rest.

The Five DORA Metrics as a Reference Architecture for Delivery Measurement

DORA metrics measure software delivery along two axes: throughput (how fast changes reach production, via change lead time, deployment frequency, and failed deployment recovery time) and instability (how often releases cause harm, via change fail rate and deployment rework rate). They are team-level indicators of a delivery system, not individual productivity scores.

DORA metrics reference architecture: commit to production with throughput and instability measurement points

Figure 1: Where each DORA metric is measured along the path from commit to recovery.

The diagram shows a single pipeline with three measurement intervals. Change lead time spans the commit-merged event to the successful production deployment. Deployment frequency counts production deployment events over a window. Failed deployment recovery time spans the moment a deployment is classified as failed to the moment service is restored. Change fail rate and rework rate are ratios computed over deployments, so they need a reliable classification of which deployments failed and which were unplanned reactions to an incident.

Throughput is about flow, instability is about harm

The most useful mental model is that throughput and instability are paired brakes on each other. If you push throughput with no regard for instability, you ship faster and break more. If you protect instability by adding approval gates, you slow throughput and typically increase batch size, which tends to raise the risk per release. Elite delivery is the combination, and the research basis for that claim is the repeated finding that the top performers do well across all metrics rather than trading one for another. The DORA page says exactly this, noting that top performers do well across all five and low performers do poorly.

This is why reporting a single metric is a design error. Deployment frequency alone rewards trivial deploys. Change fail rate alone rewards deploying rarely. The pair constrains behavior in a way neither does separately. Treat the five metrics as one instrument with five gauges, and never publish one gauge to a leadership audience without the others beside it.

Definitions that survive an argument

The hardest part of DORA is not the arithmetic; it is the definitions your organization must settle once and write down. Each requires a decision.

Change lead time starts at “committed to version control”. Does that mean the first commit on a branch, or the merge to the main line? DORA’s wording points to the commit, but many teams instrument from the merge or from pull request open time because those events are cleaner. Whichever you choose, state it, because first-commit lead time includes the developer’s own working time and merge-to-deploy lead time does not. The two can differ by days on teams with long-lived branches. Pick one and keep it for at least a year so trends mean something.

Deployment frequency must define “deployment” and “production”. A deployment is a release of a change to production that real users can reach; a canary at one percent counts if you define it so, but a deploy to staging does not. For multi-service organizations, decide whether the unit is a service, a team, or the product, since a team owning twelve microservices will deploy more often than one owning a monolith without being twelve times better.

Failed deployment recovery time needs a failure definition. The DORA phrasing, a deployment that fails and requires immediate intervention, translates in practice to a rollback, a hotfix, a forward fix shipped under incident rules, or a manual patch. Your criteria should be mechanical, such as “a deployment followed by a rollback event or an incident ticket linked to that deployment within a set window”, so classification does not become a negotiation.

Change fail rate is the failed deployments divided by total deployments for the same window. It inherits the failure definition, and it is only as good as your incident-to-deployment linkage.

Deployment rework rate is deployments that are unplanned and result from a production incident, divided by total deployments. It captures the cost of failure that change fail rate misses: a deployment can succeed technically and still trigger a chain of corrective releases. Teams that already log incident-linked deployments can compute it from the same data.

The official pages do not state numeric performance bands for the five-metric model, so be wary of any article that prints exact “elite” thresholds for rework rate. Earlier DORA reports did publish performance clusters for the four keys, and those figures varied from year to year, which is itself a reason to benchmark against your own trend line rather than against a table in a blog post.

Why platform teams should own the instrument, not the score

A platform team is the natural owner of the measurement pipeline because it already owns CI/CD, the deployment tooling, and the service catalog. But owning the instrument is different from owning the score. The score belongs to the stream-aligned teams being measured, and the platform team’s job is to make it cheap, consistent, and visible to them first. When the measurement is handed to teams as a self-service view before it is handed upward as a report, adoption goes up and defensive behavior goes down. We discuss the stakes of getting that handoff wrong in the gaming section below.

SPACE and DevEx: What the Five Metrics Leave Out

The SPACE framework says no single metric captures developer productivity and recommends selecting several metrics across at least three of five dimensions, including at least one perceptual measure such as a survey. DevEx, a companion framework, reduces the focus to three drivers: feedback loops, cognitive load, and flow state.

SPACE framework dimensions mapped to example measures including DORA delivery metrics

Figure 2: The five SPACE dimensions with example measures, showing where DORA metrics fit.

The mapping above is deliberately opinionated. Change fail rate and recovery time belong under performance, because they reflect the outcome of work and the reliability of what shipped. Deployment counts and pull request volume belong under activity, which is exactly why the SPACE authors warn against using activity measures alone to reward or penalize people: they can reflect overwork or a poor system as easily as productivity. Lead time and handoffs belong under efficiency and flow. Review wait time sits under communication and collaboration. And satisfaction, the one dimension DORA’s telemetry cannot see, comes from surveys.

The rules the SPACE paper actually gives

The ACM Queue article gives usable selection rules, and they are stricter than most dashboards honor. Use at least three metrics from at least three dimensions. Include at least one perceptual measure. Expect the metrics to pull against each other, because that tension produces a more balanced picture and discourages optimizing one area at the expense of the whole. Keep the set small, since too many metrics cause confusion. Recognize that metrics signal what an organization values, so choosing them changes behavior. Check for bias, and protect privacy by reporting anonymized, aggregate results at the team or group level.

That last rule is the one most often broken, and it is the one that matters most for avoiding gaming. The moment a metric attaches to an individual’s name in a performance review, the incentives to manipulate it become personal and immediate. Team-level aggregation does not eliminate gaming, but it moves the incentive from private advantage to shared outcome, which is easier to manage.

DevEx: the three levers behind the numbers

The DevEx paper by Abi Noda, Margaret-Anne Storey, Nicole Forsgren, and Michaela Greiler defines developer experience as how developers feel about, think about, and value their work, and names three dimensions. Feedback loops are how quickly tools and people respond to a developer’s action; slow builds, long test runs, and waiting for review all live here. Cognitive load is the mental effort a task demands, driven up by poor documentation, disorganized systems, and context switching. Flow state is the focused, energized immersion that interruptions and unplanned work destroy.

The useful insight for DORA practitioners is that the DevEx dimensions are the causes and the DORA metrics are the effects. A long lead time is a symptom; slow feedback loops are a plausible cause. A high change fail rate is a symptom; high cognitive load around an under-documented service is a plausible cause. The paper recommends combining workflow data, such as build times and review turnaround, with perceptual data from surveys, and breaking results down by team and persona. That recommendation alone separates a diagnostic program from a surveillance program.

A minimal balanced scorecard

For a platform team supporting several product teams, a scorecard that satisfies the SPACE rules can be small. From DORA, take change lead time, deployment frequency, change fail rate, and failed deployment recovery time (rework rate if your incident linkage is reliable). From the DevEx side, add one flow measure such as median pull request review wait, and one survey item such as “I can ship a small change to production without fighting the tooling”, scored quarterly. That is six or seven numbers from four SPACE dimensions, including a perceptual one, which is enough to satisfy the paper’s rules without drowning anyone.

The scorecard is a conversation starter, not a verdict. When lead time rises and the survey item falls together, you have corroborated evidence of a feedback-loop problem. When lead time rises and the survey item holds steady, you may be seeing a benign shift, such as a team deliberately taking on larger changes. The survey is what lets you tell the difference.

Instrumenting DORA Metrics from GitHub, GitLab, and CI Events

The cleanest way to compute DORA metrics is to treat delivery as an event stream: capture merge, build, deployment, rollback, and incident events into a raw table, then derive metrics with SQL models. Do not scrape dashboards, and do not compute from pull request timestamps alone, because deployment events are the ground truth for “in production”.

Instrumentation pipeline computing DORA metrics from Git, CI/CD, and incident events

Figure 3: An event-sourced pipeline from repository webhooks and CI events to team-level metric marts.

Three sources feed the pipeline. Repository webhooks from GitHub or GitLab provide push, merge, and pull request events. The CI/CD system provides build and deployment events, ideally including the commit SHA, environment, and status. The incident tool provides incident open and resolve timestamps and, critically, a link to the deployment that caused each one. Both GitHub and GitLab expose deployment objects through their APIs and webhooks, so if your pipeline records deployments through those native objects, you get a standard event shape. Check the current API documentation for exact payload field names before building, since they change between versions and this post does not reproduce them from memory.

The data model

Keep it small. Two tables carry most of the weight. A changes table has one row per merged change with change_id, team, service, first_commit_at, merged_at, and deployed_at. A deployments table has one row per production deployment with deployment_id, team, service, deployed_at, status, and is_failure and is_unplanned_rework flags that your classification step populates. A third incidents table carries incident_id, opened_at, resolved_at, and caused_by_deployment_id. The link from incident to deployment is the single most valuable field you will collect, and it is the one most teams lack at first.

How do you populate that link? The sturdiest approach is mechanical: a deployment is a failure if a rollback event or a hotfix deployment for the same service follows within a fixed window (for example 24 hours), or if an incident record cites it. Humans can then override the classification in a reviewed table, with the override reason retained. The window is a policy decision, so version it alongside the SQL.

Change lead time and deployment frequency in SQL

The following query computes weekly median change lead time and deployment frequency per team in a Postgres-compatible dialect. It uses the merge-to-deploy definition; swap merged_at for first_commit_at if you adopt the stricter commit-based definition.

WITH weekly AS (
  SELECT
    team,
    date_trunc('week', deployed_at) AS week,
    percentile_cont(0.5) WITHIN GROUP (
      ORDER BY EXTRACT(EPOCH FROM (deployed_at - merged_at))
    ) / 3600.0 AS median_lead_time_hours,
    COUNT(*) AS changes_shipped
  FROM changes
  WHERE deployed_at IS NOT NULL
  GROUP BY 1, 2
),
deploys AS (
  SELECT
    team,
    date_trunc('week', deployed_at) AS week,
    COUNT(*) AS deployments
  FROM deployments
  WHERE environment = 'production'
  GROUP BY 1, 2
)
SELECT w.team, w.week, w.median_lead_time_hours,
       w.changes_shipped, d.deployments
FROM weekly w
JOIN deploys d USING (team, week)
ORDER BY w.team, w.week;

Use the median or a percentile, never the mean, for lead time. A single change that sat for a month will drag a mean far from what most developers experience, and lead-time distributions are heavily right-skewed. Reporting the 50th and 90th percentiles together shows both the typical case and the tail where the pain lives.

Change fail rate, rework rate, and recovery time

These three come from the deployments and incidents tables.

SELECT
  team,
  date_trunc('month', deployed_at) AS month,
  COUNT(*) AS deployments,
  AVG(CASE WHEN is_failure THEN 1.0 ELSE 0.0 END) AS change_fail_rate,
  AVG(CASE WHEN is_unplanned_rework THEN 1.0 ELSE 0.0 END) AS rework_rate
FROM deployments
WHERE environment = 'production'
GROUP BY 1, 2;

SELECT
  d.team,
  date_trunc('month', d.deployed_at) AS month,
  percentile_cont(0.5) WITHIN GROUP (
    ORDER BY EXTRACT(EPOCH FROM (i.resolved_at - i.opened_at))
  ) / 60.0 AS median_recovery_minutes
FROM incidents i
JOIN deployments d ON d.deployment_id = i.caused_by_deployment_id
WHERE d.is_failure
GROUP BY 1, 2;

Monthly windows suit the instability metrics because failures are rare events. A team that deploys ten times a week and fails twice a quarter will show a noisy rate on a weekly chart: zero for eleven weeks and then ten or twenty percent in the twelfth. Aggregate wider, or show the rate with its denominator so readers see that one failure in eight deployments is not the same signal as one in eight hundred.

A Python collector for GitHub deployments

If you want the raw events without waiting for a vendor, a small collector against the GitHub REST API works. The sketch below pages through a repository’s deployments and their statuses. Verify the endpoint paths and pagination behavior against the current GitHub REST documentation before you rely on it; they are stable in my experience but this is code you should test, not trust.

import os, requests, datetime as dt

TOKEN = os.environ["GITHUB_TOKEN"]
HEADERS = {"Authorization": f"Bearer {TOKEN}",
           "Accept": "application/vnd.github+json"}

def get_all(url, params=None):
    while url:
        r = requests.get(url, headers=HEADERS, params=params, timeout=30)
        r.raise_for_status()
        yield from r.json()
        url = r.links.get("next", {}).get("url")
        params = None

def production_deployments(owner, repo):
    base = f"https://api.github.com/repos/{owner}/{repo}/deployments"
    for d in get_all(base, {"environment": "production", "per_page": 100}):
        statuses = list(get_all(f"{base}/{d['id']}/statuses"))
        final = statuses[0]["state"] if statuses else "unknown"
        yield {
            "deployment_id": d["id"],
            "sha": d["sha"],
            "created_at": d["created_at"],
            "final_state": final,   # success, failure, error, ...
        }

def lead_time_hours(commit_time_iso, deploy_time_iso):
    f = lambda s: dt.datetime.fromisoformat(s.replace("Z", "+00:00"))
    return (f(deploy_time_iso) - f(commit_time_iso)).total_seconds() / 3600

Two cautions about this approach. First, a deployment status of failure from the CI system means the pipeline failed, which is different from DORA’s failed deployment, one that reached production and needed intervention. Do not conflate a red pipeline with a failed deployment, or your change fail rate will measure CI flakiness. Second, deployments that bypass the API, such as a manual hotfix pushed from a laptop, will be invisible. Audit for them by comparing the deployment table against production artifact versions at least quarterly.

Making the pipeline trustworthy

Instrumentation fails quietly. Five checks keep it honest. Reconcile the deployment count against an independent source such as the artifact registry or Kubernetes rollout history. Alert when a team’s event volume drops to zero, since that usually means a broken webhook rather than a team that stopped shipping. Version the failure-classification rules. Publish the SQL so any team can reproduce its own number. And keep a data dictionary that states, in one sentence each, the exact definition of the five metrics your organization uses. A team that can reproduce its number from open SQL will argue about the engineering, not the arithmetic.

For organizations that run on managed platforms, some of this is available prepackaged: Google Cloud documents a Four Keys open-source project and several observability vendors ship DORA dashboards. Treat these as accelerators for the pipeline above, not replacements for your definitions, and confirm what each product counts as a deployment and a failure before comparing its numbers across teams.

Deeper Analysis: Gaming, Goodhart, and Reading the Metrics Honestly

Gaming happens when a metric becomes a target tied to reward or punishment. Teams then optimize the number rather than the outcome: splitting changes to shrink lead time, padding deployments to raise frequency, or narrowing incident definitions to lower change fail rate. The defense is balanced metrics, team-level reporting, and using the numbers as a diagnostic.

Gaming loop: pressure on a single metric produces dashboard gains without real outcome change until counter-metrics expose it

Figure 4: How targeting one metric produces dashboard improvement without real improvement, and the countermeasures that break the loop.

Common gaming patterns, one per metric

Each DORA metric has a characteristic manipulation, and knowing them lets you design guardrails before the incentive appears.

Lead time. The cheapest way to shorten change lead time is to shrink what counts as a change. A developer who splits a feature into twelve trivially small pull requests will show twelve short lead times, while the feature’s end-to-end delivery time is unchanged. Smaller batches are a genuine DORA-endorsed practice, so some of this is legitimate; the tell is when lead time improves but the time from idea to user value does not. Pair lead time with a cycle-time view at the epic or feature level to catch it.

Deployment frequency. Frequency is trivially inflated by deploying configuration changes, documentation, or no-op releases. If the metric counts every pipeline run to production, an empty commit is a deployment. Guard it by counting only deployments that carry at least one change from the changes table, and by reporting changes shipped per deployment alongside frequency.

Change fail rate. The standard manipulation is definitional: reclassify incidents as “known issues”, fix forward without opening an incident, or require a severity threshold so high that most failures fall below it. A related move is deploying less often in bigger batches, which lowers the failure count only by raising the stakes of each failure. Guard it with a mechanical failure definition (rollback or linked incident) that no one on the measured team can edit, and watch the denominator.

Recovery time. Closing incidents early, before service is genuinely restored, shortens the clock. Resolving by declaration rather than by verification is common when a team is graded on the number. Guard it by deriving the end timestamp from a health signal such as an SLO recovery, rather than from a human clicking “resolved”.

Rework rate. This one is hard to game upward but easy to hide: shipping a corrective change as part of a scheduled release rather than an unplanned one removes it from the numerator. Review how “unplanned” is determined, and prefer a rule tied to incident linkage over a team’s self-labeling.

Goodhart, Campbell, and the structural fix

Goodhart’s law and the related Campbell’s law describe the same dynamic: once a quantitative indicator is used for decisions with consequences, it is subject to corruption pressure and distorts the process it was meant to monitor. You cannot solve this by choosing a better metric, because any metric tied to consequences will bend. The structural fixes are about how the number is used.

First, report at the team level and never rank individuals on delivery metrics; the SPACE paper recommends anonymized, aggregate reporting for privacy and bias reasons that coincide with anti-gaming reasons. Second, never tie compensation or promotion to a DORA number. Third, publish trends and context, not league tables comparing teams with different architectures, regulatory burdens, and risk profiles. A firmware team shipping to field devices and a web team shipping a static site have legitimately different deployment frequencies, and a table that ranks them is measuring architecture, not effort. This issue is acute for industrial and cyber-physical systems, where releases are gated by safety processes; teams building autonomous decision engines on digital twins should expect a lower deployment frequency ceiling and judge themselves against their own trend line.

Fourth, use the metrics to ask questions in retrospectives rather than to deliver answers in quarterly business reviews. “Our lead time doubled in March; what changed?” invites an honest answer. “Your lead time is in the bottom quartile” invites a defensive one. Fifth, deliberately keep a survey channel open, because a perceptual measure is the hardest to fake at scale and the first to reveal that a number has detached from reality.

Detecting a number that has detached from reality

Gaming leaves fingerprints in the data. Watch for these patterns:

  • Deployment frequency rises sharply while changes shipped per deployment falls toward one or zero.
  • Lead time drops while the number of pull requests per feature climbs.
  • Change fail rate falls to near zero while the incident count or on-call page volume does not.
  • Recovery time shrinks while the number of reopened incidents grows.
  • The distribution of lead times becomes suspiciously uniform, which can indicate batch timestamp manipulation.

None proves bad faith. They are prompts for a conversation. The correct first response to an anomaly is curiosity about the system that produced it, not an accusation, since in most cases the team is responding rationally to an incentive someone else created.

AI Assistants and the Metrics: What Is Verified and What Is Not

AI coding assistants change what DORA metrics mean, because they inflate the volume of changes while review, testing, and deployment capacity stay fixed. The 2025 DORA report on AI-assisted software development, based on nearly 5,000 respondents as summarized by secondary reviewers, found AI adoption correlated with higher throughput and also with higher instability.

The primary source worth reading is the 2025 DORA research page, whose headline framing is that AI is an amplifier of an organization’s existing strengths and weaknesses, with the greatest returns coming from a strategic focus on the organizational system rather than the tools. A secondary summary reports that around 90 percent of respondents use AI at work and that the median user spends roughly two hours a day with it, that over 80 percent of users report productivity gains, and that about 30 percent have little to no trust in the code it generates. I did not independently verify those figures against the full report, so treat them as reported rather than confirmed.

Why the amplifier framing matters for measurement

If AI amplifies, then the same tool will move your metrics in opposite directions depending on the system around it. A team with strong automated testing, small batches, and fast review will convert faster code generation into faster, safe delivery. A team with a slow, manual review process will experience a pile-up: pull requests arrive faster than humans can read them, review becomes the bottleneck, and the pressure to approve quickly raises the change fail rate and rework rate downstream. A secondary source summarizing DORA’s work states that code review often becomes the bottleneck as AI accelerates code creation, which matches the queueing logic regardless of the exact numbers.

The DORA AI Capabilities Model, as described by the program, lists seven capabilities that amplify AI’s benefits: a clear and communicated AI stance, a healthy internal data ecosystem, internal data accessible to AI, strong version control practices, working in small batches, a user-centric focus, and high-quality internal platforms. Notice that several are classic delivery practices, and one is directly the platform team’s job. For platform teams the message is that the paved road is the amplifier’s quality control.

What to measure differently when AI is in the loop

Several adjustments follow from first principles, and I label them as engineering judgment rather than findings.

Add a review-queue metric. If changes arrive faster, median time from pull request open to first review and to approval becomes a leading indicator of both lead time and instability. Track the share of changes merged with no human comments, which is a crude proxy for rubber-stamping, and treat a rise as a warning rather than an accusation.

Watch rework more closely than before. Rework rate and change fail rate are the instability gauges most likely to move first if AI-generated changes pass review superficially. Compare a cohort of AI-assisted changes with unassisted ones only if you can label them reliably, and be careful, since developers choose when to use assistants and selection bias will distort a naive comparison.

Do not interpret a throughput increase as a productivity increase without the instability numbers beside it. The DORA finding of higher throughput alongside higher instability is exactly the pairing the five-metric model was built to expose. For deeper treatment of pipeline controls on generated code, including policy checks and provenance, see our guide to AI-generated code guardrails in CI/CD.

Finally, resist the temptation to measure AI usage itself as productivity. Prompt counts, accepted-suggestion rates, and lines generated are activity metrics in SPACE terms, and the framework’s warning about relying on activity alone applies with extra force: these numbers rise when developers are told to use the tool, regardless of whether outcomes improve.

From metric movement to action: a decision matrix

A metric that does not trigger a decision is decoration. The matrix below pairs common movements with the most plausible causes to investigate and a first action. These are hypotheses to test, drawn from the definitions and the SPACE and DevEx frameworks, not findings from a study.

Signal Plausible causes to check First action Companion measure
Change lead time rising Slow builds, review queue, long-lived branches, manual approval gates Break lead time into coding, review wait, and deploy wait stages PR review wait, survey item on feedback loops
Deployment frequency falling Batching releases, freeze windows, flaky pipeline Check changes per deployment and pipeline success rate Change size, build reliability
Change fail rate rising Weak test coverage, large batches, unreviewed generated code Inspect failed deployments for common root cause Rework rate, review depth
Recovery time rising Poor rollback path, missing runbooks, alert noise Rehearse rollback, automate it for the top services On-call load, SLO burn
Rework rate rising Fixes shipped as follow-ups, shallow acceptance testing Review incident-linked deployments for patterns Escaped defect source
All throughput up, instability up Speed gains outrunning verification, including AI-driven volume Add quality gates and review capacity before more speed Review queue, survey on trust
Metrics good, survey falling Overwork, burnout, hidden toil Talk to the team before celebrating Satisfaction and well-being

Read across rows rather than down them. The last two rows are the most important in practice, because they are the cases a single-metric dashboard hides entirely.

A worked example with illustrative numbers

The figures below are invented for illustration and are not benchmarks. A platform team supports a payments service. In a baseline month it makes 40 production deployments, of which 4 required rollback or hotfix, giving a change fail rate of 10 percent. Median change lead time is 30 hours, and median recovery time is 90 minutes. A new rollout of an AI assistant lands, and the next month shows 64 deployments, a median lead time of 22 hours, but 11 failures, a change fail rate of about 17 percent, with median recovery of 140 minutes.

A throughput-only dashboard reports a 60 percent increase in deployment frequency and a 27 percent drop in lead time, a triumph. The balanced view reports that failures nearly tripled in absolute terms and that recovery slowed. The review-queue metric then shows median time to approval fell from 9 hours to 3, and the share of changes approved with no comments rose sharply. The diagnosis is that review depth, not generation speed, was the real constraint. The action is to add automated checks and a review-load limit per reviewer, not to praise or blame anyone. That is the intended use of the instrument.

Trade-offs, Gotchas, and What Goes Wrong

DORA metrics have real limits, and an honest program states them. The first is scope: they measure delivery of changes to production, not whether the changes were worth building. A team can post excellent numbers while shipping features nobody uses. SPACE’s performance dimension, which includes impact such as adoption and customer satisfaction, exists to cover that gap, and a product-outcome measure belongs on the scorecard.

The second limit is context sensitivity. Regulated industries, embedded and firmware products, and on-premise software ship under constraints that cap deployment frequency regardless of engineering quality. Comparing a medical-device team’s frequency with a SaaS team’s is a category error. Compare each team with its own history and with its own constraints.

The third is data quality. Metrics built on weak incident-to-deployment linkage are fiction with decimals. If fewer than most of your failures are linked, the change fail rate understates reality and the recovery time sample is biased toward the incidents someone bothered to record. Improve the linkage before you interpret trends.

The fourth is the unit-of-analysis trap. A single deployment frequency for an organization hides a distribution: a few teams deploy constantly and many rarely. The average describes nobody. Report distributions and per-team trends.

The fifth is survey fatigue and bias. Perceptual measures are only valuable if people believe the results will not be used against them. A quarterly survey with visible follow-up actions sustains response rates; one that vanishes into a slide deck kills them. The SPACE paper also warns about cultural differences in survey responses, so compare a team with itself over time rather than across regions.

The sixth is metric proliferation. Once a pipeline exists, every stakeholder requests another cut. Resist it. The SPACE authors advise keeping the metric set small precisely because too many numbers cause confusion and discouragement. Add a metric only when you can name the decision it will inform, and retire one when it stops informing anything.

Finally, beware of tool-driven redefinition. Vendors implement DORA metrics with their own interpretations of deployment, failure, and lead time start. When you compare numbers from two products, or from a vendor against your SQL, expect discrepancies that are definitional. Document your canonical definition and reconcile the tool to it, not the reverse. The same discipline applies to audit trails more generally; see how identity and logging failures surface in practice in our write-up on AI agent audit logging and identity lessons from government site incidents.

Practical Recommendations

Start small and start with definitions. Write a one-page data dictionary for the five metrics before you write any SQL, and get the engineering leads to sign it, because a definition agreed in advance is much cheaper than a definition litigated after a bad number appears. Choose merge-to-deploy or commit-to-deploy for lead time and hold it for a year.

Build the event pipeline before the dashboard. Capture deployments from your CI/CD system with the commit SHA, link incidents to deployments, and store raw events so you can recompute when definitions change. Compute the metrics in versioned SQL and publish the code. Give each team its own view first, and let them see and challenge the numbers before anyone above them does.

Report distributions and pair metrics. Always show lead time with its 90th percentile, frequency with changes per deployment, and change fail rate with its denominator. Never display a throughput metric without an instability metric beside it. Add one flow measure and one survey item, so the scorecard covers at least three SPACE dimensions and includes a perceptual signal.

Decide up front how the numbers will not be used. Write down that DORA metrics are not inputs to individual performance reviews, compensation, or cross-team rankings. Review them in retrospectives as a diagnostic. If you are in an AI-assisted environment, add a review-queue metric and watch rework and change fail rate with extra care.

A short checklist to close:

  • Definitions for all five metrics written, signed, and versioned.
  • Deployment events captured with commit SHA and environment.
  • Incidents linked to deployments with a mechanical failure rule.
  • Medians and 90th percentiles used, never means alone.
  • At least three SPACE dimensions covered, including one survey item.
  • Team-level reporting only; no individual leaderboards.
  • Counter-metrics for each headline number.
  • Quarterly audit of deployment completeness and classification rules.

Frequently Asked Questions

What are the five DORA metrics?

The five DORA metrics are change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. The first three measure throughput and the last two measure instability. They describe how quickly changes reach production and how often releases cause harm, and DORA presents them as a set to be read together rather than as separate targets, since top performers do well across all five.

What is the difference between DORA metrics and the SPACE framework?

DORA metrics are a specific set of five software delivery measures. SPACE is a broader framework with five dimensions: satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow. SPACE advises picking several metrics across at least three dimensions, including a perceptual measure. DORA measures mostly fit inside SPACE’s performance, activity, and efficiency dimensions, while satisfaction and collaboration need surveys and other data.

How do you calculate deployment frequency and lead time for changes?

Deployment frequency is the count of production deployments over a period, for example per week per team. Lead time for changes is the elapsed time from a commit (or merge, if you define it so) to that change running in production, reported as a median and a 90th percentile. Both depend on reliable deployment events with commit identifiers, so capture them from your CI/CD system rather than inferring them from pull requests.

How can teams prevent gaming of DORA metrics?

Report at the team level, never rank individuals, and keep the metrics out of compensation and promotion decisions. Pair each headline metric with a counter-metric, such as changes per deployment beside frequency. Use mechanical failure definitions that the measured team cannot edit, and keep a survey channel open. Treat anomalies as prompts for curiosity about the system, not as accusations against people.

Do AI coding assistants improve DORA metrics?

The 2025 DORA report, as summarized by secondary sources, found AI adoption correlated with higher software delivery throughput and also with higher instability, and DORA frames AI as an amplifier of existing strengths and weaknesses. That means results depend on your review, testing, and platform practices. Track instability and review-queue metrics alongside throughput, and read the primary report for exact figures.

Is a high deployment frequency always good?

No. Frequency is easy to inflate with trivial or empty deployments, and it varies legitimately with architecture and regulation. It is meaningful only alongside change size, change fail rate, and recovery time. A team deploying constantly with rising failures is not performing well. Compare each team against its own trend and constraints rather than against an industry table.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *