Velocity Is Not Productivity: What to Measure Instead
Last updated September 2026
Engineering velocity vs productivity is a difference of kind, not degree. Velocity counts output inside one team’s process, so it cannot be compared across teams or trusted alone. Productivity is output plus quality, flow, and business impact, measured together. One controlled trial found developers 19% slower while believing they were 20% faster.
Engineering velocity vs productivity: what is the actual difference?
Velocity is a planning instrument. A team estimates work in its own units, completes some of it in a sprint, and the sum is velocity. That number is useful for one purpose: forecasting how much the same team, working the same way, on the same kind of work, can take on next sprint. It is unitless, self-referential, and inflatable by anyone who understands how it is computed.
Productivity is an outcome claim. It says the organization converted engineering effort into working software that customers use and that does not fall over. That claim cannot be made from a count of completed units, because the count is silent on three things that decide whether the work was worth doing: whether it works, whether the team can sustain the pace, and whether anyone needed it.
The confusion is expensive because the two respond to pressure in opposite directions. Push velocity and it goes up almost immediately, through larger estimates, thinner tests, smaller tickets, and deferred cleanup. Push productivity and nothing moves for a quarter. So the metric that is easy to move becomes the metric that gets managed, and the organization optimizes the instrument instead of the outcome.
Why does velocity feel like productivity when it is not?
Because perception of speed is unreliable, and there is now a measured gap to prove it. In Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR ran a randomized controlled trial with 16 experienced developers working on 246 real issues in repositories they already maintained. Issues were randomly assigned to allow or forbid AI tooling. The developers expected AI to make them 24% faster. They were in fact 19% slower on the AI-permitted issues. After finishing, having lived through the slowdown, they still believed they had been sped up by 20%.
That is a roughly 40-point swing between felt and measured throughput, in a small study of skilled engineers on their own code. It is the cleanest available evidence that self-reported and activity-shaped speed signals are not measurements. Any metric that a team can feel is a metric a team can be wrong about.
The industry-scale numbers point the same way. Google’s 2025 DORA State of AI-assisted Software Development report, based on responses from nearly 5,000 technology professionals, found AI adoption at 90%, a median of two hours a day working with AI, and more than 80% of respondents reporting that AI increased their productivity. In the same population, 30% said they trusted AI-generated code a little or not at all. Self-reported productivity went up while trust in the output did not. DORA’s own framing is that AI amplifies what an organization already has, which means a team with weak feedback loops now ships its weaknesses faster.
What should you measure instead of velocity?
Four metrics that pull against each other, so that gaming one shows up in another. This is the counterbalanced design behind both the SPACE framework and the DX Core 4, and the tension is the point, not a side effect.
| Dimension | Metric | What it catches | What it cannot see alone |
|---|---|---|---|
| Speed | Merged changes per engineer per week | Flow through the system, batch size discipline | Whether the changes were any good |
| Effectiveness | Developer experience index, survey based | Friction, wait states, lost focus time | Whether output actually moved |
| Quality | Change failure rate | Cost of speed, review and test adequacy | Whether the team is shipping at all |
| Impact | Percent of engineering time on new capability | How much capacity debt and toil consume | Whether the new capability was needed |
The reasoning is set out in The SPACE of Developer Productivity by Forsgren, Storey, Maddila, Zimmermann, Houck, and Butler, published in ACM Queue. Its central claim is blunt: productivity cannot be reduced to a single dimension, or a single metric. The paper names activity counts, commits and pull requests among them, as prone to error and never to be used in isolation, and recommends capturing metrics across at least three dimensions at once. Velocity is one metric in one dimension, which is exactly the failure mode the paper describes.
What are the 2026 benchmark numbers for these metrics?
This is the part usually missing. Most guidance on this topic tells you to replace velocity with cycle time and change failure rate, then gives you no number to compare yourself against, which leaves you measuring without a verdict. DX publishes percentile benchmarks for the DX Core 4 across its customer base, and these are the current figures for all companies.
| Metric | P50 | P75 | P90 |
|---|---|---|---|
| Merged changes per engineer per week | 3.8 | 4.67 | 5.53 |
| Developer experience index | 66 | 76 | 83 |
| Change failure rate | 4.02% | 3.46% | 3.05% |
| Percent of time on new capability | 57.83% | 61.72% | 65.06% |
Three things are worth noticing in that table. First, the spread on merged changes is narrow: the difference between a median organization and a 90th percentile one is under two changes per engineer per week, so a team claiming a doubling of output is almost certainly measuring something else. Second, change failure rate improves as the percentile rises, which is the counterbalance working: the fastest organizations are not buying speed with defects. Third, even at the 90th percentile roughly a third of engineering time is not going into new capability, so any plan that assumes 80% feature capacity is a plan built on a number nobody achieves. In its published benchmark set, DX also breaks these figures down by company size, and the larger the organization the lower the throughput, which is the honest context for any cross-company comparison.
How do you read engineering velocity vs productivity signals when they conflict?
Conflicting signals are the normal case, and the reading rules for them are rarely written down. Four patterns cover most of it.
- Speed up, quality flat, experience down. The team is absorbing the cost personally through overtime, skipped review, and interrupted focus. It holds for about a quarter, then the senior engineers leave. Treat a falling experience score as a leading indicator of the throughput drop that arrives two quarters later.
- Speed up, change failure rate up. You did not get faster, you moved work from before the merge to after it. Net productivity is usually negative once the incident and rework hours are counted.
- Speed flat, impact ratio falling. Capacity is being eaten by maintenance and toil. The fix is in the system, not the people, and it usually means the codebase has accumulated the kind of structural debt described in our guide to engineering delivery metrics that predict a missed deadline.
- Everything flat, stakeholders unhappy. The metrics are fine and the work is not wanted. This is the failure no engineering metric detects, and it is the argument for keeping a business impact dimension in the set at all.
The operational discipline that makes this readable is boring: collect on a fixed cadence, never attribute an individual metric to an individual person, and review the four numbers together in one 25-minute block rather than in four separate meetings. Our own delivery telemetry, PULSE, collects more than 400 data points a week per engagement for exactly this reason, because a weekly series is what makes a two-point move legible as a trend rather than noise. The mechanics of that review fit in a weekly engineering review of 25 minutes.
How do you measure productivity on a vendor or outsourced team?
Harder, and also the case most of this literature ignores. A vendor controls its own tooling, so you cannot instrument its laptops, and its incentive under a time-and-materials contract is hours, not outcomes. Three controls work without access to their systems.
Measure at the boundary. Change failure rate on what lands in your repository, lead time from accepted scope to production, and the share of delivered work that needed rework within 30 days are all computable from your own version control and incident data. Ask for the experience dimension contractually, as a quarterly survey of the engineers actually on your account, since a vendor rotating staff through your codebase will show it there first. Then check continuity directly: ask how many of the engineers who started are still on the account, because retention is the one input you cannot fake with process. It is why we publish ours at 96%, with more than half of our engineers joining through internal referral, and why a Delivery Pod is priced from $15,000 a month with 30 days notice rather than as a headcount lease. If the codebase is already in a state where none of these numbers can be computed, that is a measurement problem before it is a delivery problem, and an AI Code Rescue exists to produce the baseline.
Book a Code Review
If your velocity chart is healthy and your release still slipped, the chart is not lying, it is answering a different question than the one you asked. A code review gives you the four numbers above computed from your actual repository and incident history, a read on which of the conflict patterns you are in, and a ranked list of what to fix first, done by a senior engineer rather than a dashboard. The $4,950 Rescue Audit is credited toward remediation if you proceed. Book a Code Review and measure the outcome instead of the instrument.