Velocity Is Not Progress: Why Your AI Speed Metrics Are Measuring the Wrong Thing

Velocity Is Not Progress

Earlier this year, the COO of a financial services firm put his AI adoption dashboard on the screen with obvious pride. It was a good dashboard. Seat utilization at ninety-four percent. Weekly active users trending up for six straight months. Prompts per user, documents generated, estimated hours saved, all of it broken out by function with a confident green band across the top. Nine months of work and a little over four million dollars.

Then came a single question: what has the company decided differently because of any of this?

He thought about it for a while. Finally he offered that the finance team was producing the quarterly commentary in two days instead of five. Did anyone upstream use the commentary any differently now that it arrived three days earlier? No. The board packet still went out on the same Thursday.

That is the whole problem in one exchange. The firm had gotten measurably faster at a step that was not the constraint, and it had no instrument capable of noticing.

The Dashboard Measures What the Vendor Can See

There is a mechanical reason AI adoption dashboards look the way they do. The metrics on them are the metrics the tooling emits. Licenses activated, sessions per week, queries submitted, tokens consumed, self-reported minutes saved. Your vendor instruments what your vendor can observe, which is usage, and usage is a proxy for nothing in particular.

This is not a new failure. Every technology program in the last thirty years has reached for utilization as a stand-in for value, because utilization is available on day one and value is not available for eighteen months. What is new is the scale of the mismatch. A seat license is cheap, adoption curves move fast, and the resulting chart climbs steeply enough that it feels like evidence. Organizations are reading a steep line as a return.

The question that cuts through it is the one put to that COO, and it is worth asking out loud in your next steering committee: name a decision this organization now makes differently. Not a task that got faster. A decision. If the room cannot produce one after nine months and seven figures, the velocity on the dashboard is not progress. It is motion.

Acceleration Does Not Fix a Broken Workflow, It Industrializes It

Speed compounds value only when the underlying work was worth doing and the process around it was sound. When a workflow was ambiguous, duplicative, or poorly bounded before AI, acceleration does not clean it up. It produces the ambiguity faster and in greater volume.

The engineering data on this is the most useful evidence available, because software delivery is the one domain where both speed and quality are instrumented well enough to separate. DORA’s 2024 research found that a twenty-five percent increase in AI adoption was associated with a one and a half percent decrease in delivery throughput and a seven and two tenths percent decrease in delivery stability.1 Faster individual output, slower and less reliable delivery of the thing customers actually receive.

That result is not an argument against the tools. It is a precise illustration of what happens when you accelerate production inside a system whose constraint sits downstream of production. The bottleneck in most enterprise work is not generation. It is review, alignment, handoff, and decision. Pour more volume into a pipe that narrows at the review step and you have not increased flow. You have built a queue and called it productivity.

The Rework Is Real, and It Is Billed to Your Best People

The cost of accelerating a weak workflow does not appear as a line item. It appears as other people’s time.

Researchers at BetterUp Labs and Stanford’s Social Media Lab surveyed 1,150 U.S. desk workers and found that forty percent had received AI-generated work in the prior month that looked finished but was not usable, and that each instance consumed an average of one hour and fifty-six minutes to untangle and redo. Translated to salary, that is roughly $186 per employee per month, or more than nine million dollars a year at a ten thousand person organization.2

Two things about that number matter more than its size. The first is that none of it shows up on the adoption dashboard, because the dashboard counts the document that got produced, not the three hours someone spent determining it was wrong. The second is where the cost concentrates. Catching a plausible-looking error requires knowing what right looks like, which means the rework is absorbed almost entirely by your most experienced people. The senior engineer, the controller, the twenty-year operator who can tell at a glance that the number is off. Those are the people your organization can least afford to convert into full-time error catchers, and they are exactly who the invisible tax is levied on.

Leaders and Employees Are Describing Different Realities

The measurement gap has a perception gap sitting on top of it. The Upwork Research Institute surveyed 2,500 workers and executives and found that ninety-six percent of C-suite leaders expected AI to raise productivity, and eighty-one percent of leaders at companies that had deployed it reported that productivity had in fact risen. In the same study, more than three quarters of employees said the tools had decreased their productivity or added to their workload in at least one way, and forty-seven percent said they had no idea how to produce the gains their employers were expecting.3

Both groups are reporting honestly. They are reading different instruments. Leadership is reading the dashboard, which shows usage. Employees are reading their own calendars, which show the review burden the dashboard does not capture. When those two accounts diverge this far, the organization has a measurement problem, not a communication problem, and no amount of enablement messaging will close it.

Measure the Road, Not the Car

The replacement is not another dashboard. It is three questions asked with enough discipline that the answers become comparable quarter over quarter.

What decisions changed, and did they get better? Pick the five or six decisions that genuinely drive performance in a function: which accounts get pursued, how inventory gets positioned, which claims get escalated. For each one, record whether AI is now in the path, whether the decision is being made sooner, and whether the outcome improved. This is harder than counting prompts and it is the only measure that maps to the P&L.

How well does context move? Most AI output is weak because the model was given a thin slice of a situation that lives across six systems and four people’s heads. Context flow is measurable in plain terms: how long it takes a new analyst to assemble what they need to answer a standard question, how often a request bounces for missing information, how much of the input to a decision exists in a form a person or a model can actually retrieve. Organizations that improve this get compounding returns on every tool they buy afterward. Organizations that skip it keep buying faster cars.

What is the rework rate? Sample the work. Take thirty AI-assisted deliverables from last month and ask the recipients how many needed material rework, and roughly how long it took. You will have a number within a week. Track it. A rising volume of output with a rising rework rate is not a productivity gain, and the only way to know is to look.

Build the Road First

The organizations that will be meaningfully ahead in five years are not the ones with the highest seat utilization today. They are the ones that fixed the review step, made context retrievable, and got honest about which of their workflows deserved acceleration in the first place. That work is unglamorous, it does not produce a steep line in the first quarter, and it is the entire game.

Pull up your own AI dashboard this week and read it as a stranger would. Ask what each number would look like if your fastest-moving team were producing more work that nobody could use. If the dashboard would look exactly the same, you are not measuring progress. You are measuring the speedometer and calling it the map.

What decision does your organization make differently today than it did a year ago? If the answer takes more than a minute to find, start there.

References

  1. Google Cloud / DORA, Accelerate State of DevOps Report 2024. The 2024 DORA research found that a 25 percent increase in AI adoption was associated with a 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability, even as individual developers reported productivity gains. Roughly three quarters of respondents reported already using AI for some portion of their work.
  2. Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano, and Jeffrey T. Hancock, “AI-Generated ‘Workslop’ Is Destroying Productivity,” Harvard Business Review, September 2025, reporting research by BetterUp Labs and the Stanford Social Media Lab. A survey of 1,150 U.S. full-time desk workers conducted in August and September 2025 found that 40 percent had received AI-generated work in the previous month that appeared complete but lacked the substance to advance the task, with an average of 1 hour and 56 minutes spent addressing each instance. Based on respondent salaries, the authors estimate an invisible tax of roughly $186 per employee per month, or over $9 million annually at an organization of 10,000 people.
  3. Upwork Research Institute, “From Burnout to Balance: AI-Enhanced Work Models for the Future,” 2024. A survey of 2,500 C-suite executives, full-time employees, and freelancers across the United States, United Kingdom, Australia, and Canada found that 96 percent of C-suite leaders expected AI to increase productivity and 81 percent of leaders at AI-deploying companies reported that it had, while 77 percent of employees said AI tools had decreased their productivity or added to their workload in at least one way and 47 percent said they did not know how to achieve the productivity gains their employers expected.
  4. MIT Media Lab NANDA initiative, “The GenAI Divide: State of AI in Business 2025.” The report’s widely cited finding that approximately 95 percent of enterprise generative AI pilots produced no measurable P&L impact has circulated primarily through press coverage rather than a published methodology, so it is best treated as directional. Its central claim, that the gap is a workflow and learning problem rather than a model capability problem, is consistent with the better-documented findings above.

Jesse Jacoby

Jesse Jacoby is a recognized expert in business transformation and strategic change. His team at Emergent partners with Fortune 500 and middle market companies to deliver successful people and change programs. Jesse is also the editor of Emergent Journal and developer of Emergent AI Solutions. Contact Jesse at 303-883-5941 or jesse@emergentconsultants.com.


Leave a Reply

Your email address will not be published. Required fields are marked *


About us

Emergent Journal is a collection of business articles containing practical methods, tools, and tips for driving change and implementing business strategies from a people and change perspective. It is published by Emergent, a consulting firm headquartered in Denver and serving Fortune 500 clients across North America.

Learn More About EJ




Most Popular