TL;DR
People can badly misjudge how much AI speeds them up. Judge the tool on what it delivered, measured against a baseline taken before the tool arrived, and keep self-reported speed in a record of its own.
What the paper develops
Ask people how much faster AI made them. The answer is almost always a confident number. That number is a feeling, and a feeling is not operating evidence. When a program is funded, scaled, or defended on a self-reported speedup, the organization has let a feeling stand in for evidence.
The finding that should unsettle that habit is not that AI is slow. In a 2025 randomized controlled trial, the research nonprofit METR had sixteen experienced open-source developers complete 246 real tasks, with and without AI. Beforehand they expected a 24 percent speedup. Afterward they still believed they had gained 20 percent. The clock said they took 19 percent longer. The durable lesson is the gap between what capable people believed about their own speed and what the work recorded. That belief survived direct contact with the evidence. The trial does not set a universal effect for every task or tool. It measured expert developers, on codebases they knew well, with early-2025 tools. The gap is the point.
The same split appears across a whole software delivery process. Google's DORA program reported in 2024 that teams using AI scored higher on individual productivity, flow, and job satisfaction, while software delivery stability and throughput went down. People felt more productive while delivery became less stable, in the same dataset. Both findings are snapshots. Together they name the variable a governance standard has to account for. Felt speed and measured output are not the same signal, and the first runs ahead of the second.
People are poor witnesses to their own speed
METR's own February 2026 follow-up is blunt about the self-report problem. The developers’ estimates of their own speedup, it notes, "can be quite unreliable." Wider AI adoption "has made it more difficult to measure task-level productivity," because the comparison against unaided work is eroding. The feeling of fluency — an instant answer, a filled page — is a poor instrument for elapsed time and reworked output. A survey that asks employees how much AI helped is measuring enthusiasm and tool preference, not delivered help.
So the durable claim is not about speed at all. Self-reported speedup is a low-quality signal that systematically runs ahead of measured output. An organization scaling AI on that signal is scaling on sentiment. A portfolio does not fund sentiment on purpose. It does so by default, when no one built the alternative.
What to measure instead
The alternative is ordinary. First, capture a baseline before the tool changes the work. Then choose a work-output signal that reflects delivered value: cycle time, rework, error rate, or finished units meeting standard. Choose it so that using the tool more cannot move it by itself. Finally, review felt experience and delivered outcome in separate columns. When the columns agree, use can grow. When they diverge, the review has found the case a satisfaction survey would have counted as a win.
The hard part is deciding which numbers do not count. Adoption rate, active users, prompts per employee, and self-rated speed can all rise while the work stays the same. They belong on an adoption dashboard, not in the evidence a portfolio uses to deepen use. A contributed CIO column makes the distinction plainly. Governance machinery can sit above the work, while the value is created or lost inside the workflow. That is why a gain on one task may not reach the result the portfolio funded.
None of this slows adoption. It lets a leader say yes to deeper use on evidence, and redirect rather than abandon a tool whose value is leaking inside a workflow. Felt productivity is real and worth having. It is simply not the same thing as measured productivity, and only one of the two can tell a portfolio where its money went.
What to do next
Before scaling an AI-assisted workflow, require three things: a baseline from before the tool, an output measure that heavier tool use cannot inflate, and a review that keeps felt speed and delivered output in separate columns.
Inside the white paper
- Why people's sense of speed runs ahead of what the work actually shows
- A three-part standard: a baseline before the tool, an output measure heavier tool use cannot inflate, and two separate columns
- A one-page standard naming the seven things to fix before you scale a workflow
Sources and notes
- Joel Becker, Nate Rush, Beth Barnes, and David Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," METR, July 10, 2025 — the randomized trial where developers were 19% slower with AI yet believed they were 20% faster.
- METR, "We Are Changing Our Developer Productivity Experiment Design," February 24, 2026 — the follow-up noting self-report is unreliable and that wider adoption is eroding the ability to measure task-level productivity.
- Google Cloud DORA, "Accelerate State of DevOps Report 2024" — the system-level snapshot where AI adoption raised satisfaction while delivery stability and throughput fell.
- Nicole Forsgren and colleagues, "The SPACE of Developer Productivity," ACM Queue, 2021 — why productivity cannot be measured by a single metric or activity count.
- NIST, "AI Risk Management Framework" — the MEASURE function — measurement as an operating control that gives management decisions a traceable basis.
- Maria Korolov, "Why Is It So Hard to Measure the ROI of AI?" CIO, 2026 — explains why an AI ROI program needs a baseline and workflow-level evidence.
- Pete Johnson, "The AI ROI Gap Isn't a Model Problem. It's a Workflow Problem," CIO, 2026 — distinguishes governance activity from evidence that the task-level workflow improved.