TL;DR
Govern AI on evidence of the work — a baseline, a work-output signal, and a review that separates felt speed from delivered outcome — because self-reported speedup is not operating evidence.
What the paper develops
Ask people how much faster AI made them and you will get a confident number. The number is a feeling, and a feeling is not operating evidence. When a program is funded, scaled, or defended on a self-reported speedup, the organization has let a feeling stand in for evidence.
The finding that should unsettle that habit is not that AI is slow. In a 2025 randomized controlled trial, the research nonprofit METR had sixteen experienced developers complete 246 real tasks with and without AI. They expected a 24 percent speedup, and afterward still believed they had gained 20 percent. The clock said they took 19 percent longer. The durable lesson is the gap between what capable people believed about their own speed and what the work recorded. This study does not set a universal effect for every task or tool.
The same divergence appears at the level of a delivery system. Google's DORA program reported in 2024 that AI adoption raised individual productivity, flow, and job satisfaction while it lowered software delivery stability and throughput. People felt more productive, and the system delivered less stably, in the same dataset. Both findings are snapshots; together they establish the variable a governance standard has to account for — felt speed and measured output are not the same signal, and the first runs ahead of the second.
Why a feeling cannot govern
METR's own February 2026 follow-up is blunt about the self-report problem: those estimates 'can be quite unreliable,' and wider AI adoption 'has made it more difficult to measure task-level productivity' as the unaided comparison erodes. The internal sensation of fluency — an instant answer, a filled page — is a poor instrument for elapsed time and reworked output. A survey that asks employees how much AI helped is measuring enthusiasm and tool preference, not delivered help.
So the durable claim is not slower or faster. It is that self-reported speedup is a low-quality signal that systematically outruns measured output, and that an organization scaling AI on that signal is scaling on sentiment. A portfolio does not fund sentiment on purpose. It does so by default, when no one built the alternative.
A small standard that measures the work
The alternative is the ordinary discipline of measuring an operating change. First, capture a baseline before the tool changes the work. Then choose a work-output signal that reflects delivered value — cycle time, rework, error rate, or finished units meeting standard — and cannot rise merely because people use the tool more. Finally, review felt experience and delivered outcome in separate columns. When the columns agree, reliance can grow. When they diverge, the review has found the case a satisfaction survey would have counted as a win.
The hard part is deciding which numbers do not count as evidence. Adoption rate, active users, prompts per employee, and self-rated speed can all rise while the work stays the same; they belong on an adoption dashboard, not in the evidence a portfolio uses to deepen reliance. A CIO analysis makes the distinction plainly: governance machinery can sit above the work while the value is created or lost inside the workflow. That is why a local task gain may never reach the result the portfolio funded.
None of this slows adoption. It lets a leader say yes to deeper reliance on evidence, and redirect rather than abandon a tool whose value is leaking inside a workflow. Felt productivity is real and worth having. It is simply not the same thing as measured productivity, and only one of the two can tell a portfolio where its money went.
The operating move
Before scaling an AI-assisted workflow, require a pre-tool baseline, a work-output signal that using the tool more cannot inflate, and a review that keeps felt speed and delivered outcome in separate columns.
Inside the white paper
- Why self-reported speedup runs ahead of measured output — and stays there even after people live through the work
- A three-part measurement standard: a pre-tool baseline, a work-output signal activity cannot inflate, and a two-column separation review
- A one-page AI work-evidence standard that names the seven things to fix before a workflow is scaled
Sources and notes
- Joel Becker, Nate Rush, Beth Barnes, and David Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," METR, July 10, 2025 — the randomized trial where developers were 19% slower with AI yet believed they were 20% faster.
- METR, "We Are Changing Our Developer Productivity Experiment Design," February 24, 2026 — the follow-up noting self-report is unreliable and that wider adoption is eroding the ability to measure task-level productivity.
- Google Cloud DORA, "Accelerate State of DevOps Report 2024" — the system-level snapshot where AI adoption raised satisfaction while delivery stability and throughput fell.
- Nicole Forsgren and colleagues, "The SPACE of Developer Productivity," ACM Queue, 2021 — why productivity cannot be measured by a single metric or activity count.
- NIST, "AI Risk Management Framework" — the MEASURE function — measurement as an operating control that gives management decisions a traceable basis.
- Maria Korolov, "Why Is It So Hard to Measure the ROI of AI?" CIO, 2025 — explains why an AI ROI program needs a baseline and workflow-level evidence.
- Pete Johnson, "The AI ROI Gap Isn't a Model Problem. It's a Workflow Problem," CIO, 2025 — distinguishes governance activity from evidence that the task-level workflow improved.