AI operating governance · Field note

Measure AI by the workflow it changes

TL;DR

Usage shows who opened the tool. It does not show whether cycle time, quality, or rework improved.

What the paper develops

Most AI tools can report active users, sessions, and feature use within days. Those numbers answer a useful question: Did people open the tool? They do not answer the question leaders funded the program to solve: Did recurring work get faster, better, or less costly?

In a Gartner survey, 49 percent of respondents named estimating and showing AI value as the main barrier to adoption. A preliminary MIT NANDA study, reported by Fortune, found little or no measurable profit-and-loss effect for most pilots in its dataset. The study was preliminary and the report was secondary, so the figure is a warning, not a universal failure rate.

Measure the workflow in three layers

First, check whether AI reached the people and tasks in the target workflow. Tool-use data shows breadth and concentration. It does not show value by itself.

Second, check whether the workflow result moved. Use records the operation already keeps: cycle time, throughput, errors, rework, escalations, approval time, or another result leaders already care about. Pair speed or volume with quality so a visible gain does not hide later work.

Third, test whether AI is a credible reason for the change. Compare the workflow before and after, compare similar teams, sample quality, or check a manager's account against the records. State what the comparison can and cannot support.

Set the rules before seeing the answer

Name the workflow owner, data owner, and AI program owner. Write the baseline and follow-up periods, the measures, the comparison method, the level of tool use needed for a fair test, and the people who must confirm the result.

End with a choice

Every review should end with one of four choices: continue, hold, change, or stop. Record the choice, owner, reason, next check, and any signal that would reopen it.

Measure without watching every person

Use the least personal data that can answer the workflow question. Avoid keystroke logs, screen recordings, and a record of every prompt. Tell people what is measured, why it matters, who can see it, and which decision it supports.

Start with one recurring workflow

Choose one workflow with enough volume to measure. Ask: Which result changed, by how much, against what baseline, and with what evidence that AI contributed? If the program cannot answer, it may still be a useful experiment. It is not yet proof that the operation improved.

What to do next

Choose one recurring workflow. Set the baseline, measures, comparison, and four possible decisions before the team sees the result.

WORKFLOWCONTROL EVIDENCEHUMAN OWNER

Inside the white paper

  • Three evidence layers: tool use, workflow results, and why they changed
  • How to set the baseline and evidence rule before the test
  • How to measure results without watching every employee

Sources and notes

  1. Gartner, "Gartner Survey Finds Generative AI Is Now the Most Frequently Deployed AI Solution in Organizations," May 7, 2024 — Gartner reports that 49 percent of 644 respondents named estimating and showing AI value as the main adoption barrier.
  2. Gartner, "Gartner Identifies Four Emerging Challenges to Delivering Value from AI Safely and at Scale," October 21, 2024 — Gartner reports self-reported time savings from more than 5,000 digital workers and notes that the gains were uneven.
  3. Sheryl Estrada, "MIT report: 95% of generative AI pilots at companies are failing," Fortune, August 18, 2025, reporting findings from Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari, "The GenAI Divide: State of AI in Business 2025," MIT NANDA — Fortune reports preliminary MIT NANDA findings; the paper treats the 95 percent figure as a warning from one dataset, not a universal failure rate.
  4. Gartner, "Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025," July 29, 2024 — Gartner predicted that at least 30 percent of generative AI projects would be dropped after proof of concept by the end of 2025; this was a forecast.
  5. OpenAI, "The State of Enterprise AI: 2025 Report," 2025 — OpenAI reports large differences in specialized-GPT use between firms it labels frontier and the median enterprise; it is a vendor report about use, not results.
  6. Microsoft, "Copilot controls measurement and reporting," Microsoft Learn, updated February 25, 2026 — Microsoft separates Copilot readiness and adoption reporting from business-value and return-on-investment reporting.
  7. Rachel Schlund and Emily M. Zitek, "Algorithmic versus human surveillance leads to lower perceptions of autonomy and increased resistance," Communications Psychology 2, article 53, June 6, 2024 — Schlund and Zitek report lower autonomy and more resistance under algorithmic monitoring in four studies; two lab tasks also showed lower performance.
  8. Jared Spataro, "Our commitment to privacy in Microsoft Productivity Score," Microsoft 365 Blog, December 1, 2020 — Microsoft removed user names and associated actions from Productivity Score and shifted toward organization-level measures after criticism.
  9. Microsoft, "Copilot Business Impact Report," Microsoft Learn, updated February 17, 2026 — Microsoft's Business Impact report places business outcome data beside Copilot use data; the comparison does not by itself prove causation.
  10. National Institute of Standards and Technology, "AI Risk Management Framework Core," excerpt from Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023 — The NIST AI RMF Core calls for analysis, tests, comparison, monitoring, and a record of risks that cannot be measured.
  11. Gartner, "Gartner Survey Finds 45% of Organizations With High AI Maturity Keep AI Projects Operational for at Least Three Years," June 30, 2025 — Gartner reports associations among AI maturity, longer-running projects, business trust, and defined success measures; the survey does not prove causation.