TL;DR
A capable AI agent clears its pilot, and the decision to switch it on lands with one leader who will answer for everything it does next. What has to be true before it is allowed to touch real work?
What the paper develops
A capable AI agent, software that takes actions on its own, clears its pilot. The decision to put it into production lands on one leader's desk. The demo was clean and the model is genuinely capable. Approving it means handing the agent a license, a live workflow, and real customers and real data to act on. At that moment there is no settled answer to a plain question. The agent makes a judgment call, escalates the wrong case, or returns an output that looks finished and is wrong. Who in the organization owns that result?
The gap is not hypothetical. Gartner projects that more than 40 percent of agentic AI projects will be canceled by the end of 2027. The reasons it cites are rising costs, unclear value, and weak risk controls. Gartner does name immature models. Its headline reasons, though, point past the model to how agents are chosen and governed.
The Agent Operating-Model Canvas
My claim is narrow and testable. An agent belongs in production only after a named human can answer seven things about it.
1. Role and scope
What is it for, and where does it stop? Scope creep is how agents end up making decisions nobody designed them to make.
2. What it may decide, and what it may not
What may it decide alone, what needs a sign-off, what must go up the line? Left unwritten, authority defaults to the widest reading anyone can give it.
3. How it escalates and stops
A path for cases it should not handle, and a halt that exists before it is needed. One bad assumption at speed becomes a thousand bad outputs.
4. What has to challenge it
Before anyone trusts an output, what had to disagree with it and be overruled? Confident wrong answers are the expensive ones.
5. What a good result looks like
Define good before the agent produces anything. Otherwise review collapses into a check for plausibility, which is the check a fluent model defeats best.
6. What evidence it leaves
A trail that lets a human work out what it did and why. That is what turns standards like NIST's into daily practice.
7. Who owns the result
One named person who owns the outcomes. Not a committee. Not the vendor. If no name attaches, it does not ship.
The real ceiling
The canvas reads like a checklist you fill in once per agent. That invites an obvious next move: fill it in again and keep adding agents. The seventh element breaks that plan. Every agent needs one named human owner, and no human can answer for an unlimited number. So the canvas that governs one agent also caps how many you can run.
Human oversight is finite, and it runs out before model capability does. McKinsey reports that teams of two to five people can already supervise 50 to 100 agents. It also expects how far agents spread to be capped by how much oversight people can provide. Add agents without adding oversight and you do not grow output. You grow exposure. So fund the oversight as a standing cost, and make a completed canvas the gate an agent has to clear before it is funded.
What to do next
Before your next agent goes live, apply the simplest version of the test. Name the person who would stand up in a review and account for its decisions. If that name does not exist, or is really a group that adds up to no one, the deployment is premature, however capable the model is.
The white paper works through all seven canvas elements and the failure each one prevents. It adds a symptom-to-gap table for reading trouble back to its missing element, three governed agents from portfolio work, and the funding-gate move that keeps the count of live agents tied to the people available to answer for them.
What to do next
Before an agent goes live, name the human who answers for it. They should be able to answer seven questions: its role and scope, who decides what, how it escalates and stops, how its output gets challenged, what counts as acceptable, what evidence it must leave behind, and who is accountable. Write the answers on one page, the Agent Operating-Model Canvas, and make a completed canvas the gate an agent has to clear before it is funded.
Inside the white paper
- The Agent Operating-Model Canvas: seven things a named human must answer before an agent ships
- Why the real limit is human oversight, not model capability
- How to tie the number of live agents to the owners who can answer for them
Sources and notes
- Alex Singla, Alexander Sukharevsky, Bryce Hall, Lareina Yee, Michael Chui, and Tara Balakrishnan, "The state of AI in 2025: Agents, innovation, and transformation," McKinsey & Company, November 5, 2025. Verified July 10, 2026. mckinsey.com
- Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," press release, June 25, 2025. Verified July 10, 2026. gartner.com
- David Mallon, Brad Kreit, and Natasha Buckley, "Rethinking operating models for humans with agents," Deloitte Insights, April 2, 2026. Verified July 10, 2026. deloitte.com
- Alexander Sukharevsky, Alexis Krivkovich, Arne Gast, Arsen Storozhev, Dana Maor, Deepak Mahadevan, Lari Hämäläinen, and Sandra Durth, "The agentic organization: Contours of the next paradigm for the AI era," McKinsey & Company, September 26, 2025. Verified July 10, 2026. mckinsey.com
- National Institute of Standards and Technology, "AI Risk Management Framework Core," excerpt from AI RMF 1.0, 2023. Verified July 10, 2026. airc.nist.gov
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1, July 2024. Verified July 10, 2026. doi.org