AI operating governance · Field note

Every agent needs a human operating model

TL;DR

A capable agent clears its pilot, and the go-live decision lands on one leader who will answer for everything it does next. What has to be true before that agent is allowed to act on real work?

What the paper develops

A capable agent clears its pilot, and the decision to put it into production lands on one leader's desk. The demo was clean and the model is genuinely capable. Approving it means handing the agent a license, a live workflow, and real customers and real data to act on. And at that moment there is no settled answer to a plain question: when this agent makes a judgment call, escalates the wrong case, or returns an output that looks finished and is wrong, who in the organization owns the result?

The gap is real at scale. Gartner projects more than 40 percent of agentic AI projects will be canceled by the end of 2027, citing rising costs, unclear value, and weak risk controls. Those cancellations rarely trace back to model quality. They trace back to the operating model: agents dropped into work that was never redesigned to hold them accountable.

The Agent Operating-Model Canvas

The claim is narrow and testable: an agent belongs in production only after a named human can answer seven things about it.

  • Role and scope. What is it for, and where does it stop? Scope creep is how agents end up making decisions nobody designed them to make.
  • Decision rights. What may it decide alone, what needs sign-off, what must escalate? Left unwritten, authority defaults to the widest possible reading.
  • Escalation and stop. A path for cases it should not handle, and a halt that exists before it is needed. One bad assumption at speed becomes a thousand bad outputs.
  • Challenge protocol. Before an output is trusted, what had to disagree with it and be overruled? Confident wrong answers are the expensive ones.
  • Acceptance criteria. Define good before the agent produces anything, or review collapses into a plausibility check — the check a fluent model defeats best.
  • Evidence and logging. A trail that lets a human reconstruct what it did and why. This is what turns frameworks like NIST's into daily practice.
  • Accountable human owner. One named person who owns the outcomes. Not a committee. Not the vendor. If no name attaches, it does not ship.

The real ceiling

The canvas reads like a per-agent checklist, which invites an obvious next move: fill it in once for each agent and keep adding agents. The canvas quietly breaks that plan, because it raises the question most leaders underestimate — how many agents can one human actually own? Oversight is finite, and it runs out before model capability does. McKinsey observed small teams supervising 50 to 100 agents, and concluded that the ceiling on agentic scale will be human oversight capacity rather than model capability. Add agents without adding oversight and you do not scale output; you scale exposure. So fund the oversight as a standing cost, and make a completed canvas the gate an agent has to clear before it is funded.

The operating move

Before your next agent goes live, apply the simplest version of the test: name the person who would stand up in a review and account for its decisions. If that name does not exist, or is really a group that collectively means no one, the deployment is premature, however capable the model is. The white paper develops all seven canvas elements with the failure each one prevents, a symptom-to-gap table, three governed agents from portfolio work, and the funding-gate move that keeps the count of live agents tied to the human ownership available to answer for them.

The operating move

Before an agent goes live, require a named human who can answer seven things about it — its role and scope, decision rights, escalation and stop, challenge protocol, acceptance criteria, evidence duty, and accountable owner — and make that completed canvas the gate an agent must clear before it is funded.

WORKFLOWCONTROL EVIDENCEHUMAN OWNER

Inside the white paper

  • The seven-element Agent Operating-Model Canvas a named human must be able to answer before an agent ships
  • Why human oversight capacity, not model capability, is the binding constraint on how many agents you can safely run
  • The funding-gate move that keeps the count of live agents tied to the owners who can answer for them

Sources and notes

  1. Alex Singla, Alexander Sukharevsky, Bryce Hall, Lareina Yee, Michael Chui, and Tara Balakrishnan, "The state of AI in 2025: Agents, innovation, and transformation," McKinsey & Company, November 5, 2025. Verified July 10, 2026. mckinsey.com
  2. Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," press release, June 25, 2025. Verified July 10, 2026. gartner.com
  3. David Mallon, Brad Kreit, and Natasha Buckley, "Rethinking operating models for humans with agents," Deloitte Insights, April 2, 2026. Verified July 10, 2026. deloitte.com
  4. Alexander Sukharevsky, Alexis Krivkovich, Arne Gast, Arsen Storozhev, Dana Maor, Deepak Mahadevan, Lari Hämäläinen, and Sandra Durth, "The agentic organization: Contours of the next paradigm for the AI era," McKinsey & Company, September 26, 2025. Verified July 10, 2026. mckinsey.com
  5. National Institute of Standards and Technology, "AI Risk Management Framework Core," excerpt from AI RMF 1.0, 2023. Verified July 10, 2026. airc.nist.gov
  6. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1, July 2024. Verified July 10, 2026. doi.org