TL;DR
Every AI agent adds work that a named person must review, own, and explain. Oversight may be the first limit a company hits, and a plan can easily leave it uncounted.
What the paper develops
You are told to add more AI agents, find more uses, and get more work done. The plan may count model cost and technical skill. It may not ask a harder question: How many AI decisions can your people still review, explain, and stand behind?
Every agent adds output. It also adds work for people who check important results, handle exceptions, and answer when the AI is wrong. The claim here is narrow: human oversight can become the first limit a company hits as it adds agents, and a plan can easily leave it uncounted. The evidence does not show that it always does, and it gives no number a company can borrow. If the owner's review hours stay well within the time the owner has as agents or cases are added to a workflow, oversight is not the limit for that workflow.
McKinsey's authors expect human oversight to cap how widely companies adopt AI agents. They also say that, in their experience, a team of two to five people can already supervise 50 to 100 specialized agents. That is a forecast and an early observation with no stated method, not a safe benchmark. It still points to the question leaders need to answer: Where is the limit for this team and this work?
Estimate how much one person can supervise
The span of supervision is how much AI work one accountable person can genuinely own, counted in one unit per workflow, such as agents, cases a week, or decisions. The span gets smaller when errors are costly, hard to see, or hard to reverse. It can be larger when errors are easy to spot and fix. If no one can estimate it yet because the workflow is new, start at the narrowest span: one owner, one workflow, a person checking every case or a large sample, and no new agents or cases until the records exist. Widen only after four weeks of records show room.
Use a weekly test. Can the named owner explain what the agent did, what was checked, what failed, and what changed? A no can mean the owner is out of time, or that the records are too thin to answer from. Compare the owner's hours with what the owner expected before narrowing the span. If the hours are fine, fix the records first.
Review by risk
Oversight does not mean reading every output. Low-risk work may need sampling. High-risk work may need a person to approve each case. New or uncertain work may need tighter review until the team knows its typical failures. Write down what is sampled, what always gets checked, what triggers a deeper review, and who can stop the workflow.
Use tools, but keep a human owner
Critic agents, guardrails, monitoring, tests, and incident response can help a small team supervise more work. They can find common errors and send unusual cases to people. They still need their own tests and monitoring. A named person must be able to change or stop them.
Treat review time as shared capacity
Every live AI workflow draws on a shared group of reviewers and experts. Each funding request should name the owner, expected review load, type of review, experts needed for hard cases, and signals that would pause or stop the work. When that pool is full, move attention away from stable, low-risk work before adding more demand.
Start with one workflow
The owner writes a one-page record with the portfolio owner who funds the work. As a planning estimate, not a measured figure, expect a few hours to write it and a few minutes a week to log review time. Each week, log the owner's review and follow-up hours beside the live volume behind them. After four weeks, check whether exceptions reached the right person and whether the owner can explain what changed. If the owner had room, widen the span one step and compare hours before and after. If time ran over, narrow it before another agent goes live.
Better models and automated checks may shrink the review load, and the weekly log will show it if they do. Oversight may not be your limit, but you will not know until you count.
What to do next
Log one workflow's review hours beside its live volume for four weeks, then widen or narrow the number of live agents before adding another.
Inside the white paper
- How to estimate how much AI work one person can own
- How risk-based review and automated controls extend a team
- What the evidence supports, and where it stops
Sources and notes
- Alexander Sukharevsky, Alexis Krivkovich, Arne Gast, Arsen Storozhev, Dana Maor, Deepak Mahadevan, Lari Hämäläinen, and Sandra Durth, "The agentic organization: Contours of the next paradigm for the AI era," McKinsey & Company, September 26, 2025 — McKinsey's authors forecast that human oversight will cap how widely companies adopt AI agents, and say that in their experience a team of two to five can supervise 50 to 100 agents. No sample or method is given, so the paper does not treat the figure as a benchmark.
- M. L. Cummings and P. J. Mitchell, "Operator Scheduling Strategies in Supervisory Control of Multiple UAVs," Aerospace Science and Technology 11 (2007): 339–348 — Cummings and Mitchell report cognitive saturation and automation bias in a simulation with one operator supervising several semi-autonomous unmanned aircraft; the study did not measure performance as aircraft were added.
- M. L. Cummings and Stephanie Guerlain, “Developing Operator Capacity Estimates for Supervisory Control of Autonomous Vehicles,” Human Factors 49, no. 1 (2007): 1–15 — Cummings and Guerlain report lower performance and awareness at 16 missiles, with 8 and 12 also tested, in one simulation with 42 U.S. Navy personnel; the result does not set a business AI limit.
- David Mallon, Brad Kreit, and Natasha Buckley, "Rethinking operating models for humans with agents," Deloitte Insights, April 2, 2026 — Deloitte says oversight may include agents guarding other agents that are then guarded by humans.
- Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, and Kamie Roberts, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1, July 26, 2024 — The NIST Generative AI Profile suggests ongoing monitoring, red-team exercises, independent evaluation, incident response, and structured feedback.
- National Institute of Standards and Technology, "AI Risk Management Framework Core," AI RMF 1.0, 2023 — The NIST AI RMF Core organizes AI risk work into govern, map, measure, and manage functions and treats monitoring as ongoing.
- Raja Parasuraman and Dietrich H. Manzey, "Complacency and Bias in Human Use of Automation: An Attentional Integration," Human Factors 52, no. 3 (2010): 381–410 — Parasuraman and Manzey review mostly laboratory evidence that people monitor automation less closely when other tasks compete for attention and the automation is consistently reliable; the studies do not set an enterprise AI staffing ratio.