TL;DR
An AI assistant approved for one team's use is not approved for the next team's. The second team can depend on the first team's experience for a use nobody tested, or reopen the whole debate and make every new use as slow as the first. The paper recommends an evidence pack that travels with the system and says what it covers.
What the paper develops
Consider an AI assistant that drafts summaries of customer contracts for a sales-operations team. A risk reviewer approved that use after the team tested it. Now the procurement team asks to use the same assistant on supplier contracts. The first approval covers customer contracts, and its tests and notes sit in one team's files. The procurement lead can depend on the first team's experience for a use nobody tested, or reopen the whole debate, which makes every new use as slow as the first. This is an illustration, not an account of a real company.
This paper's position is that the proof has to be written so it can travel and be checked again. The record is an evidence pack. In it, the system is the AI assistant or tool itself, and a use is one task for one group of people. The pack says what the system did, on what basis, where it has and has not been tested, and who answers for depending on it, and it is kept current as the system and its uses change. The paper starts after the first approval, which "The Demo-to-Production Gap" covers. In this paper's design, that approval and its tests become the first use's part of the pack, and the second use answers the four questions again plus whichever of that paper's seven production questions, such as ownership or stop authority, the new use changes.
A strong pilot often ends up on low-stakes work at the edge of the business. That is the paper's observation, and no source here counts how often it happens. A tempting response is to ask for a better model. The paper reads the stall differently: the people accountable for important work cannot show that they can depend on the system for that work. A new model is a change like any other, so the tests run on the old one no longer cover it.
This is a reading, not a measured cause. McKinsey's survey found security and risk concerns to be the top barrier to scaling agentic AI, and Deloitte's found that few companies have a mature governance model for autonomous agents. McKinsey asks what holds leaders back and Deloitte asks how mature their governance is. Neither measures why a particular pilot stalled. The reading would be wrong if a pilot with a current pack, a named owner, and tested limits were still kept out of important work, or if a pilot moved into important work on a better model alone, with no new tests.
A use means one task, for one group of people, with one kind of input. For each use, the owner who answers for it has to show four things: what the system did, the basis for its result, where it has and has not been tested, and who answers for depending on it. Because the answers belong to a use, a second use starts with four new questions, and the first team's answers only partly carry over.
The pack has two parts, kept by two people. The system part is true for every use: the version and supplier, what has changed, test results, incident history, and known weak spots. The system steward, the person or team that runs the system for everyone, keeps it and keeps a register of uses. The use part covers one use: the purpose, the cases tested and not tested, how human review works, and the owner. It ends with one line saying what the pack covers and what it does not. A reviewer who did not build the pack reads it against the four questions and tries to find an answer that is missing or untested.
When the second team asks, its owner checks that the system part is current, writes a use part, and tests a sample of its own cases. If all four answers are there, the use goes ahead. If one is missing, the use stays a trial on low-risk work. If one is weak, the use is limited to the cases the tests cover. Nothing here measures how much time the review saves; the paper expects it to look at the differences instead of the whole system.
A pack filed once is a snapshot. The paper sets five triggers for testing again: the system changes, the inputs change, the use widens, something goes wrong, or a fixed interval passes with none of those. A better model is a trigger like the rest. The owner decides what to retest, and the reviewer samples the changes the owner called small. A trigger with no retest behind it counts as a missing answer for whatever it touches. The triggers are the paper's design, drawn on NIST guidance that treats AI risk work as continuous and asks for re-evaluation when a model is adapted to a new domain.
The pack becomes paperwork if no decision ever uses it, so each page must answer one of the four questions. It is not a certification: no single standard governs AI evidence packs, and a complete pack can still sit beside a system used past its tested limits. A small team with one use can write one page and split it when a second use appears. The paper's cost figures are planning estimates, a day or so for the system part and a few days for a use part. No study tests whether packs shorten a second review or reduce incidents, and NIST says assurance cases for software were still under study.
Pick the system that already has one approved use and is most likely to get a second. Have its steward assemble the system part and start the register of uses, have the first owner write the use part and end it with the covers line, and name a reviewer from outside the team. When the next team asks, hand over the pack and ask that team's owner to answer the four questions for the new use. If the system has no approval for its first use yet, begin with "The Demo-to-Production Gap."
What to do next
Pick the system that has one approved use and is likely to get a second. Have its steward assemble the system part of the pack and a register of uses, have the first owner write the use part ending with what it covers, and name a reviewer. When the next team asks, it answers the four questions for its own use.
Inside the white paper
- Why a second team's request starts four new questions about the same system
- An evidence pack in two parts, what each part holds, and who keeps it
- Five triggers for testing again, the objections, and what no source tests
Sources and notes
- Gabriel Morgan Asaftei, Roger Roberts, Abby Sticha, and Cécile Prinsen, "State of AI trust in 2026: Shifting to the agentic era," McKinsey & Company, March 25, 2026 — McKinsey's 2026 AI Trust Maturity Survey, of about 500 organizations from December 2025 to January 2026, found nearly two-thirds of respondents naming security and risk concerns as the top barrier to fully scaling agentic AI. Organizations with clear ownership for responsible AI averaged 2.6 on its four-level maturity scale against 1.8 for those without a clearly accountable function. Both are survey results and an association, not causes.
- Deloitte, "The State of AI in the Enterprise," 2026 AI report, Deloitte AI Institute — Deloitte's survey of 3,235 senior leaders, fielded August to September 2025, found that only one in five companies has a mature model for governance of autonomous AI agents. It reports leader answers, not the causes of stalled pilots.
- Thomas Rhodes, Frederick Boland, Elizabeth Fong, and Michael Kass, "Software Assurance Using Structured Assurance Case Models," NIST Interagency Report 7608, National Institute of Standards and Technology, May 2009 — NIST's 2009 report describes a structured assurance case, a documented body of evidence for claims about a system in a given application and environment. It covers software, not AI, calls the method an active topic of research, and names further work on how much evidence is enough and on maintaining and revising cases.
- Matthew Arnold, Rachel K. E. Bellamy, Michael Hind, Stephanie Houde, Sameep Mehta, Aleksandra Mojsilović, Ravi Nair, Karthikeyan Natesan Ramamurthy, Darrell Reimer, Alexandra Olteanu, David Piorkowski, Jason Tsay, and Kush R. Varshney, "FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity," arXiv:1808.07261, 2018; revised 2019 — The FactSheets paper proposes a standard document that an AI service supplier fills in on purpose, performance, safety, security, and provenance, including the intended use, the domains tested, and whether it is updated at each retraining. It is a proposal with example answers for two fictional services, not a study of adoption, and a supplier writes it before it knows a given buyer's use.
- Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, and Kamie Roberts, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1, July 26, 2024 — NIST's July 2024 generative AI profile is a voluntary companion to its risk framework. It suggests documenting a system's knowledge limits, defining who reviews periodically, and re-evaluating risk when a model is adapted to a new domain.
- National Institute of Standards and Technology, "AI Risk Management Framework Core," AI RMF 1.0, 2023 — The NIST AI Risk Management Framework Core organizes AI risk work into govern, map, measure, and manage and describes it as continuous and iterative. It is voluntary and is not a checklist.