TL;DR
A capable AI pilot impresses everyone and then stalls at the safe edge of the business, and a better model does not move it.
What the paper develops
A capable AI pilot impresses everyone in the demo, gets a funded trial, and then stalls. It runs on low-stakes tasks at the edge of the business and never crosses into the work that would justify it. The reflex is to ask for a better model, and a more capable model does not move a pilot that is stuck, because capability is not what is missing. The organization has not let the system into the work that matters because it cannot answer, for the people accountable, one plain question: can we rely on this?
What a business checks before it relies on a system
A demo shows that the system can produce a good result on a curated example. Reliance in production depends on something the demo never puts on screen: whether the people accountable can see what the system did and decide, on evidence, whether to depend on it. A capability is trustworthy, in operating terms, when a named owner can show what the system did, on what basis, within what limits it is known to work, and who is accountable for the outcome. If any of the four cannot be shown, it is not yet trustworthy for that use, whatever the demo looked like.
The evidence pack
The proof travels with the capability as an evidence pack: its purpose and intended use, its measured performance and the conditions that held under, its safety and security properties, the provenance of its data and models, the tested limits of its validation, the human oversight design, and the named owner. The idea has a research lineage in AI FactSheets, and the NIST AI Risk Management Framework supplies the operating loop, govern, map, measure, and manage, that keeps a system's use under review as it operates. Packaged this way, evidence compounds: a capability that earned reliance in one workflow arrives at the next with its proof attached, and the review covers only what changed.
Keeping it current
A pack assembled once and filed is a snapshot, and AI systems do not hold still. Models drift, inputs shift, and use expands past the envelope the system was validated for. Someone has to own keeping the pack current, watching for drift, and refreshing the validation as use grows. A living pack is also what lets an organization move faster with less risk, because it can extend a trusted capability into new work without re-proving it from zero.
Where evidence discipline goes wrong
Evidence has its own failure modes. Evidence theater is volume that counts activity, dashboards and decks that never answer the four questions. False precision is a number that implies more validation than was performed, hiding the gap it should expose. The launch artifact is a pack built to clear an approval and then never maintained, decaying while the system changes. Each produces the appearance of evidence without its function; the test in every case is whether the named owner can still answer for the system today.
Start with one high-consequence use
Do not try to earn trust across the whole AI portfolio at once. Pick the single use that carries the highest consequence if it is wrong and build one living, owned, inspectable evidence pack for it. That first pack becomes the template, and the discipline spreads by demand as other teams see what reliance required. Return to the pilot parked at the edge and the reflex to ask for a better model: what actually moves it is evidence a named owner can show and keep current as the system runs.
The operating move
Pick the single highest-consequence AI use and build one living, owned, inspectable evidence pack for it: what the system did, on what basis, within what limits, and who is accountable. Let that first pack become the template.
Inside the white paper
- The four questions a named owner has to answer before a business relies on an AI capability
- The evidence pack that carries the proof into production, and how it compounds across deployments
- How evidence discipline fails - theater, false precision, and the unmaintained launch artifact
Sources and notes
- Gabriel Morgan Asaftei, Roger Roberts, Abby Sticha, and Cécile Prinsen, "State of AI trust in 2026: Shifting to the agentic era," McKinsey & Company, March 25, 2026. mckinsey.com
- Matthew Arnold, Rachel K. E. Bellamy, Michael Hind, Stephanie Houde, Sameep Mehta, Aleksandra Mojsilovic, Ravi Nair, Karthikeyan Natesan Ramamurthy, Darrell Reimer, Alexandra Olteanu, David Piorkowski, Jason Tsay, and Kush R. Varshney, "FactSheets: Increasing Trust in AI Services through Supplier’s Declarations of Conformity," arXiv:1808.07261, 2018 (rev. 2019). arxiv.org
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1, July 2024. Verified July 9, 2026. doi.org
- National Institute of Standards and Technology, "AI Risk Management Framework Core," excerpt from AI RMF 1.0, 2023. Verified July 5, 2026. airc.nist.gov
- Deloitte, "The State of AI in the Enterprise," 2026 AI report, Deloitte AI Institute. deloitte.com
- Jessica Apotheker, Sylvain Duranton, Vladimir Lukic, Nicolas de Bellefonds, and Christoph Schweizer, "As AI Investments Surge, CEOs Take the Lead," BCG AI Radar 2026, January 15, 2026. bcg.com