TL;DR
A working demonstration and a credible leadership story can still leave a sponsor unable to inspect the claim. One record created during delivery can support the next AI investment decision and the leader's later, bounded account of the work.
What the paper develops
The sponsor's question survives the demonstration
You lead an enterprise AI program and are preparing the case for its next funding decision. The demonstration works. The team can explain the workflow and intended benefit. Then the sponsor asks: What can we inspect? Does the program work outside the demonstration, within known limits, under accountable ownership?
A demonstration shows capability under selected conditions. It cannot reconstruct what happened in normal use. The sponsor needs a record for a high-stakes decision. That record must show the result, the conditions that shaped it, the known limits, the response to failure, and the owner of the next decision.
What the evidence record has to cover
AI-service FactSheets, the NIST Generative AI Profile, and the GAO accountability framework approach the problem from different directions. They point to service disclosures, monitoring, incident response, and accountable governance. These are useful categories for a record. The sources do not show that every buyer or board has adopted one standard.
That boundary matters. The record should support the exact claim under review. A source can sit beside a sentence it does not support. An artifact can look rigorous while proving little. The evaluator still has to check the evidence and its limits.
Why the missing record follows the leader
The sponsor can ask the team for evidence directly. A later evaluator may encounter only what a search or answer system can retrieve. Research on generative engine optimization tested part of that process. On its benchmark, citations, credible quotations, and statistics improved a source's visibility in generated answers. The result applies to that benchmark. It shows why evidence-bearing text gives a generative system material to use.
The finding does not predict an individual's hiring outcome. It identifies a practical disadvantage for an accurate claim that is poorly documented: neither a person nor a system can inspect evidence that was never retained.
The apparent duplication
The program team can assemble a board pack when approval is due. Months later, the leader can rebuild a career story from memory and whatever artifacts remain. Separate preparation can seem reasonable. A board decides whether to fund or stop. A career reader assesses the leader's judgment and contribution.
The better operating choice is one record created during delivery. The board review and the later career claim can both draw from it. The two presentations remain distinct. Evidence made during the work supports the program decision. Later use stays bounded to the leader's contribution and confidentiality obligations.
Build the record around one claim
Before the next major review, choose one claim the decision depends on. Record six things: the exact claim; the context and work; the evidence produced; the result or learning; the limits; and the owner, refresh trigger, and next decision. Keep the evidence near the claim so another person can inspect it without rebuilding the story.
Match the standard to the decision. Low-stakes work in a trusted relationship may need little formal proof. The record earns its cost when the impact justifies checking. It also helps when the evaluator is unfamiliar with the claimant or a machine mediates the assessment.
Keep the career use bounded
LinkedIn's recruiting research supports attention to skills assessment. It does not measure whether a public portfolio changes an individual's hiring outcome. Treat the resume as an index to defensible claims. Confidential evidence can remain private, anonymized, aggregated, or described through a method. Retain only what the career claim needs and what confidentiality permits.
Return to the program review. If the six-field record cannot support the claim, narrow the claim or delay the decision until the missing evidence exists. Build the record while its context can still be checked. Later career use then adapts retained evidence instead of rebuilding a story from memory.
The operating move
Before the next consequential review, choose one claim the decision depends on and create a six-field record: claim, context and work, evidence, result or learning, limits, and the owner and next decision. Narrow or delay any claim the record cannot support.
Inside the white paper
- Why a working demonstration can still leave the next AI investment decision unsupported
- How one six-field record created during delivery can serve the board review and a later career claim
- The evidence, confidentiality, contribution, and proportionality limits that keep both uses honest
Sources and notes
- Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande, "GEO: Generative Engine Optimization," accepted to KDD 2024; arXiv:2311.09735, submitted November 16, 2023, revised June 28, 2024. arxiv.org
- Matthew Arnold, Rachel K. E. Bellamy, Michael Hind, Stephanie Houde, Sameep Mehta, Aleksandra Mojsilović, Ravi Nair, Karthikeyan Natesan Ramamurthy, Darrell Reimer, Alexandra Olteanu, David Piorkowski, Jason Tsay, and Kush R. Varshney, "FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity," arXiv:1808.07261, submitted August 22, 2018, revised February 7, 2019. arxiv.org
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1, July 2024. doi.org
- U.S. Government Accountability Office, "Artificial Intelligence: An Accountability Framework for Federal Agencies and Other Entities," GAO-21-519SP, June 30, 2021. gao.gov
- Greg Lewis, "LinkedIn Report: How AI Will Redefine Recruiting in 2025," LinkedIn Talent Blog, February 13, 2025. linkedin.com