Work, adoption & judgment · Field note

AI should make people better thinkers, not just faster producers

TL;DR

A polished AI answer can hide who formed the view, tested it, and owns the final decision.

What the paper develops

Your team now submits AI-assisted work faster. One recommendation reads well and reaches the review early. Then a leader asks which assumptions carry the decision and what evidence could overturn it. No one can answer without opening the chat again. The tool did not just help write the analysis. It supplied the first view, the challenge, and the wording of the final decision.

That is the risk behind an answer-first rollout. Adoption can rise while independent judgment becomes harder to see. The cost appears when the model is wrong, the case is unusual, or the tool is not available. A person who only approved a finished answer may not be ready to rebuild the reasoning.

The same model can support a different habit

Two randomized trials show why the interaction matters. Bastani and colleagues studied nearly 1,000 high-school students. Students with a standard AI tool later scored 17 percent below a no-AI group on an exam without the tool. A guarded version of the same model led students through hints and largely avoided that loss. Kestin and colleagues tested a custom AI tutor in introductory physics at Harvard. The AI group had more than twice the median learning gain of a class using strong active-learning methods.

The trials used different students, subjects, and designs. Neither tested experienced workers. They do not prove that a workplace will get the same result. They do show that access to a capable model does not decide whether people learn. The way people use it can keep reasoning active or let them skip it.

A review step needs skill and ownership

The usual phrase, “human in the loop,” is too weak on its own. It does not say whether the person can spot a poor assumption, check the evidence, or reject the answer. It also does not say who makes the final decision. A late signature is not the same as review.

Workplace evidence makes that distinction useful. In a field experiment with 758 consultants, AI users did better on tasks inside the model's tested range. On one task outside that range, they were 19 percentage points less likely to reach the correct answer. People whose results fell tended to accept the AI output and question it less. Other studies show that a wrong AI explanation can sound as convincing as a right one.

Make people form a view before they ask

For high-judgment work, start with a short pre-mortem. Before opening the tool, the worker states the answer they expect, names the assumptions that must be true, and says what evidence would change their mind. AI then challenges that first view. It can find a weak assumption, raise a counterargument, or point to missing evidence. The worker checks the challenge and owns the conclusion.

This pattern does not make the model more reliable. It makes the person's reasoning visible. It also gives a reviewer something concrete to inspect: the first view, the challenge, the evidence, and the final decision.

Support the habit in three places

Training should teach people to test assumptions and evidence, not just write prompts. IT and operations should give them a safe, usable tool and require stronger checks when the cost of error is high. Managers should ask what changed the person's view and who makes the decision. If any one of those pieces is missing, deadline pressure will push the work back toward ask, receive, and paste.

Start with one quarter of real work

Choose one function where judgment already matters, such as strategy, risk, client advice, or product decisions. Name one executive owner. For one quarter, require the three-step pre-mortem on selected work. Compare speed, rework, error, and review quality with the old method.

Then decide whether to expand, change, or stop. The goal is to keep people able to form, test, and own a view when the answer matters, without adding needless delay. AI should help people think better, not make their thinking disappear behind a polished response.

What to do next

For high-judgment work, have people state their first view, key assumptions, and what would change their mind. Then let AI challenge it.

HANDOFFSLEARNINGRECOVERY

Inside the white paper

  • What two student trials do and do not show about answer-first and reasoning-first AI
  • Why a human signature is not the same as skilled review
  • How to test a three-step pre-mortem in one quarter of real work

Sources and notes

  1. Challapally, A., Pease, C., Raskar, R., and Chari, P. (2025). The GenAI Divide: State of AI in Business 2025 (v0.1). MIT NANDA. The $30–40B enterprise investment figure and the 95% zero-return figure are from this preliminary report. Methodology: 52 structured interviews, 300+ disclosed AI initiatives, 153 senior leaders — this preliminary report estimated enterprise spending and reported that most studied organizations saw no measurable return; it does not establish the cause.
  2. Microsoft and LinkedIn, 2024 Work Trend Index Annual Report. Survey of 31,000 people, 31 markets, Feb–Mar 2024. The 78% BYOAI figure includes any personal tool use supplementary to employer-provided tools and may overstate fully unsanctioned behavior. Only 39% of AI users had received AI training from their company — the survey reports widespread use of personal AI tools at work and limited company training among AI users.
  3. SANS Institute, "Sunlight AI: Bringing Shadow AI Into the Light," December 2025. The 48% would-not-stop figure draws on October 2024 survey data — this survey report describes continued AI use despite employer bans.
  4. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O., and Mariman, R. (2025). Generative AI without guardrails can harm learning. PNAS, 122(26), e2422633122. Pre-registered RCT; ~1,000 students, 9th–11th grade; Turkey; IRB-approved. August 2025 correction updated one author's affiliation only; findings unchanged — this pre-registered student trial compared standard AI, a guarded tutor, and no AI.
  5. Dell'Acqua, F., et al. (2026). Navigating the jagged technological frontier. Organization Science, 37(2), 403–423; first circulated as HBS Working Paper No. 24-013 (2023). Pre-registered; n=758 BCG consultants. "Blindly adopt" wording is from the working paper; its retainment measure (copying AI output) was associated with better in-frontier performance. BCG co-designed a study finding BCG consultants perform better with AI; pre-registration and multi-institutional authorship partially mitigate this — this field experiment compared consultant performance on tasks inside and outside the model's tested range.
  6. Khanna, M., Wang, Z., Wei, L., and Xue, L. (2025). When Medical AI Explanations Help and When They Harm. arXiv:2512.08424. n=257 medical students; 3,855 diagnostic decisions. https://arxiv.org/abs/2512.08424. Also: Groot, T., and Valdenegro-Toro, M. (2024). Overconfidence Is Key: Verbalized Uncertainty Evaluation in GPT-4, GPT-3.5, LLaMA2 and PaLM 2. Proceedings of TrustNLP 2024, ACL — these studies examine model overconfidence and the effect of correct and incorrect AI explanations on medical students.
  7. IBM, Cost of a Data Breach Report 2025 (Ponemon Institute), July 30, 2025. Organizations with high shadow-AI levels saw $670,000 higher breach costs than those with low or no shadow AI; one in five organizations reported a breach due to shadow AI — IBM reports higher breach costs and shadow-AI breaches in its 2025 study.
  8. Netskope, Cloud and Threat Report, October 2024–October 2025. 223 incidents a month of users sending sensitive data to AI apps, as reported by Cybersecurity Dive; the count doubled year over year — this report describes incidents in which users sent sensitive data to AI apps.
  9. UpGuard survey of security leaders, 2025. Reported in Fortra, November 2025 — this secondary report cites a survey of security leaders' use of unapproved tools.
  10. Lee, H-P., et al. (2025). The impact of generative AI on critical thinking. CHI '25, Yokohama. Microsoft Research. "Task stewardship" is the authors' own framing. Limitations: self-report; cross-sectional; critical thinking not objectively measured; causality not established — this CHI study links confidence in AI with lower self-reported critical-thinking effort; it does not establish causation.
  11. BYOD-era security practice is cited only as a structural analog; no AI-specific provision study of equivalent design currently exists — the paper uses BYOD practice only as a structural analogy, not as direct evidence about AI.
  12. Kosmyna, N., et al. (2025). Your brain on ChatGPT. arXiv:2506.08872. MIT Media Lab. Preprint — not peer-reviewed as of March 2026. Formal commentary arXiv:2601.00856 identifies five methodological concerns. Cited as directionally consistent only — this EEG preprint is treated as directional evidence because formal commentary raised material method concerns.
  13. Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), 6. n=666. Correlational; recommends balanced AI integration, which the study did not test. No objective performance measures — this correlational study links frequent AI use, cognitive offloading, and critical-thinking scores; it did not test a better use pattern.
  14. Wu, S., Liu, Y., Ruan, M., Chen, S., and Xie, X.-Y. (2025). Human-generative AI collaboration enhances task performance but undermines human's intrinsic motivation. Scientific Reports. Four online experiments (Prolific); N=3,562. Yukun Liu is corresponding author — four online experiments found higher AI-assisted performance followed by lower motivation in later solo work.
  15. CybSafe and National Cybersecurity Alliance, survey of 7,000 respondents, late 2024. The 38% confidential-data-sharing figure is from Cloud Security Alliance analysis — this survey analysis reports confidential work data shared with AI without approval.
  16. Microsoft, New Future of Work Report 2025. The inference that resistance to top-down mandates extends specifically to established BYOAI interaction habits is the author's extension; not directly established by the cited source — the report discusses the limits of top-down change; applying that point to AI habits is the paper's inference.
  17. Klein, G. (2007). Performing a project premortem. Harvard Business Review, 85(9), 18–19. Validated in: Veinott, E. S., Klein, G., and Wiggins, S. (2010). Evaluating the effectiveness of the PreMortem technique on plan confidence. ISCRAM, Seattle. The adaptation here — individual cognitive priming before AI engagement — extends Klein's team-based technique and has not been independently validated in that specific form — Klein's team pre-mortem and a controlled follow-up inform the paper's unvalidated individual adaptation.
  18. Ma, S., et al. (2024). "Are you really sure?" CHI '24, Honolulu. Three studies; "Think the Opposite" was tested in Study 2 without AI assistance, and Study 3 found that confidence calibration reduced under-reliance on AI but not over-reliance. The intervention is similar to the pre-mortem prompt protocol and was developed independently. Limitation: income-prediction task in a lab setting — these studies tested confidence calibration and a think-the-opposite prompt, with mixed effects on reliance.
  19. Kestin, G., Miller, K., Klales, A., Milbourne, T., and Ponti, G. (2025). AI tutoring outperforms in-class active learning. Scientific Reports. RCT with crossover design; 233 enrolled, 194 eligible Harvard undergraduates; introductory physics; fall 2023. Effect size 0.73–1.3; p < 10⁻⁸. Critical limitations: selected population; two topics; introductory material; authors do not presume the result extends to complex synthesis and higher-order critical thinking — this student trial compared a custom AI tutor with active learning in two introductory physics topics.
  20. Curran, F. C., and Goo, M. (2025). Disciplining AI Use. University of Florida Education Policy Research Center policy brief, reporting K-12 teacher survey findings collected by Dwyer, L., and Laird, E. (2024) — this policy brief describes school responses to student AI use.
  21. HEPI, Student Academic Experience Survey 2025. UK university student AI use for assessments: 53% (2024) to 88% (2025). HEPI 2025 also reports 67% saying good GenAI skills are essential. The 38% teacher figure is from the Walton Family Foundation / Impact Research teacher survey (February 2023) — this secondary summary reports student-use and teacher-preparation survey figures; the paper uses it only as context.
  22. Autor, D., and Thompson, N. (2025). Expertise. NBER Working Paper 33941. The inference that AI is specifically automating the non-expert tasks of knowledge work is the author's application of the framework; Autor and Thompson do not make this specific claim about generative AI — this labor-economics paper explains how automation of expert and less-skilled tasks can affect employment and wages; the AI application is the author's inference.