TL;DR
Hitting the target does not prove the decision was sound or even that the outcome was successful. Judge what was achieved, how it was achieved, and whether the original decision was reasonable given what was known at the time.
What the paper develops
Getting home is not the whole outcome
Suppose a drunk driver arrives home uninjured and in record time. If the scorecard records only his arrival time and his own condition, the trip looks successful.
But on the way, he hits a cat, breaks several traffic laws, and drives through someone's yard. His decision to drive was reckless. The way he drove broke laws and damaged property. The trip harmed others. Once the whole trip is counted, "made it home" is not a successful outcome.
Portfolio reviews make a quieter version of the same mistake when they treat a delivered benefit or an early finish as proof of good judgment. They collapse three questions: What was achieved? How was it achieved? Was the original decision sound given what was known at the time?
Keep the result, the execution, and the decision separate
First score the full outcome. Did the program deliver the promised benefit within the approved cost, schedule, and risk limits? What collateral costs or harms accompanied it? A narrow target does not define the whole outcome.
Then inspect execution. Was the approved plan followed? Were controls bypassed, dependencies ignored, or risks shifted elsewhere? The route matters even when the destination is reached.
Finally grade the original decision. Was the problem framed correctly? Were realistic alternatives compared? Were assumptions, estimate ranges, dependencies, and accepted risks explicit? Was the choice reasonable given the evidence available at approval?
To make that comparison possible, preserve a short record at approval: the problem, alternatives, assumptions, estimate ranges, accepted risks, and evidence that would change the decision. At review, keep anything learned later in a separate column.
The decision grade and full-outcome score produce four combinations. Sound/good work gets celebrated; weak/bad work gets investigated. The mixed cases carry more learning. A sound decision with a bad result tests whether leaders will protect disciplined risk-taking. A weak decision with a good result tests whether they can question a success before its tactics become mandatory. The approved-versus-executed route and the variance pattern inform the attribution.
Attribute the gap before assigning credit
Start with the declared range. A result inside a well-calibrated range is uncertainty resolving. Next compare the approved plan with what was executed: changed scope, missed controls, or ignored dependencies belong to delivery governance. Finally examine comparable decisions. Repeated misses in the same direction are evidence that the estimating model or gate criteria need correction.
These categories overlap, so the review should record the attribution and the evidence behind it rather than force false certainty. The full-outcome score never edits the decision-quality grade. Both grades and the attribution record should appear in the committee packet.
One result should not settle a sponsor's reputation or an estimating model. Calibration accumulates across comparable decisions. Scattered errors around the estimate may indicate an imprecise but roughly centered process; repeated misses in the same direction are a reason to investigate systematic optimism, omitted dependencies, or benefit definitions that do not survive delivery.
Change the incentive inside the review
Review successful programs as well as failures. Report calibration across comparable decisions, not sponsor rankings from isolated outcomes. Protect a sponsor whose documented decision was sound and penalize any attempt to retrofit the record after results arrive.
The practical shift is small: add a decision-quality grade beside every full-outcome score, preserve the approval-time record, and close every material variance with one attribution record and a named update to the model, delivery system, or accountability record. That separation turns retrospective review from a verdict into an operating-control loop.
The operating move
At review, ask three questions in order: What was achieved when the whole consequence is counted? How was it achieved compared with the approved route? Was the original decision sound given what was knowable? Record two grades plus one attribution record before reallocating anything.
Inside the white paper
- The four combinations of decision quality and outcome, including the weak decision that happened to win
- A four-move retrospective that reconstructs, grades, scores, attributes, and updates without hindsight contamination
- Three tests for distinguishing declared uncertainty, execution divergence, and a decision-model error
Sources and notes
- Jonathan Baron and John C. Hershey, "Outcome Bias in Decision Evaluation," Journal of Personality and Social Psychology 54(4), 1988 — the canonical experiment showing that outcomes alter ratings of otherwise identical decisions.
- Sriraj Aiyer and colleagues, "Outcomes Affect Evaluations of Decision Quality," International Review of Social Psychology 36(1), 2023 — an open-access direct replication with 692 participants.
- Bent Flyvbjerg, "Quality Control and Due Diligence in Project Management," 2013 — the outside view and reference-class discipline for testing what was knowable at approval.
- Jan C. van Ours, "Outcome Bias in Managerial Decisions," Journal of Economic Psychology 112, 2026 — open-access field evidence from managerial replacement decisions.
- Society of Decision Professionals, "Decision Education at GM" — six elements used here as an inspection framework, not a guarantee of results.
- Michael J. Mauboussin and Dan Callahan, "Outcome Bias and the Interpreter," Credit Suisse Global Financial Strategies, 2013 — practitioner framing for separating process, skill, and luck.