The Scale, Iterate, or Kill Decision With Evidence
The calendar invite is ten weeks old. It was booked in the same week the pilot was chartered, before a single invoice exception touched the model, and its body text has not changed since: the three scale conditions, the two kill conditions, the names of the five people who signed them. Tomorrow at 9:00 that meeting happens, and the pilot team lead is doing something almost no pilot team in the MIT sample ever got to do: walking into a decision with the evidence already agreed, the options already priced, and the verdict written in the first line of a one-page memo. Meanwhile, three floors up at a company that will remain hypothetical for exactly one more section, a different pilot is entering its fourteenth month of not being decided about, at a renewal cost of "only" twelve thousand dollars a month, and nobody in that building can tell you what evidence would end it. This lesson is about the difference between those two rooms. It is the shortest distance in this program between a discipline and a career.
The Four Endings of Every Pilot
Every pilot ever run ends in one of exactly four states. Three of them are decisions. The fourth is the absence of one, and it outnumbers the other three combined.
Scaled on evidence. The Delta Table showed a real, attributable improvement, the pre-committed success criteria were met, and the organization commits rollout money with its eyes open: named sites, a calendar, a hardening budget. This is the ending everyone imagines at kickoff.
Iterated with a named fix. The evidence was close but not clean: one criterion missed, one failure mode surfaced, one confound unresolved. The pilot goes another round, but not another lap of the same track. A legitimate iteration is a new mini-charter with one hypothesis, one fix, one re-decision date. Hold that definition; it is the load-bearing wall between iteration and the fourth state.
Killed with documentation. The evidence said no, the pre-committed kill criteria fired, and the pilot was shut down in a memo that records what was tested, what was learned, and what gets kept. In this program's accounting, a documented kill is a win: the organization bought knowledge at pilot price instead of ignorance at production price.
Undead. Not scaled, because nobody is confident enough to put their name on rollout money. Not killed, because nobody wants to host the funeral. Renewed by inertia, month after month, consuming license fees, meeting slots, and the credibility of everyone attached to it. The zombie pilot is not a decision that failed. It is a decision that never happened, and that distinction is the whole lesson. When MIT found that 95 percent of enterprise generative AI pilots deliver no measurable profit-and-loss (P&L) return, the popular reading was "95 percent failed." Read the autopsy more carefully and a large share of that 95 percent did not fail in any active sense; they simply never arrived at a moment where anyone was forced to say scale, iterate, or kill. The 95 percent is substantially a census of the undead. S&P Global's finding that 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before, is what a mass zombie-clearance looks like when the budget cycle finally forces it: not a considered verdict per pilot, but a bulk exorcism. And Gartner's prediction that over 40 percent of agentic AI projects will be canceled by the end of 2027 describes cancellations arriving late and acrimonious, which is what happens when the decision was never designed and has to be improvised under budget pressure years later.
You have spent this chapter installing the antidote, mostly without ceremony. The decision meeting was booked at design time, with the criteria pasted into the invite. The baseline was frozen before launch. The Delta Table was built honestly, confidence words and all, in the last lesson. What remains is the final artifact and the final hour: the Decision Memo, and the meeting it is written for. This is also where the transformer's professional identity lives at its deepest level. Plenty of people can launch pilots. The person whose pilots always end in a decision, every one of them, whatever the verdict, is rare enough that organizations learn their name.
A pilot that ends in a decision succeeded as a process, whatever the verdict. A pilot that ends in a renewal did not end at all.
The Decision Memo: One Page, Five Blocks
The Decision Memo is this lesson's named artifact, and it is the sibling of a document you already know. Level 2 taught the one-page readiness report: verdict line first, evidence compressed, the reader's decision made easy. The Decision Memo is that same discipline standing at the other end of the pilot's life. It is one page. It has five blocks, in a fixed order, and the order is doing most of the work.
Block one: the verdict, first
Line one of the memo is the verdict, stated with its conditions attached: "SCALE, with two conditions," or "ITERATE: one named fix, re-decision on March 14," or "KILL: salvage itemized below." Not the background, not the journey, not "as you will recall." Executives read line one and allocate their attention based on it; the L2 verdict-line discipline exists because a recommendation buried on page one, paragraph four, is a recommendation the room will reconstruct incorrectly from memory. If writing the verdict in line one feels exposed, that is the feeling of accountability arriving on schedule.
Block two: the evidence block
The Delta Table from the last lesson is imported whole: baseline value, pilot value, delta, attribution note, and confidence word for each chartered metric. Two rules govern this block. First, the confidence words survive intact. If the cost delta was PROBABLE on one month-end cycle, the memo says PROBABLE; upgrading it to proven prose because the memo wants to persuade is exactly the dishonesty the last lesson spent itself preventing. Second, the memo does not re-argue the numbers, it cites them. The Delta Table is the single source of truth, and every artifact downstream of it (this memo, the steering deck, the eventual rollout charter) points at the same figures rather than growing its own. The moment two documents carry two versions of the cycle-time delta, the meeting will be spent reconciling documents instead of deciding.
Block three: the three options, costed
This block is what separates a Decision Memo from an advocacy memo, and it is the block most first-time authors skip because it feels like arguing against yourself. All three options appear, each with a real cost attached, even the two you are not recommending.
- SCALE, costed. Rollout cost, target sites, calendar, and the hardening list: the honest inventory of what must move from pilot-grade to production-grade. The manual sampling lane becomes an automated one. The prompt that one analyst maintains becomes a versioned asset with an owner. The vendor's informal assurance about model updates becomes a contract clause. Scaling is not "do the pilot, but bigger"; the next lesson, the replication playbook, exists because that sentence has buried so many first wins.
- ITERATE, costed. The named fix, its cost, and the re-decision date. All three are mandatory, and the third is the one that keeps this option honest. An iteration is a new mini-charter carrying one hypothesis: "we believe fixing X will move metric Y past its threshold, and we will know by date Z." An iterate without a re-decision date is not an iteration. It is the zombie in a lab coat, inertia wearing the costume of rigor, and it will consume quarters exactly the way the undocumented renewal does.
- KILL, costed. The wind-down cost, and, crucially, the salvage list: what survives the kill. This is the item almost everyone underestimates. A kill rarely kills everything, because the redesign work you did in chapter two was real work on the process itself. In the running example, the queue restructure alone, the resequencing and routing changes made before the model touched anything, accounts for 0.8 days of the cycle-time improvement, and it survives any verdict because it never depended on the AI. The instrumentation stays too: the baseline, the counters, the event log are permanent process assets. So does the lessons register. Itemizing the salvage is what makes a kill survivable politically, because the sponsor can announce what was kept, and true economically, because the organization really does keep it.
Block four: the recommendation, with conditions
One option, argued in three or four sentences from the evidence block, never from enthusiasm. And its conditions, each one dated and owned: not "we should monitor vendor drift" but "vendor-drift monitoring contractualized before site one go-live; owner: procurement lead; due: contract renewal, September 30." A condition without a date and an owner is a wish, and wishes do not survive contact with a rollout.
Block five: the dissent line
One sentence per unresolved disagreement, with the dissenter named: "The data steward believes the vendor-drift risk is underweighted in this recommendation." This is the L2 disagreement-log discipline arriving at its highest-stakes moment, and it is absurdly cheap insurance. If the dissenter proves right, the organization has a dated record that the concern was raised, and it learns to weight that voice properly next time. If dissent is suppressed instead, the disagreement does not disappear; it goes underground and gets re-litigated in hallways for a year. Recorded dissent also does something subtler: it makes signing the memo easier for the dissenter, because agreeing to be overruled on the record is a much smaller ask than pretending to agree.
The Worked Example: The Invoice-Exception Decision Memo, Rendered in Full
Here is the memo for the pilot this level has been building since its first lesson. Every figure is illustrative, carried forward from the running example: roughly 1,150 invoice exceptions a month, a frozen baseline of about $428,000 annual run rate, the accounts payable (AP) routine lane redesigned in chapter two, six weeks of pilot data measured against baseline in the last lesson.
DECISION MEMO: Invoice-Exception Pilot (v1.0, one page)
Verdict: SCALE, with two conditions. (1) A second month-end observation at the first scale site before site two activates; the cost delta is currently PROBABLE on a single month-end cycle. (2) Vendor-drift monitoring moves from informal assurance to contract language before site one go-live.
Evidence (Delta Table v1.2, cited whole, source of truth unchanged):
| Metric | Baseline | Pilot (6 wk) | Delta | Confidence |
|---|---|---|---|---|
| Cycle time (routine lane) | 3.4 days | 1.9 days | -1.5 days (0.8 attributable to queue restructure, 0.7 to AI step) | PROVEN |
| Cost per exception | $31 | $24 | -$7 (-23%) | PROBABLE (one month-end observed) |
| Escape rate at gate | 1.1% | 0.4% | -0.7 pts | PROVEN |
| Analyst hours redeployed | 0 | 62 hrs/mo | +62 hrs/mo | INDICATIVE (self-reported, partial) |
Options, costed:
- SCALE: $48,000, three sites, nine weeks. Hardening list: automate the sampling lane, version and assign the prompt asset, contractualize drift monitoring, train two backup gate reviewers per site. Projected annualized saving at full rollout: about $96,000 against baseline, to be re-proven per site.
- ITERATE: Not applicable. No named fix is pending; no criterion missed its threshold. Listed to show it was considered, at an estimated $9,000 per additional six-week cycle if invoked.
- KILL: Wind-down $6,000 (license exit, documentation, handback). Salvage retained regardless of AI: queue restructure worth 0.8 days of cycle time, the instrumentation and baseline pack, the rewritten SOP structure, the lessons register. Net: even the kill keeps roughly half the operational win.
Recommendation: Scale. Two of four deltas are PROVEN, the primary cost delta is PROBABLE with a defined path to PROVEN (condition one), and the kill option's own salvage math shows the redesign is sound independent of the model. Conditions: second month-end observation at site one, owner: pilot lead, due before site two activation; drift clause, owner: procurement lead, due September 30.
Dissent: The data steward believes the vendor-drift risk is underweighted and would prefer the drift clause signed before site one begins onboarding, not merely before go-live.
That is the whole memo. Notice what the ITERATE line does even while declining the option: it proves the author priced all three doors before recommending one. Notice that the dissent line cost one sentence and bought the data steward's signature. And notice that the memo's most persuasive line for a skeptical chief financial officer (CFO) is inside the KILL option: a recommendation that shows what dying would salvage is a recommendation that has nothing to hide.
The Meeting: Forty Minutes When the Design Was Done Right
The decision meeting is run against the criteria attached to the invite ten weeks ago, and the agenda is almost embarrassingly simple: read each pre-committed criterion aloud against the Delta Table, walk the three costed options, take the vote on the recommendation and its conditions. When the design work was done right, that is 40 minutes. Compare the criteria-less pilot review meeting you have certainly attended: 90 minutes of anecdote wrestling, where the demo enthusiast's story competes with the skeptic's story and the loudest recollection wins. The meeting's length is a measurement of the design's quality. Every minute past 45 is usually a criterion that should have been pre-committed and was not.
Three hard moments recur in this meeting across organizations, and each has a script worth rehearsing out loud before you need it.
The sponsor who wants to scale past the evidence
The delta is good, the sponsor is delighted, and suddenly the proposal on the table is five sites by year end instead of three sites with a checkpoint. The script: "The cost delta is PROBABLE on one month-end. The condition is a second month-end at the first scale site. We scale the confidence with the footprint." Conditional scaling is the honest yes: it gives the sponsor the win and the momentum while keeping the evidence and the exposure in lockstep. You are not the person who says no to enthusiasm; you are the person who prices it.
The stakeholder who wants to kill a pilot that passed
Sometimes the criteria were met and someone still wants it dead: turf, budget, an unrelated grievance. The script is to read the pre-committed criteria aloud, then: "These are the thresholds we all signed before launch. The evidence clears them. If we kill a pilot that met its signed criteria, no kill condition we ever write will be signable again." The constitution defends the pilot too. Pre-commitment cuts both ways, and saying so out loud is precisely what made the kill condition signable back at design time: people accept a tripwire that can fire against them only if it also protects them.
The room that wants to extend without deciding
This is the zombie moment, and it always arrives sounding reasonable: "Let's give it another quarter and see." The script: "An extension is a decision to spend twelve thousand dollars for no new information. An iteration has a hypothesis, a fix, and a re-decision date. Which of those are we choosing?" The sentence works because it forces the fourth option to confess that it is an option, and the worst one on the table: all of iterate's cost, none of its learning. In most rooms, once extension is named as a purchase of nothing, nobody wants to be recorded buying it.
And here is how the invoice-exception meeting actually went, in the running story: 35 minutes. The verdict was adopted. The one change came from the CFO's analyst, the same hostile reviewer the team had deliberately pre-briefed during the measurement lesson: she tightened condition one so that the second month-end observation must complete before site two spends money, not merely before it goes live. The team accepted in the room, because she was right and because the pre-brief meant her challenge arrived as a refinement instead of an ambush. That is the hostile-analyst investment paying its dividend in real time: the sharpest person in the room improving your conditions instead of demolishing your evidence.
The Kill, Honored
This program has said since Level 1 that a documented kill is a win. This is the lesson where that conviction gets its full mechanics, because the verdict you will most need courage for is the one this section covers.
The kill memo is the Decision Memo with the verdict block pointed the other way, plus one discipline: it is structured as an asset inventory, not an apology. Four parts. What we tested: the hypothesis, the scope, the criteria as signed. What we learned: the specific, transferable findings, including which readiness assumptions held and which failed, stated plainly enough that the next pilot team can use them. What we keep: the salvage list, itemized: the process improvements from the redesign that never depended on the AI, the instrumentation that stays and keeps paying, the vendor evaluation notes, the lessons register entry. What it cost versus what month nine would have cost: the goldmine arithmetic. A readiness-driven kill in week one saves roughly nine months and six figures against the classic stalled-pilot arc. An evidence-driven kill at week six still saves about seven of those months and most of the money. Write both numbers down; the difference between them is the price of certainty, and it is almost always a bargain either way.
Then comes the part that is pure leadership mechanics, and it matters more than the memo: the kill is announced, publicly, by the sponsor, with the savings number and the salvage list in the same breath. "We killed the contract-summary pilot at week six. That decision saved us approximately seven months and $210,000 of continued spend, and we kept the intake redesign and the baseline instrumentation, which are already saving four hours a week." When leadership behaves this way even once, it reprices honesty for every future pilot team in the organization: the team that surfaces disqualifying evidence early gets credit instead of a stigma, and the next Delta Table you receive gets more honest, not less. This is the cultural seed of what Level 4 formalizes as the stage-gate discipline and Level 5 teaches as portfolio-level kill discipline. It starts here, with one sponsor saying one kill out loud like the win it is.
And say the career frame plainly, because the reader deserves it stated as bluntly as it operates: the transformer's ledger records kills as predictions vindicated. In the readiness economy, "I have killed two pilots, and both stayed dead" is a hiring credential. It certifies that your criteria mean something, that your evidence discipline survives pressure, and that the pilots you did scale were scaled for reasons. The Level 1 career lessons made this argument in the abstract. Here it is concrete: the person across the interview table has almost certainly inherited a zombie, and they are hiring you to be the person who ends things.
The Zombie, a Complete Anatomy
You have earned the full version of the failure story, the one this program was substantially founded on. It deserves telling with some sympathy, because no one in it is a fool.
A mid-size insurer launches a document-processing pilot: policy correspondence in, structured summaries out. The demo is good. The team is capable. The original sin is quiet and happens in week zero: no frozen baseline, and no decision meeting on any calendar. All figures that follow are illustrative, and none of them are exotic.
Month six. The metrics are ambiguous, because with no baseline they could never have been anything else; every claimed improvement dissolves under the question "compared to what?" The sponsor is protective, having spent visible credibility at launch. The team is genuinely fond of the tool, which works often enough to be likable. Renewal costs "only" $12,000 a month, a number carefully sized to be nobody's problem. The steering committee, facing ambiguity with no criteria to resolve it, does the natural thing: it renews. Note what did not happen: nothing failed. A decision was simply not taken, because no mechanism existed to take one.
Month fourteen. Three renewal cycles have passed. Cumulative direct spend is over $160,000. The team's best analyst, the one who should be running the next baseline study, has spent most of a year tending the tool: patching its outputs, managing the vendor, assembling optimistic slide-lets for reviews that decide nothing. The pilot has become furniture. Killing it now would mean explaining the $160,000, so each renewal quietly raises the price of ever deciding, which is the zombie's cruelest mechanism: it compounds.
Month nineteen. A new CFO runs a zero-based review and kills the pilot in a single meeting, not by disputing its value but by asking one question: "Show me the decision criteria." There were none. There had never been any. Nineteen months of operation, and the pilot was killable in ninety seconds by a stranger, because it had never possessed the one property that protects any initiative: decidability.
The autopsy's finding is the one to memorize. The pilot was never bad. It was never decidable. And the true cost was never the $12,000 a month. It was the analyst-year, spent tending instead of building. It was the nineteen months of steering-committee attention, the scarcest resource in the building. Above all it was the three better pilots never started, because the zombie occupied the budget line, the political appetite, and the team. The zombie's true rent is paid in opportunity, and it is always the largest number in the story, and it never appears on any invoice. Gartner's abandonment forecasts and S&P Global's 42 percent scrap statistic are, at scale, this exact story repeated across an economy: cancellations arriving late, angry, and improvised, because the decision was never designed in.
You now own the cure, all of it: the meeting booked at design time, the criteria in the invite, the baseline frozen, the Delta Table honest, the memo with the verdict in line one, the scripts for the three hard moments, and the kill honored in public when it comes. In the running story, the verdict was SCALE, with two conditions, one of them tightened by a friendly adversary. The next lesson closes this chapter by answering the question a scale verdict immediately raises: how do you turn one evidenced win into three, without discovering that what you actually built was one hero team and a lucky site? That is the replication playbook.
What to Do Monday Morning
Everything in this lesson can be drafted before your pilot ends, and most of it should be.
- Draft your verdict line today, in pencil. From your current Delta Table, write line one of your Decision Memo: verdict plus conditions, one sentence. If you cannot draft it, identify which metric's confidence word is blocking you; that is your remaining measurement work, named.
- Cost all three options, including the two you do not expect to choose. Scale with its hardening list, iterate with a named fix and a re-decision date (or "not applicable" stated explicitly), kill with its wind-down number.
- Write the dissent line with its owner named. Ask each signatory directly: "What disagreement should this memo record?" One sentence per dissent. If you collect none, ask harder; a memo with no recorded dissent usually means the disagreements are waiting for the meeting.
- Rehearse the three scripts out loud: the conditional-scaling script for the eager sponsor, the constitution script for the hostile stakeholder, the extension-confession script for the room that wants another quarter. These sentences fail when improvised and work when rehearsed.
- Compute your kill's salvage list even if you expect to scale. Itemize what survives without the AI: the queue and routing changes, the instrumentation, the rewritten SOP. This exercise regularly finds that the redesign already won something material, which strengthens every option, including the one you recommend.
- Open the decision-meeting invite and confirm the criteria are still in it, verbatim, with the signatories still attending. If the meeting has drifted off anyone's calendar, that drift is the first zombie symptom, and re-pinning it costs one email today.
Key Takeaways
- Name the four endings honestly: scaled on evidence, iterated with a named fix, killed with documentation, or undead, and recognize that the undead outnumber the other three combined; MIT's 95 percent is substantially a census of pilots that were never decided, not pilots that failed.
- Write the Decision Memo as one page in five fixed blocks: verdict first, the Delta Table cited whole with confidence words intact, all three options costed, a recommendation with dated and owned conditions, and a named dissent line.
- Cost the options you are not recommending, because pricing all three doors is what separates a decision document from an advocacy document, and the kill option's salvage list is often the memo's most credible paragraph.
- Define iteration strictly as a new mini-charter with one hypothesis, one named fix, and one re-decision date; an iterate without a re-decision date is the zombie in a lab coat.
- Run the meeting against the pre-committed criteria attached to the invite, expect roughly 40 minutes when the design was done right, and treat every minute beyond that as a measurement of missing design.
- Rehearse the three hard-moment scripts: scale the confidence with the footprint, the constitution defends the pilot too, and extension is a decision to spend money for no new information.
- Honor the kill in public: structure the kill memo as an asset inventory, have the sponsor announce it with the savings number and the salvage list, and record it in your own ledger as a prediction vindicated, because in the readiness economy a documented kill is a hiring credential.
- Remember the zombie's true rent: not the monthly renewal, but the analyst-year, the committee attention, and the better pilots never started; decidability, installed at design time, is the only vaccine.
Skill.re