Killing Projects Early: The Stage-Gate Discipline
Norvik Group's scrap year did not end with a series of thoughtful decisions. It ended on a Thursday afternoon in November, in a zero-based budget review, when the chief financial officer went down a list of AI initiatives and cancelled three of the four in about ninety minutes. Roughly $2.8 million written off. Two sponsors who had defended their projects for months in status meetings found out in the room. The technology was not the problem and the people were not stupid. What Norvik lacked was any mechanism for ending a project on purpose, early, cheaply, and in public. So every project ended the only way projects end without such a mechanism: late, expensively, and in a batch, at the hands of someone with the authority to be blunt. This lesson builds the mechanism.
Why One Honest Kill Does Not Scale
Level 3 taught you how to end a single pilot honestly: the one-page decision memo with the verdict first, the forty-minute meeting where the room argued about evidence rather than feelings, the anatomy of the zombie (the pilot never killed, only "paused" or "de-prioritized pending capacity," quietly consuming a license renewal every year), and the rule underneath all of it: success and kill criteria are written before the pilot starts, because criteria written afterward are just descriptions of whatever happened.
That works for one pilot run by one transformation lead who cared enough to design for the ending. At portfolio scale it is not enough, and the reason has nothing to do with the method.
When killing a project is an individual act, it requires individual courage, and courage is scarce and personally expensive. Picture doing it eleven times. Each time you tell a director that the thing they announced to their function, staffed with two of their best people, and mentioned to the chief executive in a corridor, is going to stop. Each time you spend relationship capital you will need later, and become slightly more identifiable as "the person who kills things." After the third or fourth, something human happens: you begin, without deciding to, to find reasons why the fifth deserves another quarter. The organizational default reasserts itself, not through anyone's cowardice but through the accumulated friction of eleven separate acts of bravery. Courage does not scale. Institutions do.
A stage gate is that institution. The term comes from capital projects and new-product development, and in operator language it means a pre-scheduled decision point at which an initiative must present specified evidence to earn its next tranche of money and attention. It is not a status review. A status review asks "how is it going?", a question about effort and narrative whose honest answer is always "well, considering." A gate asks: "does the evidence in front of us meet the threshold we agreed before this item started?" That question has an answer, it can be answered in eleven minutes by people who are not experts in the item, and when the answer is no, nobody in the room has to be brave, because nobody in the room is deciding. The decision was made months earlier, by the same people, in calmer conditions. Today's forum executes it.
A gate does not make killing a project easy. It makes killing a project normal, which is the only property that survives contact with a portfolio.
Rereading the 42 percent
S&P Global found that 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before, and Gartner forecast that more than 40 percent of agentic AI projects will be cancelled by the end of 2027. Read casually, those sound like organizations killing too much. They are organizations killing too late. Nobody scraps most of a portfolio in month two. That happens in month fourteen, in a budget crisis, in one afternoon, after the licenses are signed, the integrations half-built, and the credibility spent. It is Norvik's November: not decisions, but the delayed consequences of decisions nobody made when they were cheap.
Now the counterfactual, which is the arithmetic this lesson rests on. Same organizations, same portfolios, same unpromising ideas, with a gate in front of each. A third of those projects stop in month two rather than month fourteen, at roughly a tenth of the cost, before anyone announced them or fused their identity to them. The year ends with fewer initiatives and about the same number of successes, which is the part people miss: the items you kill at gate 0 and gate 1 were never going to become successes. Early killing reduces spend, elapsed time, and the resentment that makes the next proposal unfundable. MIT's finding that only about 5 percent of custom enterprise AI tools cross from pilot into production is not an argument for fewer experiments. It is an argument for making the other 95 percent end quickly.
The artifact: the Stage-Gate Charter
The named artifact you build here is the Stage-Gate Charter: four pages, published portfolio-wide, containing four things. The four gates, each with the evidence it demands and the numeric thresholds that evidence must meet. The forum: who sits, who presents, who decides, on what calendar. The four verdicts, standardized so nobody has to invent a status. And the kill ritual: memo, announcement, salvage handoff, scoreboard. It is published before the first intake, exactly as a tender's evaluation criteria are published before bids arrive, and for the same reason: criteria published afterward are not criteria, they are rationalizations.
The Four Gates and the Evidence Each One Demands
The gates map one to one onto the artifact chain this program has built since Level 1. That is the design's quiet elegance: you are not inventing governance, you are turning a method you already have into one. A well-run item passes with almost no extra work; a badly run one cannot pass at all.
| Gate | Decision | Evidence demanded | Cost of a kill here |
|---|---|---|---|
| Gate 0: Readiness | Does this enter the portfolio? | Selection Scorecard above threshold, audit readiness scores, dependencies mapped, mechanism stated in one sentence | Hours. Nobody is invested yet. |
| Gate 1: Design | Do we build and pilot? | Charter with pre-committed success and kill criteria, measured baseline, data conditions cleared or dated, sourcing decided, full-cost business case | Tens of thousands. Design only. |
| Gate 2: Pilot Evidence | Scale, iterate, or stop? | The Delta Table with its confidence words, checked against the criteria pre-committed at gate 1 | Low six figures. No rollout. |
| Gate 3: Scale Health | Does this keep running? | Quality regime counters, refreshed value measurement, adoption against the redesigned SOP, current run rate | Run cost, recovered by sunsetting. |
Gate 0: Readiness, the cheapest gate and the one that prevents the most waste
Gate 0 controls entry to the portfolio. Before an idea becomes an "initiative" with a slide and a name, it must clear four things: a Selection Scorecard result above the published threshold; the readiness scores its function earned in the enterprise assessment, because value cannot land on terrain that cannot hold it; a dependency map, particularly on the wave-0 foundations half the portfolio quietly assumes; and a mechanism sentence naming the specific step in a specific process, the number that step moves, and how. "AI will improve customer experience" is not a mechanism. "A classifier reads the inbound email, assigns one of nine intents, and routes it, removing the 40-second triage step and the 12 percent misroute rate that costs 1.1 days of rework" is a mechanism.
Here is the point most program offices miss. Most of your kills should happen at gate 0. It is the only place in the lifecycle where stopping is a genuinely low-emotion event, because nobody has built anything, told anyone, or attached their name to it. A portfolio whose gate-0 rejection rate hovers near zero is therefore not selective; it is a queue in a governance costume, and every idea that should have died in a fifteen-minute conversation will instead die at gate 1 or gate 2 having spent real money on the way. Expect to reject a third to a half of first-intake candidates, and expect the rate to fall as proposers learn what the scorecard rewards. That falling rate is the gate teaching the organization, which is the second thing a good gate does.
Gate 1: Design, the gate that catches the doomed-by-design item
Gate 1 is entry to build and pilot: the Level 2 readiness report's verdict turned into an institution. The forum checks five things exist and are real.
- Pre-committed success and kill criteria, numeric, dated, with the confidence standard named. Not "improved accuracy" but "routing accuracy at or above 92 percent, on a stratified sample of 400 tickets, at 'clear' confidence, by week 10."
- A measured baseline, not a planned one: computed, covering the required window including at least one month-end. The highest-yield check here, because a pilot without a baseline cannot produce provable value however well it performs.
- Data conditions cleared or dated, with a named owner and a date preceding the measurement window.
- A sourcing decision made, with the build, buy, or partner reasoning recorded rather than assumed.
- A full-cost business case, including what pilots habitually omit: integration, change and training effort, and the ongoing verification tax paid by whoever checks the AI's output forever.
A doomed-by-design item cannot succeed regardless of execution quality: no baseline, so no provable result; a data dependency that will not clear in time; an upside smaller than its verification cost. Such items are invisible in status meetings, because for months they look exactly like healthy ones. They are trivially visible at gate 1, which asks for the artifacts a doomed item cannot produce.
Gate 2: Pilot Evidence, and the property that makes it fast
Gate 2 is entry to scale. The item arrives with the Delta Table from Level 3: one row per metric, the like-window baseline, the pilot value with controls applied, the delta, the controls named, the surviving caveats in plain words, and a confidence word (clear, probable, or suggestive) assigned by rules agreed before anyone saw results. The forum places that table beside the gate-1 criteria and compares.
Now the crucial property. At gate 2 the decision was already made at gate 1. The forum is not deciding whether the result is good enough; that was answered months ago, in writing, by substantially the same people, before anyone knew which way it would fall. Checking evidence against a pre-commitment is a clerical act performed to a high standard, not a judgment call, and that is what makes gate 2 fast and unemotional. An eleven-minute gate 2 ending in a kill is not a callous meeting, it is a well-designed one. Everything that destroys gate 2 is relitigation ("the mix was unusual," "the second half was better," "the vendor has a new release"), and each of those belongs in exactly one place: the costed, dated iterate option presented alongside the other two. What none of them may do is move the threshold.
Gate 3: Scale Health, the gate everybody forgets
Gate 3 is the quarterly review of things that already scaled, and most programs never build it because scaling feels like an ending. Level 3 named the failure mode: set and forget, the deployed workflow that decays quietly while everyone celebrates its launch. Model quality degrades, the upstream process changes, the exception rate climbs back, users route around the redesigned SOP (standard operating procedure, the written definition of how the work is done), and value that was real in month three is gone by month eleven, unnoticed, because the item left the portfolio review and joined the run rate. So gate 3 demands four things quarterly: the quality regime's counters (drift indicators, exception rates, escalations), a refreshed value measurement against the original baseline, adoption measured against the redesigned SOP rather than the tool's login report, and the all-in run rate including licenses, support, and human verification load. Its verdicts include one the others lack: sunset, the orderly retirement of something that worked and stopped working. A zombie after scale is worse than one before it, because it carries a license, a support burden, and a false entry in the benefits register that inflates every future business case built on the same assumptions.
The Forum, the Four Verdicts, and the Threshold Rule
Gates are not documents. Gates are meetings that happen whether or not anyone feels ready, staffed by people with defined roles. Design the forum wrong and the four gates become four opportunities for advocacy.
Composition: the presenter is not the sponsor
- The portfolio owner chairs. They own the sequence and the aggregate outcome, not any individual item, which is exactly why they can call a verdict without losing anything.
- Finance sits permanently, ideally the analyst who recomputes the business cases, which turns "avoided spend" from a rhetorical flourish into a number somebody owns.
- Function sponsors attend for their own items: business context, absorbing the verdict, owning what follows. They do not present the evidence.
- The strategist (you, or the program office) presents the evidence, not as an advocate but as a presenter of findings, in the same neutral register whether the item passes or fails.
That last line is Level 3's measurement-independence principle promoted to governance level: the person who measures should not be the person whose reputation the number decides, so the presenter must not be the person the verdict rewards or punishes. When sponsors present their own items, every gate becomes an advocacy meeting, because that is what a reasonable human does when standing in front of a decision about their own work. They do not lie. They select: lead with the good chart, contextualize the bad one, describe a missed threshold as "trending toward." Independence is not a comment on integrity, it is an acknowledgment that nobody should prosecute their own case. Two seats are situational: the data or governance lead when an item carries dated data conditions or a regulatory classification, and specialist fairness review for any item in the people-decision category (screening, ranking, evaluating, or scheduling humans) before it can pass gate 1, per Level 3's bias discipline. Those items get an extra gate, not an exemption.
Cadence: gates are met, not requested
The most consequential sentence in the charter is about the calendar. Gate day is a standing monthly forum on a fixed date. Items are presented when the calendar arrives, not when the item feels ready. Fifteen minutes per item, agenda published a week ahead, evidence pack circulated five working days before. No pack, no slot, and the missed slot is itself a recorded event.
The alternative, scheduling a gate when an item is ready, makes readiness self-declared, and self-declared readiness arrives exactly when the evidence looks best. A fixed calendar inverts the pressure: an item that misses gate day waits four weeks, in public, with the reason in the minutes. Four weeks is a real cost and a survivable one. And it compounds: an item that must wait four weeks for a gate cannot drift for six months. Drift needs an absence of appointments; twelve appointments a year cap the time any problem can hide at one cycle. Fifteen minutes sounds brutal and is deliberately so: if a gate needs an hour, either the pack was not circulated, the thresholds were not pre-committed, or the forum is relitigating.
The four verdicts
- PROCEED. Evidence meets threshold, the next tranche is released, the next gate date is set.
- PROCEED-WITH-CONDITIONS. Threshold met except for named gaps that can close while work continues. Every condition carries a dated deadline, a named owner, and a check at a specified gate day, and a condition that misses its date converts automatically to a HOLD. Without dates and owners this is PROCEED with a decorative caveat, which is why it is the most abused of the four.
- HOLD. Work pauses pending specific missing evidence, with the re-gate date set now rather than "when it is ready." A HOLD names the artifact ("eight weeks of baseline including one month-end"), never the feeling ("needs more maturity"). It is distinct from a kill: the item may be excellent and merely unproven. It is also the verdict most vulnerable to abuse, because postponing without deciding is the path of least resistance for any group of humans. Therefore holds are capped at two per item; a third would-be hold is a kill, because an item that cannot assemble its evidence across three cycles is telling you something.
- KILL-THIS-VERSION. The item stops in its current form with a salvage list attached. The phrasing comes from Level 3: you are not killing the ambition, the problem, or the sponsor's judgment, you are killing this version, built this way, on this data, at this cost. A better version may re-enter at gate 0 whenever the world changes, and several should.
The evidence threshold discipline, and the anti-drift provision
This program has told you three times that criteria are set before the evidence arrives: as a checklist question in Level 1, as pilot design in Level 3, and now in constitutional form. The repetition is not padding; it is the third place where the whole structure fails without it. Thresholds are set at gate entry, never at gate review. When an item passes gate 1, the numeric criteria it must meet at gate 2 go into the record with the measurement method and confidence standard, and on gate day the forum's authority is to check evidence against those numbers, not to adjust them in light of the evidence. A threshold adjustable in the presence of results is not a threshold, it is a description.
Thresholds do legitimately change sometimes, so the charter permits change through one narrow door with a camera on it: any change requires the forum's approval as a distinct agenda item and is logged in the item's record with date, requester, old value, new value, and stated rationale. Five columns, permanently. Walk slowly through the pathology that log prevents, because it is how zombies are actually manufactured. Nobody proposes abandoning a success criterion. What happens is a sequence of small, individually defensible adjustments:
- Gate 1 sets routing accuracy at 92 percent.
- Week 6: "given the data-quality issues we found, 88 is a fair bar for this phase." Reasonable. Everyone agrees.
- Week 12: "the ticket mix in the pilot window was unusually hard, so 84 on this mix equals 88 on a normal mix." Plausible, possibly true.
- Week 18: "accuracy is not really the right measure anyway; what matters is agent satisfaction, which is strongly positive."
- Week 24: "the pilot is directionally positive and the team has learned a lot."
No single step is dishonest and each would pass a friendly conversation. The endpoint is an item with no measurable success criterion, no possible failure condition, and an indefinite life: the zombie, arrived at through five agreeable Tuesdays. The log defeats it with pure visibility. All five moves sit in one column in date order, and a chair reading "92, 88, 84, measure changed, narrative" out loud does not need to make an argument. The sequence makes it. This is the cheapest control in the charter and the one most likely to be quietly dropped when somebody edits the template.
The Kill Ritual: Making an Ending a Public Good
Everything so far is machinery. This is culture, and it decides whether the machinery survives its second year. A gate that produces kills in an organization that experiences kills as humiliations will be quietly disarmed within eighteen months, because humans route around pain. So the charter specifies how a kill is performed, the way an organization specifies how a safety incident is reported. Four mechanics that work as a set.
1. The kill memo, published portfolio-wide
Within a week of a kill verdict, a one-page memo goes to everyone with visibility of the portfolio. Not a private note to the sponsor, not a line quietly changed in a tracker. Five blocks: what we tested (mechanism sentence and design in three lines), what we learned (the Delta Table's honest result, including the parts that were genuinely good), what we keep (the salvage list with named recipients), what we spent (actual all-in cost, not the license cost), and what we saved by stopping now (the counterfactual, as arithmetic).
That last block changes how the memo reads: "Stopping at gate 1 cost $40,000 in design and baseline work. The same discovery at month nine, after build, integration, licensing, and the change program, would have cost approximately $380,000. Stopping now returned about $340,000 of avoided spend." Have finance compute it from the item's own approved business case (remaining committed spend plus run costs that would have been incurred, minus anything salvageable anyway), mark it an estimate, footnote the method. A number computed by finance is credible in a way that a number produced by the person recommending the kill can never be. And the arithmetic reframes the kill from loss into return: without it the memo says "we spent $40,000 and got nothing"; with it, "we spent $40,000 to avoid spending $380,000," a 9x return on an experiment, which is factually what happened.
2. The sponsor announces, not the strategist
The memo goes out under the sponsor's name and the sponsor says the words in whatever forum they normally use: the function town hall, the operations meeting, the weekly leadership call. Not the program office. Not you. Name the mechanism rather than gesturing at values. Leadership behavior teaches faster than policy, because employees update on what they observe happening to people, not on what they read in a deck. When a director says "we tested this, it did not clear the bar we set ourselves, so we stopped it, and here is what we kept," and that director is visibly not punished (still in the room, still funded, still leading, publicly thanked in the portfolio report), everyone watching learns something no psychological-safety training can teach: that here, stopping something is a normal professional act performed by senior people. One such announcement is worth ten memos about candor. For it to hold, the sponsor must get something concrete, and the charter should say what: their name on the avoided-spend number, first claim on the salvage, and priority consideration for their next item at gate 0. That is not a bribe but a rational allocation, since a sponsor who has stopped their own work on evidence is a better bet than one who has not.
3. The salvage handoff, with names attached
Every kill produces assets: redesign insights, the measured baseline, cleaned or newly documented data, the instrumentation, and people who now know how to run an AI pilot properly. Left in the memo those are rhetoric, so each salvage line is routed: asset, named beneficiary, date, and what they do with it. "Instrumentation retained: event counters stay live on the ticket queue, owner S. Aydin, feeding the service dashboard from next month." A salvage list without names is a consolation prize; a salvage list with names is a transfer of value, and everyone reading can tell the difference.
4. The portfolio scoreboard, where kills are a reported metric
Now the artifact that changes executive behavior. The quarterly portfolio report shows kills in the same typeface, on the same line, with the same prominence as scales:
This quarter: 2 scaled, 1 iterating, 2 killed at gate, $340k of avoided spend.
That line makes killing a countable output rather than an absence of output, and gives the executive committee a number to ask about other than "how many are green." It also sets up the diagnostic that eventually does the cultural work by itself: a portfolio reporting zero kills is reporting a broken gate. Either intake is so conservative that nothing interesting is being attempted, or the gates pass everything, and both are worse than the kills would have been. Executives learn to read the line quickly, and the quarterly question shifts from "why did this one fail?" to "what is our avoided-spend run rate, and is the gate-0 rejection rate holding?" A committee asking that has stopped treating kills as bad news, and at that point the individual courage this lesson opened by calling unscalable is no longer required. Nobody needs bravery to report a number leadership expects to see.
The Strategist's Ledger, grown up
Level 2 gave you the Messenger's Ledger: a private record of your consequential calls, their forecasts, and their outcomes, on the theory that an assessor's real authority is their forecast record. At portfolio scale it gains a column and answers the most dangerous question a program office ever faces, usually from a new CFO in month seven: "what does this function actually produce?" Log every gate-0 and gate-1 rejection with its finance-computed avoided-cost estimate, every gate-2 kill with its counterfactual, and every gate-1 prediction against the gate-2 outcome, so your forecast accuracy is measurable rather than asserted. After four quarters it reads: fourteen items rejected at gates 0 and 1, an estimated $1.9 million not committed, three scaled with measured value, two killed at gate 2 with $410,000 avoided, gate-1 predictions correct on nine of eleven. That is not a defense. It is a profit-and-loss statement for a governance function, and the only kind of answer that survives a zero-based review, which is the event that started Norvik's story.
Norvik's First Three Gate Days, and the Theater Next Door
All figures here are illustrative, built to show the shape of the arithmetic rather than to report a real company's results. Norvik's roadmap holds eleven items across four waves, with wave gates already in the sequence. Gate day is the second Thursday of each month, 90 minutes, chaired by the portfolio owner.
Gate day 1: the intake
Nineteen candidates arrive at gate 0, drawn from function wish lists, two vendor proposals, and the pipeline the readiness assessment surfaced. Eight are rejected in one session: five for no stated mechanism ("AI-powered forecasting," "intelligent document handling," "a chatbot for internal policy questions": asked which step in which process moves which number, none could answer in a sentence, and all five go back with the mechanism template and an open invitation to re-enter, which two later accept in better shape); two dependency-blocked pending wave 0, both waiting on the customer master data remediation, deferred with a named re-entry date rather than rejected on merit, because a deferral counted as a kill corrupts the scoreboard and demoralizes a sponsor who did nothing wrong; and one routed out of the portfolio, a proposal to score internal job applicants, which falls in the people-decision category and goes to specialist fairness review before it may occupy a gate slot at all.
Eleven items are admitted and become the portfolio. The rejection rate is 42 percent, which the chair notes is high and appropriate for a first intake, since a first intake is partly a backlog of accumulated wishes rather than a pipeline of designed items. Finance logs avoided spend for the five mechanism-less rejections only, at a conservative average of $95,000 each: $475,000, marked an estimate, method footnoted. The deferrals and the routed item are logged at zero, because deferring is not saving.
Gate day 3: one condition, one hold
Two items present at gate 1. The first is the portfolio's deep bet, a supply chain exception-handling workflow with the largest projected value and the longest dependency chain. Its discovery stage is complete: charter with numeric criteria, baseline measured across ten weeks, sourcing decided as a partner build. Two data gaps remain: field completeness on the vendor master sits at 71 percent against a required 90, and the data-sharing agreement with the logistics provider is unsigned. Verdict: PROCEED-WITH-CONDITIONS, both conditions owned and dated ahead of the measurement window and checked at gate day 5. The item keeps moving; the conditions keep their teeth.
The second is a quick win, an accounts payable coding assistant with modest but reliable value. Its charter is good. Its baseline is three weeks long, contains no month-end, and came from a spreadsheet export rather than the system of record. Verdict: HOLD, missing evidence named precisely (eight weeks of system-extracted baseline including one month-end), re-gate in four weeks, hold count one of two. Sit with what that HOLD replaced. In the Norvik of eighteen months earlier, this item appears in the monthly pack as green with an amber note reading "baseline work continuing in parallel," launches on schedule, runs four months, produces a genuine improvement, and proves nothing. The cost of the hold is four weeks. The cost of what it replaced is an unprovable pilot, which is the structural definition of the 95 percent. The sponsor was annoyed. Four weeks later the item passed with a real baseline, and at gate 2 it produced a number nobody could argue with.
Gate day 6: the eleven-minute kill
A wave-1 item, the customer service triage classifier, presents at gate 2. Its gate-1 pre-commitment was explicit: routing accuracy at or above 92 percent on a stratified sample of 400 tickets at "clear" confidence, handle-time reduction of at least 0.8 days, and a written kill condition stating that accuracy below 88 percent at gate 2 ends the version.
The Delta Table says accuracy 84 percent and handle time down 0.3 days at "suggestive" confidence, with a caveat that the pilot window's ticket mix was unusually complex. The presenter, who is you, reads it in four minutes without editorializing. Finance confirms spend to date at $118,000. The chair reads the gate-1 kill condition aloud. The sponsor says the mix caveat is real and asks whether an iterate option exists; it has been costed at $140,000 for twelve more weeks, with an honest note that the accuracy gap sits in intent categories the training data covers poorly and would likely close by roughly four points, not eight. The chair calls it: KILL-THIS-VERSION. Elapsed time, eleven minutes. Nobody in that room was brave. The decision was made at gate 1 by people who did not yet know how it would turn out, and the forum executed it. That is the whole design.
The following week the sponsor announces it at the service function's operations meeting and the kill memo publishes the same day: spend $118,000; counterfactual, computed by finance from the approved case, of roughly $358,000 had the item reached full rollout and failed there; avoided spend $240,000. Four named salvage lines: the event instrumentation stays live and now feeds the service dashboard, the eight-week handle-time baseline transfers to the service redesign team, the cleaned intent taxonomy goes to the knowledge base owner, and the queue restructure the redesign introduced (a change made to prepare for the classifier, not by it) is retained permanently and is worth an illustrative 0.6 days of cycle time on its own. The killed pilot therefore delivered three-quarters of the handle-time improvement its successful version was supposed to deliver, through pure process redesign, and Norvik keeps that forever. The quarter's scoreboard reads: 1 scaled, 1 killed at gate, $240k avoided.
Here is the finding that surprises people, stated explicitly because it is why this lesson exists: Norvik's program was measurably more credible after the kill than before it, through three unsentimental mechanisms. First, the CFO now believes the number on the item that scaled, because he watched the same instrument return a negative on a different item; a measurement system that has never reported a failure is not evidence of success, it is evidence of a broken instrument, and every experienced finance executive knows it. Second, sponsors saw that a kill is survivable, and gate-1 submission quality jumped the next cycle, with two items arriving with baselines nobody had asked for yet. Third, the board's question changed from "is this another scrap year?" to "what is the avoided-spend run rate?" That is not a softer question. It is a harder one, asked of a program the board has decided is real.
The failure story: gate theater
A peer organization installed stage gates the same year, with more enthusiasm and a better slide template. Then it made three quiet choices, none of which appeared in any minutes. Gates were scheduled when items were ready, which is sensible, respectful of teams, and fatal, because readiness became self-declared and every item arrived when its evidence looked best. Thresholds were called "guidelines," written in the charter and described at kickoff as indicative, and therefore never missed, since a guideline cannot be missed, only approached in a spirit of continuous improvement. Each item was presented by its own sponsor, efficient because they knew it best, and also meaning the only person with detailed knowledge of the evidence had a personal stake in the verdict.
Eighteen months later the numbers were remarkable: a 100 percent gate pass rate across roughly thirty reviews, and two successes out of nine items. Then the CFO ran a zero-based review and killed six items in one afternoon. Look at what the gates bought. The batch kill arrived exactly as it would have without gates, on roughly the same timeline, with the same acrimony, plus two extra costs: eighteen months of spend on items a working gate would have stopped in month two, and a governance apparatus (illustrative: about 340 person-hours of preparation and review) that produced no decision it did not also permit. Gates without pre-commitment, calendar, and independence are ceremony, and ceremony is more expensive than no governance at all, because it manufactures false comfort. An organization with no gates knows it is flying blind and stays nervous. An organization with theatrical gates believes it has controls, and that belief is what lets a portfolio run eighteen months without one hard conversation.
Three questions diagnose any gate process in under a minute, including yours. Who sets the gate date, the calendar or the item? What actually happened, in the record, the last three times a threshold was missed? Who reads the evidence out loud in the room? If the answers are "the item," "nothing," and "the sponsor," you do not have a stage gate. You have a recurring meeting with a serious name.
Your gates now decide which items proceed and which stop. But look at Norvik's two deferrals, the deep bet's dated data conditions, and that vendor master sitting at 71 percent completeness, and you can see what the gates keep pointing at. Gartner found that 63 percent of organizations lack or are unsure of AI-ready data practices, and forecast that through 2026, 60 percent of AI projects without AI-ready data will be abandoned. Most of your portfolio depends on foundations no individual item can build for itself. Chapter 4.3 builds them: the data program, the governance operating model, the platform, and vendor risk.
What to Do Monday Morning
- Write the four gates' evidence thresholds and publish them before your next intake. One page per gate, naming the specific artifacts (scorecard threshold, mechanism sentence, measured baseline, Delta Table, quality counters) and the numeric bars. Criteria published after proposals arrive are negotiable, and everyone knows it.
- Put gate day on the calendar as a standing monthly forum, fixed date, named chair, 15 minutes per item, five-working-day pre-read rule. Book the next six months today, because the calendar, not the charter, is what makes gates real.
- Separate the presenter from the sponsor, in writing, before the first contentious item, and state that the presenter presents findings in the same register whether an item passes or fails.
- Draft the kill-memo template with its counterfactual arithmetic block, and agree the avoided-spend method with finance now, while nothing is at stake. That method blessed in a quiet week is worth more than any argument you could win in a loud one.
- Add kills and avoided spend to the portfolio scoreboard you report upward, starting this quarter even if the number is zero, so the first "2 killed at gate, $340k avoided" reads as normal rather than as news.
- Open the threshold-change log and the Strategist's Ledger. Five columns for the first, four for the second. Both are empty today and both become your defense, and your evidence of value, within two quarters.
Key Takeaways
- Institutionalize the kill instead of relying on courage: one lead can end one pilot honestly, but eleven kills need a public mechanism, because bravery is scarce and the organizational default returns the moment it runs out.
- Read the 42 percent as a timing failure rather than excess discipline: those companies killed late, in a batch, after the money was spent, when gated items would have stopped in month two for a fraction of the cost.
- Build the four gates on artifacts you already produce: gate 0 readiness (scorecard, readiness scores, dependencies, mechanism sentence), gate 1 design (pre-committed criteria, measured baseline, dated data conditions, sourcing, full-cost case), gate 2 pilot evidence (Delta Table against the pre-commitment), gate 3 scale health (quality counters, refreshed value, adoption, run rate).
- Expect most kills at gate 0, where stopping costs only a conversation, and treat a near-zero gate-0 rejection rate as proof you are running a queue rather than a portfolio.
- Protect gate 2's speed by remembering its decision was made at gate 1: the forum checks evidence against a pre-commitment instead of relitigating it, which is why an eleven-minute kill signals good design rather than callousness.
- Fix the calendar and split presenter from sponsor: gates are met, not requested, an item that waits four weeks for a gate cannot drift for six months, and an item presented by its own advocate turns every review into a pitch.
- Log every threshold change with date, requester, old value, new value, and rationale, because zombies are manufactured by a sequence of reasonable downward adjustments that only becomes visible when the sequence sits in one column.
- Run the kill ritual as a set: published memo with counterfactual arithmetic, sponsor-led announcement that visibly costs the sponsor nothing, salvage handoff with named beneficiaries, and a scoreboard reporting kills and avoided spend beside scales, because a portfolio reporting zero kills is reporting a broken gate.
Skill.re