←
AI Readiness & Process Transformation
Visionary · M19 · lesson 19 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

The Kill Discipline at Enterprise Scale

15 min

The most important sentence spoken in Norvik Group's second Portfolio Sunset Review arrives eleven minutes in, from the person with the most to lose by saying it. The head of customer service, whose function is two years into a genuinely successful AI program, looks at the four tests on the screen and says: "Before we go round the table, I want to put the weekly reporting assistant up for sunset. It is mine, it is in my objectives, and I do not think it should still be running." Nobody argues, because she is right. Nobody is surprised either, because six months earlier this room watched the chief executive do exactly that to the initiative he announced at the January town hall, and then say why, calmly, with nobody's name attached to it. The first sunset review produced zero volunteers and a lot of careful language. The second produces three. The agenda did not change. The conduct of the most senior person in the room did.

Three Things That Need Killing, and Only One of Them Is Hard

You have been taught to stop things twice already, and both lessons hold. Level 3 gave you the decision memo for a single pilot: verdict first, evidence underneath, scale or iterate or kill, on criteria written before the pilot started. Level 4 turned that act of individual courage into machinery: four stage gates, a gate day in the calendar, a kill ritual, and a scoreboard reporting items scaled, items stopped, and money avoided.

Between them those lessons handle two of the three things an enterprise must be able to kill. Pre-launch ideas are easiest: a proposal that has not started has no team, no users, and no line in anyone's objectives, and Level 4's early gates stop these cheaply. In-flight pilots are harder but tractable, because a pilot that misses its pre-committed criteria has a written threshold it failed to clear, which is why you wrote the threshold down. A pilot that fails its criteria stops, and the criteria are not renegotiated in the meeting where they are applied. Then there is the third category, which nothing you have built so far will ever surface on its own.

The running system nobody is looking at

A running system is a workflow that passed every gate, scaled, entered the operating model, and has been producing value in production for eighteen months or three years. It is not failing in any way that raises an alarm; it shows as green. And it may still deserve to be switched off, for four ordinary reasons that arrive quietly.

  • Its value decayed. Volume shifted, a policy change put a manual step back into the flow, or the surrounding process improved and the AI step now saves less.
  • Its cost trajectory rose. Consumption grew with adoption, support load with users, maintenance with every model version and re-test cycle.
  • Its process changed underneath it. The report it feeds was retired, the product line it serves was sold, the regulation it was built around was rewritten. It does its job perfectly for a job that no longer exists.
  • Its economics inverted. The vendor restructured pricing at renewal, or what you built became a bundled feature in a platform you already pay for.

Now the structural point, and the reason this lesson sits at Level 5. Nothing in your governance apparatus is designed to notice any of this, because gates face forward. A gate is a question asked of something that wants to proceed, triggered by ambition, when an item requests its next tranche of money. A system that passed its gates two years ago is asking for nothing, so it never appears on a gate agenda again. Your committee reviews exceptions, your portfolio board reviews the pipeline, and your dashboard reports totals that still carry the decayed system's original claim. Every instrument you own points at what is arriving; the estate behind you is unobserved by design. So examining a running system is never automatic. Somebody has to choose to look, at something that is not complaining, in a way that can only produce work and awkwardness for whoever chose. That choice is the discipline.

Why a running system is fundamentally harder to stop

The difficulty is not analytical. It is political, and it reduces to one asymmetry: a running system has a constituency and a pre-launch project does not.

Count what accumulates around a workflow in two years of production. A named owner whose competence is publicly associated with it. Users who reorganized their day around it. A place in someone's annual objectives, probably the sponsor's. An appearance in a board deck as evidence the program works, so stopping it reads as an admission the deck was wrong. A support rota, a monitoring dashboard, a register entry, a compliance calendar line. It has become part of the furniture, and furniture is not evaluated. It is dusted. A pre-launch proposal has a sponsor and a hope, and that is the whole constituency. Which is why organizations with excellent gate discipline can still be quietly rotting: they are very good at the easy kill and have never performed the hard one.

The compounding is brutal. A zombie running system does not merely waste its own budget. It consumes a slot in the portfolio's WIP limit (work in progress: the cap on how many things can be in flight at once, enforced because attention, not money, is the scarce resource), a share of the support and monitoring capacity the whole estate draws from, and governance calendar, because it must still be audited like anything else. Most expensively, it consumes capable people who could be scaling something that works. Each of those costs grows as the estate grows, which is why accumulation is invisible in year one and close to fatal by year four.

An organization with stage gates can stop projects. An organization with kill discipline can stop things that are already working, already staffed, and already in the board deck.

The Artifact: The Portfolio Sunset Review

The artifact you leave holding is a standing meeting with an unusual scope. Gate day looks at everything proposed. The Portfolio Sunset Review looks at everything running. It is the only forum whose default question is "should this continue?" rather than "should this proceed?", and that inversion is what makes it work.

Cadence: semi-annual. Quarterly is too often (systems need time to produce trend data). Annual is too rare (a system with inverted economics burns twelve months of budget first). Scope: every AI-touched system in the register, without exception, because the ones most likely to have quietly decayed are exactly the ones nobody would nominate. Attendance: the portfolio owner in the chair, the function heads who own the systems, the value measurement lead, someone from risk who can speak to control burden, and the program's most senior sponsor personally, every time.

Prep pack: one page per system, produced by the program office and not by the system's owner, from data the program already holds: value measurement against original claim, twelve-month cost actuals, incident history, usage trend, control obligations, and the date of the last honest re-measurement. The owner reviews and corrects it; the owner does not author it. Who owns the narrative determines whether a review is real.

The four tests

Every system gets the same four questions in the same order. Value comes first, because a system with persistent value survives a mediocre answer on any other test, and a system without it fails regardless of how elegant it is.

TestThe question it asksFailure signature
1. Value persistenceIs the original benefit still present, re-measured inside six months on the original baseline method?"It is definitely still saving time", with no measurement since the scale-up case
2. Cost trajectoryWhat does it cost to run now, from twelve months of actuals, and which way is that moving?A cost line estimated once, at launch, never reconciled to actuals
3. Strategic fitDoes the process it serves still exist, and has a better instrument appeared?Nobody has looked at the market since the build-versus-buy decision
4. Control burdenWhat governance and compliance overhead does it carry, and is that proportionate?A low-value system carrying a heavy control regime because it was classified once

Value persistence reuses the Level 3 improvement loop and this level's persistence check as a stopping instrument rather than a reporting one. The most common finding in a first review is not a scandal: it is a system whose value decayed by a third and whose owner had not noticed, because nothing in their week required them to look.

Cost trajectory carries the point operators trained on traditional systems get wrong. Conventional software has a decaying cost profile; AI-touched systems have a rising one. Consumption scales with use, support scales with users and with the ambiguity of outputs, and maintenance is continuous, because every model version, prompt, schema, and policy change triggers a re-test of the evaluation harness and, for anything sensitive, a re-run of human review sampling. That is how the economics invert without anybody deciding anything. Ask for the direction of travel, not the level.

Strategic fit has two halves. A remarkable number of AI systems outlive the workflow they were built for, because process change and decommissioning are owned by different people who do not meet. And Level 4's annual sourcing review applies to internally built systems too, the ones nobody re-tests: MIT found externally partnered solutions succeeded roughly twice as often as internal builds, and only about 5 percent of custom tools crossed from pilot into production. Being overtaken is not a failure of the team; it is a market working.

Control burden catches the enterprise phenomenon with no small-company equivalent: the forgotten system nobody uses much and everybody must still audit. An oversight step, a quarterly evidence pull, an access recertification, a model change notification, and four hours a quarter of an owner's time, against usage of eleven transactions a month. Worse than the wasted cost, a low-attention system with real permissions is exactly the profile that produces an incident nobody was watching for.

The three outcomes

  1. CONTINUE. The four tests pass. The system stays, having cost forty minutes and gained a fresh re-measurement. Most systems land here, which is a good sign rather than evidence the review is soft.
  2. REMEDIATE. One test failed for a fixable reason. The outcome is never "keep an eye on it": it is a named fix, a named owner, and a dated re-review, normally inside ninety days, held as a short separate session. A remediation without a date is an escape hatch with better manners.
  3. SUNSET. The system stops, and the verdict is only credible with a decommissioning plan behind it.

Decommissioning properly

A sunset is a project, not an announcement. Give it five elements and six to eight weeks. The process reverts or moves: decide whether the work returns to the prior manual method, transfers elsewhere, or stops entirely, and write which, because a sunset that leaves the work homeless produces a shadow process within a month. Data is handled per retention policy: retention schedule, legal hold check, documented disposal or archival. The audit trail is preserved: decisions the system made while live stay auditable as long as retention requires, because regulators do not accept "we turned it off" as an answer about last year's decisions. Users are told well in advance: named date, named alternative, named person to ask. People are redeployed, visibly and immediately: the subject of the next section.

Get this wrong and the damage exceeds the zombie you removed. A botched sunset does more cultural harm than a system you should have stopped and did not, because it teaches one sticky lesson: being associated with a stopped initiative is dangerous. Once people believe that, every owner becomes an advocate and every review becomes defensive. The mechanics of stopping are easy. The choreography is the job.

The Reallocation Reflex

Here is where kill discipline usually collapses, for a rational reason: in most organizations stopping something is a pure loss for everyone in the room. The function loses budget it will not get back, the owner loses scope, the team loses a project and gains uncertainty, and the only beneficiary is the corporate centre, abstractly, later. Under those incentives honest people find honest reasons to continue.

Money and capacity return to the portfolio envelope, not the general fund. This is the funding rule from earlier in this level applied to sunsets rather than gate kills: run cost removed from a decommissioned system goes back into the envelope from which the next tranche is released. Say it out loud in the review, with the number: "this sunset returns roughly 40,000 a year to the envelope, which funds most of the integration work on the item at gate three." A function head who watches their released budget fund something visible, rather than vanish into a corporate saving, behaves differently next time. That is not generosity; it points the review's incentives at its purpose.

The freed people go somewhere visibly better, immediately. When an initiative stops, almost nobody reads the decision memo. Everybody notices what happened to the two people who were working on it. That observation, not your governance charter, is what the organization actually learns. Two outcomes, two permanent lessons:

  • The team is assigned within a week to the highest-profile scale-out in the portfolio, in the same communication as the sunset. Lesson learned: working on something that gets stopped is not a career event, and the people who ran it are trusted with the best work we have.
  • The same team is left "between assignments" for a quarter while somebody works out where they fit. Lesson learned: get attached to something that cannot be stopped. That one is learned once and never unlearned, even if every word of your communication was supportive, because people believe placements over prose.

So make the placement decision before the review: part of the prep pack for any system heading toward sunset is a named landing place for its people. If you cannot answer where they go, you are not ready to decide, because you are about to teach something you did not intend to teach.

The Leader's Personal Conduct

Everything above is procedure, and by now your organization has plenty of procedure. Here is the uncomfortable truth that puts this lesson at Level 5: the procedures will be quietly bypassed unless the most senior person in the room repeatedly and personally demonstrates that stopping is safe. Bypass does not look like defiance. It looks like a review where every pack has a slightly optimistic value line, every marginal case lands on REMEDIATE, and the meeting finishes early with nothing stopped. Everyone has behaved correctly and the forum has become theatre, within two cycles, because nobody senior ever paid a personal price for stopping something. Five behaviours prevent it, and they are actions rather than personality traits.

Kill your own initiative first, and publicly

The first sunset the organization sees should be yours. Find the initiative most closely associated with you, the one you announced, and apply the four tests to it in front of the same forum that will apply them to everyone else. If it fails, stop it and say why in plain words. If it passes, publish the evidence in the detail you would demand of anyone else. The mechanism is simple and slightly brutal: the sponsor who stops their own pet project has purchased the right to ask anyone else to stop theirs. Until then, every request you make asks someone else to take a risk you have not taken, and function heads detect that asymmetry easily. Their response is not rebellion but polite, well-evidenced, indefinite delay. It is why this lesson opened as it did: the head of customer service volunteered because she had watched the chief executive go first, and that cost had already been paid by someone senior enough to make it safe.

Report kills in the same breath and the same tone as wins

Level 4 gave you the scoreboard: items scaled, items stopped, money avoided, one table. At enterprise scale that table must be spoken, by the leader, not merely tabulated by the program office. There is a real difference between a kill count on slide fourteen and a chief executive opening a quarterly review with "we stopped two systems this half, here is what they were costing us, and here is where that money and those people went." Tone is the payload. Kills reported apologetically, buried in an appendix, or prefaced with "unfortunately" teach the room that kills are bad news being managed. Kills reported in the register used for wins teach that a documented stop is an outcome the program produces on purpose. That rule has followed you since Level 1 (a documented kill is a win, a zombie is a career risk); this is its final scale, demonstrated by the most senior leader in a room that includes the board.

Never let a kill be described as a failure of the people who ran it

This is a language discipline, enforced in real time, in the meeting, by the most senior voice present. Every sunset has a cause and it is almost never the competence of the team. Name the actual one: the design was wrong (we built the right thing for a process we had misread), the readiness was absent (the data foundation was not there, which we now know and have priced), the market moved (what we built is now a bundled feature in a platform we already pay for), or the process changed. Each is true, specific, non-personal, and teaches something reusable. "It did not work out" teaches nothing and invites everyone to fill the vacuum with a name.

State the exception precisely, because leaders who over-apply this rule become incapable of managing. If a system genuinely failed through execution, that is a private performance conversation in the normal management process, not a public narrative attached to a governance decision. Mixing the two contaminates both.

Thank the person who surfaces bad news, before asking anything about it

When someone tells you a running system has stopped delivering, your first response is not a question. It is thanks, said out loud, in front of whoever is present. The questions come afterward, and they are about the system rather than the person's judgment in raising it. The reason is mechanical: bad news about a running system is voluntary information, and nothing in your process would have produced it that day. If the messenger's first experience is interrogation, everyone watching updates their estimate of what surfacing bad news costs, and next time you will get the same information from a dashboard, six months later, after the money is spent.

Refuse the three escape hatches, by name

Rooms that cannot bring themselves to stop something do not say so. They reach for one of three devices, all of which sound like good governance. Name them to your leadership team before the review, and have the closing sentence ready.

Escape hatchWhat it sounds likeWhat it isThe sentence that closes it
The indefinite extension"Let's give it another six months to bed in."A decision to continue disguised as a decision to wait, with no test attached."Happy to extend. What measurement, by what date, would make us stop it? If we cannot name one, we are not extending, we are continuing."
The descope-and-rename"Let's narrow it to the core use case and relaunch it as the Insights Hub."A zombie under new branding: costs intact, constituency reassured, clock reset."Then it is a new item and goes to gate one with a fresh baseline and business case. The existing system sunsets on a date we set today."
The revisit-next-quarter"Let's park this and revisit next quarter."The one that repeats forever, because next quarter's agenda is set by next quarter's pressures."Parked with a date, an owner, and the evidence we will bring, or decided now. I am content with either, but not with this conversation again in April."

None of the three refuses the request; each converts a vague continuation into a dated, evidenced commitment. The leader does not have to be the person who says no. The leader has to be the person who will not accept an undated yes.

Worked Example: Nine Systems, and the Estate That Ate a Program

Norvik Group is the 2,400-person business-to-business services and distribution company this level has followed: six functions, a scrap year behind it, now a governed operating model with a portfolio, a funding envelope, and several scaled workflows in production. Every figure below is illustrative. Nine running systems are in scope at the second Portfolio Sunset Review, which takes two hours and twenty minutes.

Six continue: value re-measured, costs reconciled, process current, control burden proportionate. Two get their first honest re-measurement in over a year, and one comes in slightly above its original claim. The review is not a hunt.

Two remediate. A supplier-onboarding assistant whose measured value fell about 35 percent after an upstream change to the vendor master feed added a manual verification step nobody connected to this system's performance. Named fix: restore the automated feed mapping, owned by the data steward, re-review in ten weeks. And a document-extraction workflow whose vendor restructured pricing at renewal from a flat per-seat licence to per-document consumption, roughly doubling annual run cost against an unchanged benefit. Named fix: renegotiate with a volume-tier commitment and price two alternatives in parallel, owned by the sourcing lead, re-review in twelve weeks. Both carry dates. Neither becomes "we will keep an eye on it."

One sunsets. A year-one reporting assistant that drafted commentary for a weekly operations report. The report was retired five months earlier in an unrelated process change, and nobody connected the two facts, because the person who retired the report worked in operations and the person who owned the assistant worked in the program. It had been running, unused but fully governed, for five months at an illustrative 1,900 dollars a month in licence and consumption, plus a quarterly compliance review costing half a day of a risk analyst's time. Nobody did anything wrong; there was simply no mechanism that would have asked the question.

The sunset takes six weeks. Week one: decision published, the four users notified, alternative named. Weeks two to four: data handled per retention schedule, prior outputs archived to the standard evidence store, register entry updated to decommissioned with the decision reference. Weeks five and six: access removed, contract terminated at the notice date, closure note filed with actual costs removed. The two-person capability is assigned, in week one and in the same announcement as the sunset, to the highest-priority scale-out in the portfolio. Visibly. Within a week.

The quarterly business review three weeks later opens with it. Not slide fourteen: slide two, spoken by the chief executive in the tone used for the quarter's two scaled workflows. What was stopped, why, what it cost, where the money and the people went. Then the year's cumulative reallocation as one line: an illustrative 340,000 dollars of avoided spend from stage-gate kills plus 84,000 dollars of annualized run cost removed by sunsets, both returned to the portfolio envelope and visible as funding for named items now in flight.

The result that matters is measured indirectly. At the following review, three function heads arrive having nominated their own systems as sunset candidates before the packs circulated. The first review produced zero. Nothing in the process changed. What changed was that the organization had watched, in sequence, a chief executive stop his own initiative, a team from a stopped system land on the best work in the portfolio inside a week, and a sunset reported in the same register as a win. No memo produces that and no charter substitutes for it.

The failure story: the accumulation

Now the other path, more common and harder to see from inside, because every decision along it is defensible. A large professional services firm builds an impressive AI estate over four years. Its gate discipline on new work is excellent: real intake, real gate day, documented kills, a scoreboard. It has, however, no mechanism for retiring old work. Nobody decided against having one; the question never came up.

By year four it operates 23 AI-touched systems. A belated internal review sorts them: 7 delivering their original value measurably; 9 delivering materially less (decayed benefit, changed process, drifted usage, risen cost); 4 essentially unused and fully governed, audited and monitored and used by almost nobody; and 3 with no identifiable owner, because the named owner left and the system quietly became nobody's.

The maintenance, support, monitoring, and compliance burden of those 23 systems consumes roughly a third of the AI function's capacity, which explains a fact leadership has discussed for two years without solving: portfolio throughput has been falling while headcount has been rising. New items take longer to reach a gate, scale-outs slip, and the best engineers spend their week on re-tests, evidence pulls, and incident triage for systems nobody would fund today. The diagnosis, in good faith, across three planning cycles, is "we need more delivery capacity." They hire, and throughput falls again, because the new capacity is absorbed by the sixteen systems delivering less than they should. The problem is not capacity. The organization only ever added.

Set that against the research. S&P Global found 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before, and Gartner expects over 40 percent of agentic AI projects to be cancelled by the end of 2027. Nobody scraps most of a portfolio in an orderly semi-annual review, so those numbers are mass late kills: a budget crisis, one afternoon, decided by someone with the authority to be blunt, taking the good systems out with the dead ones because nobody in the room can tell them apart. McKinsey's 88 percent using AI regularly against only about 39 percent able to attribute any EBIT impact, mostly under 5 percent, describes exactly the estate above, and MIT's 95 percent is the same story at pilot stage. One sentence holds it: gates without sunsets produce an estate, and an estate eventually eats the program that built it.

With the estate honest and the capital freed, a different question becomes available. Everything so far has been about doing existing work better and cheaper. The final lesson of this chapter asks what AI makes possible that was not possible before: new products, new services, and new business models.

What to Do Monday Morning

  1. List every running AI-touched system, with its owner and the date of its last honest value measurement. Use the AI system register, and expect the list to be incomplete and the dates worse than you feared. A system whose last measurement predates its scale-up is not measured; it is remembered.
  2. Put the first Portfolio Sunset Review in the calendar with a date, a chair, and an invitation list, then run the four tests (value persistence, cost trajectory, strategic fit, control burden) against every system. Have the program office write the packs.
  3. Decide in advance where the people from any sunset will go. If you cannot name a landing place for a system's team, you are not ready to decide about that system, because the placement is what the organization learns from.
  4. Kill or publicly re-justify your own most-cherished initiative, in front of the same forum, first. Until you have, every request you make of a function head asks for a risk you have not taken.
  5. Name the three escape hatches to your leadership team before the review: the indefinite extension, the descope-and-rename, and the revisit-next-quarter. Give the closing sentence for each and let everyone recognize them the moment they appear. A hatch named in advance is much harder to walk through.

Key Takeaways

  • Distinguish the three things that need killing: pre-launch ideas (stage gates handle these), in-flight pilots (pre-committed kill criteria handle these), and running systems, which nothing notices because gates face forward and a system that passed its gates two years ago is never re-examined unless somebody looks.
  • Recognize the asymmetry: a running system has a constituency (owner, users, objectives, a board-deck mention, a support rota) while a pre-launch proposal has only a sponsor and a hope.
  • Install the Portfolio Sunset Review as a standing semi-annual pass over everything running, scoped to every system in the register, with packs written by the program office rather than the owners.
  • Apply the four tests in order: value persistence (measured recently, not remembered), cost trajectory (AI systems have a rising cost profile, so economics invert without anyone deciding), strategic fit (does the process still exist, has the market commoditized it), and control burden (low value plus heavy controls is a hidden tax and an unowned exposure).
  • Close every case with CONTINUE, REMEDIATE with a named fix and a dated re-review, or SUNSET with a decommissioning plan covering the process, data retention, audit trail, user notice, and people.
  • Build the reallocation reflex so freed money returns to the portfolio envelope rather than the general fund, and freed people move within a week to the most visible work available, because organizations learn from placements, not memos.
  • Practise the leader's five behaviours: stop your own initiative first and publicly, report kills in the tone used for wins, name the non-personal cause while keeping execution issues private, thank the messenger before asking a question, and refuse the three escape hatches with a dated alternative.
  • Remember the arithmetic of accumulation: gates without sunsets produce an estate, an estate consumes the capacity that would otherwise scale what works, and organizations that only ever add end up scrapping most of a portfolio in one afternoon.