←
AI Readiness & Process Transformation
Visionary · M8 · lesson 8 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

From Pilots to Operating Model: The Transformation Playbook

15 min

Two years in, the program is working. Nine use cases have been through the gates, three are scaled and running in production, two were killed early with the evidence written down, and the last quarterly scorecard showed a defensible number in the value column for the first time. The transformation lead has earned the room. Then, in a corridor after the review, the chief financial officer (CFO) says something friendly that lands like a diagnosis: "You know, if you took the other job, I honestly don't know who would run any of this." He means it as a compliment. It is not one. It is the most accurate description anyone has offered of what has actually been built: not a capability, but a person doing an excellent job, surrounded by artifacts that only work while she is there to work them. This lesson is about the conversion that fixes that, and it is the last hard thing this program will ask you to do.

The Wall Every Good Program Hits

Every successful AI program eventually meets the same wall, and it is never technical. The tools work by then. The method works. Evidence accumulates, the portfolio delivers, the skeptics have gone quiet. The wall is that all of it is still a program: a temporary structure, funded annually out of an initiative line, staffed by people on loan from their real jobs, dependent on a handful of individuals who carry the method in their heads, and structurally unable to outlive its sponsor's tenure or the next reorganization. It is not fragile because it is bad. It is fragile because that is what programs are: scaffolds, and scaffolds come down.

You already know the headline statistic of this whole field: MIT's 2025 study found that roughly 95 percent of enterprise generative AI pilots delivered no measurable profit-and-loss (P&L) return, and that only about 5 percent of custom-built tools ever crossed from pilot into production. That finding has an organizational twin that nobody has put a number on, and you are going to spend Level 5 preventing it: most AI programs never cross from initiative into operating model. They run, they produce real value for two or three years, and then they dissolve, leaving a shared drive full of frameworks nobody maintains and a scaled workflow or two quietly decaying because the improvement ritual that kept them honest stopped happening when the person who chaired it moved on.

McKinsey's 2025 State of AI numbers describe the same divide from the outside. Eighty-eight percent of organizations use AI somewhere. Only about 39 percent can point to any earnings-before-interest-and-taxes (EBIT) impact from it, and among those, most report the impact at under 5 percent. The high performers, roughly 6 percent of the sample, are about three times more likely to have fundamentally redesigned workflows rather than layering tools onto existing ones. Read that finding structurally rather than as a project fact: you cannot fundamentally redesign workflows, repeatedly, across a large organization, on the strength of a temporary team and an annual budget argument. Sustained redesign is what an operating model produces. The 6 percent are not people who ran better pilots. They are organizations that made the redesign capability permanent.

What changes at this level

Everything up to now has taught you to produce outcomes: Level 2 assessment, Level 3 one process end to end, Level 4 the organizational program (audits, portfolio, data and skills foundations, change plan, governance charter, value scorecard). All of that was you doing the work well. The Level 5 question is different and colder. Not "what did we deliver?" but "what would still function if every person currently doing it were replaced tomorrow?" That is the operating-model question, and it is the only one whose answer survives a promotion.

The conversion has a precise shape. Four capabilities exist in every functioning AI program, whether or not anyone has named them: discovery (finding and qualifying what to do next), delivery (turning a qualified candidate into a running, scaled operation), governance (deciding what is allowed, gated, and accepted), and value tracking (proving what the whole thing produced). In a program, all four are projects: done when someone asks, staffed by whoever is free, documented in whatever the last person used. In an operating model, all four are standing functions, each with a permanent owner, a funding source in the base budget, a cadence on the corporate calendar, defined inputs and outputs, and defined handoffs to the other three. The transformation leader's job, from here, is that conversion: turning four temporary capabilities into four permanent ones, so that the organization's ability to absorb AI stops being an initiative and becomes a property of how it operates.

The Four Capabilities: As a Project, as a Function

The before-and-after framing carries this entire lesson, so take each capability twice: what it looks like while it is still a project, and what it looks like once it has become a function. The tell in each case is the same kind of thing, and it is worth noticing the pattern: a capability has become a function when it keeps operating without anyone deciding to operate it.

Discovery

As a project: a use-case workshop gets run when someone senior asks for one. Facilitators are borrowed, a list of thirty ideas is produced, six are scored, two get funded, and the other twenty-four sit in a spreadsheet until the next workshop generates a list that heavily overlaps the old one. Discovery happens in bursts, driven by demand from above.

As a function: there is a standing intake with published criteria that anyone in the enterprise can submit to, and a scored pipeline that always has qualified candidates ready for the next unit of delivery capacity. Demand does not arrive only from executives; it flows in continuously from the shadow-AI census (the inventory of tools people are already using unofficially) and from the friction streams your champion network and process audits generate. Candidates are scored on the value-by-readiness logic from Level 4 before anyone is asked to fund them. The owner is usually the readiness or transformation lead. The tell that discovery has become a function is that intake keeps running when nobody is asking for it. Chapter 5.2 builds this in depth as the use-case engine; here you only need it as one of the four boxes on the blueprint.

Delivery

As a project: a team is assembled per pilot, mostly out of goodwill and the sponsor's ability to borrow people, and disbanded when the pilot ends. The method lives in the heads of whoever ran the last one. Every new effort renegotiates its own staffing from scratch, which means the negotiation itself, not the work, sets the pace.

As a function: there is a standing delivery capability whose standard operating procedure (SOP, the documented way a task is performed) is the Level 2 to Level 3 method: baseline, charter, redesign, pilot with pre-committed kill criteria, scale, embed. Transformer is a named role with a job description and a career path (typically from process analyst into transformer into process owner or portfolio lead), not a title someone wears while on secondment. And the capability has a stated capacity that the portfolio plans against rather than negotiates for.

The operating unit is the pod: a small cross-functional team (a transformer, a process owner or delegate, a data or systems person, and part-time subject-matter time from the function) that carries one use case from charter through to scaled operation, then hands it to its permanent process owner with the runbook, the metrics, and the improvement ritual attached. A pod is not a project team, because it does not dissolve into nothing; it re-forms around the next candidate with most of its learning intact.

Here is the insight that reorganizes how you think about the whole portfolio: delivery capacity is the binding constraint on the entire enterprise's AI throughput. Not ideas, of which there are always too many. Not budget, which follows evidence. Not technology, which is bought. The number of use cases your organization can carry from charter to scaled operation in a year is set by how many pods you can staff with people who know the method, and treating that as a permanent capability with real headcount rather than a rolling series of secondments is the single thing that raises the number. BCG's 10-20-70 rule (10 percent of the effort in algorithms, 20 percent in technology and data, 70 percent in people and process) is usually quoted at project level, but it is truer here: the 70 percent is where your capacity lives.

Governance

As a project: a committee is formed for the AI initiative. It meets when the initiative has something to review, its authority derives from the sponsor's authority, and its calendar is the program's calendar. When the program pauses, so does it.

As a function: the Level 4 governance charter, decision rights, and stage gates operate on their own calendar, independent of any single program, with the regulatory register maintained continuously rather than refreshed in a panic. The register matters more each year: European Union AI Act general-purpose AI (GPAI) obligations have applied since August 2, 2025, transparency duties for AI-generated content arrive December 2, 2026, high-risk Annex III obligations December 2, 2027, and embedded Annex I products August 2, 2028. Those dates will land on whoever holds the function, and they do not care whether your program is between sponsors.

The tell that governance has become a function is that a new AI use case anywhere in the enterprise automatically enters governance without anyone deciding to involve them. Procurement routes it. The project intake form triggers it. The architecture review flags it. When involvement depends on someone remembering that the AI committee exists, governance is still a project, and the shadow estate is already growing outside its view.

Value tracking

As a project: a business case is written to get approval, and a claim is made at the end to declare success. Between those two documents, nothing is measured on a schedule. Baselines are collected once, live in a slide, and are lost when the deck is superseded.

As a function: the Level 4 value scorecard is produced on a quarterly cadence by a named owner with measurement independence, meaning the person who reports the numbers does not report to the person whose success the numbers describe. Baselines are preserved and versioned in a place with an owner. Value is re-proven on cadence for systems that were scaled two years ago, because a scaled workflow can decay quietly and the scorecard is what catches it.

This is the capability most often left as a project, and its absence is why organizations cannot answer the most damaging question they will be asked. Three years in, a new executive says: "What has AI actually done for us?" If value tracking is a function, the answer is a document produced last quarter by someone whose job it is. If it is still a project, the answer is a scramble through old decks, and the scramble itself is the answer everyone hears. S&P Global found that 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before. Some of those deserved scrapping. Some worked and could not prove it.

CapabilityAs a projectAs a standing functionThe tell
DiscoveryA workshop when someone asksStanding intake, published criteria, scored pipeline fed by census and friction streamsIntake runs when nobody is asking
DeliveryA team per pilot, disbanded afterPods, the method as SOP, transformer as a career role, stated capacityThe portfolio plans against capacity instead of negotiating for people
GovernanceA committee for the initiativeCharter, decision rights, gates and regulatory register on an independent calendarNew use cases enter governance automatically
Value trackingA business case, then a claimQuarterly scorecard, independent owner, versioned baselines, value re-proven for scaled systemsLast quarter's numbers exist without anyone chasing them

The Five Conversion Moves

Knowing the destination is not the same as knowing the mechanism. Each capability crosses from project to function through the same five moves, and a capability that has had three of them done to it is not two-thirds converted; it is still a project with better documentation. Do all five, one capability at a time.

1. A permanent owner with the capability in their objectives. Not "also doing AI." The distinction is testable: open the person's objectives for the year and see whether the capability is named there with a measure attached. If it appears only in their manager's verbal understanding of what they do, the capability is one reorganization or one busy quarter away from stopping. Permanent owners can be part-time (0.2 of a full-time equivalent, or FTE, is a real answer for governance coordination), but they cannot be informal.

2. Budget in the base. Move the capability's funding from project money to run-rate: from the initiative line that gets re-argued every planning cycle to the operating budget that gets carried forward unless someone actively cuts it. This is the single clearest signal that an organization has decided the capability is permanent, because it changes the default from "prove it again" to "justify stopping it." Lesson 3 of this chapter builds the multi-year funding logic in full. For now, note the asymmetry: a project budget dies of neglect, and a base budget survives it.

3. Documented method that survives its authors. The artifacts you built in Levels 2 through 4, the assessment instruments, the charter template, the baseline pack, the gate criteria, the scorecard definitions, become the capability's operating procedures: versioned, owned, dated, and reviewed on a schedule. This is the same discipline the SOP lesson taught one level down, applied to your own function rather than to a business process. The test is uncomfortable and worth running: hand your method to a competent person who has never met you and see whether they can run a gate review from it. If they need you in the room, you have documentation, not a method.

4. Cadences on the corporate calendar. Gate day, the quarterly value scorecard, the governance meeting, the improvement rituals for scaled workflows. Put them where the quarterly business review and the budget cycle live, with the same booking discipline and expectation of attendance. A rhythm on the corporate calendar continues through leadership changes; a rhythm that exists because someone enthusiastic sends the invite stops the first month that person is on holiday during a crisis. Half of what people call program decay is just meetings that stopped being scheduled.

5. Interfaces defined. This is the move most often skipped and the one that quietly undoes the other four. Each capability hands off to the others at specific moments, and each handoff needs a stated trigger, a stated artifact, and a stated receiver:

  • Discovery to delivery, at gate 1. What discovery hands over (a scored candidate with a baseline, a named process owner, and a draft charter) and what delivery commits to in return (a slot, a pod, a date). Without this, discovery hands over enthusiasm and delivery receives an argument.
  • Delivery to the process owner, at scale. The runbook, the metrics definitions, the exception paths, the improvement ritual, and an explicit acceptance moment. Without this, the pod never really leaves, and your capacity silently drops as scaled work accumulates.
  • Everything to governance, continuously. Automatic entry at intake, at gate, at deployment, and at material change. Not a courtesy notification.
  • Everything to value tracking, on the quarterly clock. Baselines at charter, realized value at scale, re-proof on cadence afterward, in the scorecard's format rather than in prose.

Undefined interfaces are exactly where standing functions leak back into project behavior. When a handoff has no trigger, someone has to notice that it is time, and "someone noticing" is the operating logic of a project. Every undefined interface is a place where the operating model quietly reverts to depending on a person's attention.

The Honest Sequencing, and the Blueprint

An organization cannot stand all four capabilities up at once, and should not try. Attempting the whole conversion in one planning cycle produces four half-built functions and a credibility problem, because half a function behaves exactly like a project while costing like a commitment. The order that works, and why, is the most practical advice in this lesson.

First: governance and value tracking. They are the cheapest to convert, often a fraction of an FTE each plus a calendar entry, because the artifacts already exist from Level 4 and the work is mostly cadence and ownership rather than headcount. More importantly, they make everything else legible. Governance makes the estate visible, including the parts you did not know about. Value tracking makes the results provable. Convert these two first and every subsequent conversion argument gets easier, because you are now arguing from a scorecard rather than from conviction.

Second: delivery. This is the expensive one, and it is expensive in the way organizations find hardest: it requires permanent headcount, trained people, and a career path that human resources has to actually create. It is also the one that takes longest to show, because a transformer takes two or three use cases to become genuinely good. Start it once you have a scorecard to fund it with.

Last: discovery. This ordering surprises people, and it is the one to hold firm on. A discovery function is useless until delivery can absorb what it finds. Open a standing intake against a delivery bottleneck and you manufacture exactly two things: a backlog and a reputation for asking people to submit ideas that nothing happens to. The second is expensive to repair, because the friction streams and the census depend on people believing that reporting friction leads somewhere. Keep discovery informal and demand-driven until delivery has capacity to plan against, then open it.

The artifact: the Operating Model Blueprint

Here is what you take from this lesson: the Operating Model Blueprint, one page, four rows. Each capability gets five columns: its owner (a named person and role, not a function), its funding source (base or project, stated honestly), its cadence (the calendar entries that make it real), its key inputs and outputs, and its handoffs to the other three. It is the one page that describes the end state your transformation is actually building toward, and it doubles as a diagnostic: fill it in for today, honestly, and every cell you cannot complete is a conversion you have not made yet.

Two rules keep it useful. Write the current state beside the target state in the same table, so the gap is visible rather than aspirational. And review it twice a year with your sponsor: the blueprint's real audience is the executive who funds the conversion, and a table showing three "project" entries in the funding column is a more persuasive budget argument than any narrative you can write.

Eighteen Months of Conversion: A Worked Example

Return to the enterprise from Level 4: a 2,400-person business-to-business services and distribution company, six functions, one scrapped year of AI spending in its history, and now a running portfolio with governance that works. All figures below are illustrative and rounded, chosen to show the shape of the arithmetic rather than to predict yours.

Months 1 to 3: governance converted. The Level 4 charter already existed, so conversion was mostly formal: the compliance manager took the coordinator role as a named objective at 0.2 FTE, gate day and the monthly governance meeting moved onto the corporate calendar alongside the quarterly business review, and procurement's intake form gained a mandatory AI question that routes automatically. Cost of conversion: about 0.2 FTE and two conversations with procurement. Within a quarter, three tools the program had never heard of entered governance through the procurement trigger. That was the return: not a decision, but visibility that used to depend on gossip.

Months 3 to 6: value tracking converted. A 0.5 FTE analyst was funded, and the structural choice that made everything credible was where the role reported: into finance, not into the transformation program. Measurement independence is not a personality trait, it is a reporting line. The analyst owned the baseline library (versioned, with a retention rule), produced the quarterly scorecard, and re-proved value for the two workflows scaled the previous year. That re-proof found one running at about 70 percent of its claimed benefit because an exception path had grown, which the improvement ritual then fixed. A program would never have looked.

Months 6 to 15: delivery converted, and this was the hard one. Two transformer positions were created as permanent roles, funded from the base, with a written career path from process analyst through transformer to process owner. One was filled internally from the champion network, one hired. The pod model formalized what had been improvised: each pod paired a transformer with borrowed function capacity, roughly 0.6 FTE of process-owner and subject-matter time per active use case, plus systems support from the existing technology team. The measurable result of treating delivery as a capability rather than a series of secondments: concurrent capacity rose from 3 efforts to 5, and the median time from charter to gate 1 fell because the negotiation for people stopped happening. Nine months, two headcount, and a job architecture, for a 67 percent throughput increase.

Months 15 to 18: discovery converted, last and deliberately. The standing intake opened only once delivery could absorb it. Its first quarter's pipeline arrived pre-filled: eleven candidates from the friction streams the champion network had been logging, six from the shadow-AI census, four from the process audits, and only three from executive requests. The 0.2 FTE of coordination sat with the transformation lead. Nobody ran a workshop to fill it, which is precisely what it means for discovery to have become a function.

CapabilityOwnerFundingCadenceKey outputs
DiscoveryTransformation lead (0.2 FTE)BaseMonthly pipeline reviewScored candidate list, gate 1 packs
DeliveryDelivery lead, 2 transformers (2.0 FTE) plus borrowed pod capacityBase and function timeWeekly pod stand-up, monthly gate dayScaled workflows, runbooks, handover packs
GovernanceCompliance manager (0.2 FTE)BaseMonthly meeting, continuous registerDecision log, risk register, regulatory calendar
Value trackingFinance analyst (0.5 FTE, reports to finance)BaseQuarterly scorecardBaseline library, realized-value report, re-proofs

The arithmetic a CFO can accept. The conversion moved roughly 2.9 FTE into the base budget, about 276,000 dollars a year at an illustrative fully loaded 95,000 per FTE, on top of the borrowed function capacity that was already being spent informally. Against that, the portfolio was delivering an illustrative 412,000 dollars of realized run-rate value, verified by an analyst who does not report to the program, with capacity for five concurrent efforts instead of three. That is not a spectacular ratio and it does not need to be. It is a defensible one: a standing cost with an audited return, in the format the rest of the business uses for functions it intends to keep. Once a transformation capability can be discussed in those terms, the argument stops being about whether AI works and becomes how much capacity to fund, which is the only budget conversation worth having.

The failure story: the program that stayed a program

A mid-sized manufacturer ran a genuinely excellent AI program for two years under a capable transformation director. Eleven use cases through the gates. Real, measured value. Governance that made sensible decisions. A trained cohort of five people who knew the method and had used it. By any project standard, an unambiguous success, and it was recognized as one: the director was promoted to run operations for a different division.

The program was folded into the information technology (IT) portfolio "for efficiency." Nobody made a bad decision. Nobody cut anything. Here is what four quarters looked like. Gate day stopped, because it had no owner and no calendar entry outside the director's. The value scorecard's last edition is now nine months old, because the analyst who produced it reported into the program and was reassigned when the program's line disappeared. Three of the five trained transformers returned to their home functions or left, because transformer was never a role, it was a thing they were doing. The improvement rituals lapsed, and two scaled workflows are decaying quietly, one of them producing exceptions that a supervisor now handles manually without reporting them: that is how a scaled AI workflow reverts to a manual process without a single meeting acknowledging it. And in the fourth quarter, a newly appointed executive asked for an AI strategy, unaware that one had existed and worked.

The artifacts are all still on the intranet. The charter, the gate criteria, the scorecard template, the baseline library, the eleven case write-ups. Unmaintained, undated, and correct. Nothing failed here except permanence. Note that this manufacturer was never in the 95 percent; it did the hard work and got the value. It simply built that value into a person rather than into the organization, and then the person was promoted, which is what happens to people who do excellent work.

A program is a person's achievement. An operating model is an organization's property. Only one of them survives a promotion.

Gartner's forecast that over 40 percent of agentic AI projects will be canceled by the end of 2027 is usually read as a warning about immature technology and vendor exaggeration, and partly it is. But sit with the manufacturer's story and you will see a second mechanism hiding inside that statistic: a meaningful share of cancellations are not technical failures at all. They are perfectly good work that lost its owner.

What to Do Monday Morning

The conversion starts with an honest diagram, not a proposal. Do these five in order.

  1. Draw your four capabilities on one page and mark each one "project" or "function" today, with evidence. Evidence means a named owner with the capability in their written objectives, a funding line you can point to, and a calendar entry that exists. Not "we basically do that." Most organizations two years in mark three or four as projects, which is the normal starting position.
  2. Name a permanent owner for the two cheapest conversions. Governance and value tracking, in that order. Ask for a fraction of an FTE and a named objective, not a headcount requisition. This is a conversation with two managers, not a business case.
  3. Get one capability's funding moved from project to base budget in the next planning cycle. Pick the one with the clearest running cost and the least ambiguity, usually value tracking. Base-budget status is the signal everything else follows, and the first one is by far the hardest.
  4. Put the four cadences on the corporate calendar: gate day, the quarterly scorecard, the governance meeting, and the improvement rituals for every scaled workflow. Book them twelve months out, in the same system that holds the quarterly business review, with named chairs and named deputies.
  5. Write the interface between discovery and delivery before you open intake. One paragraph: what discovery hands over at gate 1, what delivery commits to in return, and what happens to candidates delivery cannot take. If you cannot write that paragraph, your intake is not ready to open, and opening it anyway will cost you the credibility of your friction streams.

An operating model needs sustained executive backing across multiple budget years, and that means one narrative that holds up in three very different rooms: the executive team, the board, and the finance function that has to carry the base-budget line. That is the next lesson.

Key Takeaways

  • Recognize the wall every successful AI program hits: it is not technical, it is that the program remains a temporary structure with annual funding, borrowed staff, and a dependence on individuals that cannot survive a sponsor's promotion or the next reorganization.
  • Read MIT's finding organizationally as well as technically: only about 5 percent of custom tools cross from pilot to production, and most AI programs never cross from initiative to operating model, dissolving into artifacts nobody maintains.
  • Name the four capabilities that exist in every AI program whether or not anyone owns them (discovery, delivery, governance, value tracking) and diagnose each as a project or a standing function using the tell: a function keeps operating when nobody decides to operate it.
  • Treat delivery capacity as the binding constraint on enterprise AI throughput, and raise it by making transformer a permanent role with a career path and the pod the standing unit, rather than repeatedly negotiating secondments.
  • Apply all five conversion moves to each capability: a permanent owner with it in their objectives, budget moved into the base, documented method that survives its authors, cadences on the corporate calendar, and interfaces defined between the four.
  • Sequence the conversion honestly: governance and value tracking first because they are cheap and make everything legible, delivery second because it needs headcount and a career path, discovery last because a standing intake feeding a delivery bottleneck manufactures backlog and distrust.
  • Build and maintain the Operating Model Blueprint, one page with owner, funding source, cadence, inputs and outputs, and handoffs for each capability, written in current state beside target state so the remaining conversions are visible to the executive who must fund them.
  • Remember the manufacturer whose excellent two-year program dissolved in four quarters after its director was promoted: nothing failed except permanence, and permanence is the only thing this level is about.