←
AI Readiness & Process Transformation
Visionary · M14 · lesson 14 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Scaling What Works: From Function to Enterprise

15 min

The steering meeting is going well, which is when the dangerous sentence arrives. The exceptions workflow at the Manchester shared-services site has held its numbers for two straight quarters: backlog down from nine weeks to four, verified extraction accuracy inside its threshold, a named owner who reports the process without mentioning AI. The COO looks at the chart, looks at you, and says the thing every transformation leader waits for and then dreads: "Good. Roll it out to the other five." Someone adds that it is basically a copy exercise. And you know what the room does not: the workflow on that chart is not software. It is software plus a person, a transformer who sits forty feet from the team, fixes broken documents before anyone files a ticket, and knows which two suppliers send scanned faxes, none of which is written down. Move the software to Site 2 and you move half the system. MIT put a number on how that ends: of the custom AI tools enterprises build, roughly 5 percent cross from pilot into production. This lesson is about the other 95 percent, and about the five things that must exist before a verified win travels.

The Hothouse and the Climate

Start with the honest reason the 5 percent figure is so brutal. It is not that scaling is technically hard: the technology usually travels fine, copied to a second site in an afternoon. The attrition happens in the gap between the conditions a pilot enjoys and the conditions an enterprise actually provides. A pilot runs in a hothouse. The glass is controlled: a hand-picked team who volunteered, a transformer present and personally invested, a scope narrowed to the friendliest slice of the work, daily attention from someone whose reputation depends on it working. Under those four conditions almost anything grows. Workflows survive their first bad week because somebody stays late, and edge cases get handled by a human who happened to be sitting there, so quietly that nobody logs them as edge cases.

An enterprise is a climate. Nobody controls it. Site 4 has users who did not volunteer, working a process variant the pilot never saw, on documents in a format the pilot never met, with a champion who is also the shift supervisor and has forty minutes a week. The author is nine hundred miles away. At four o'clock on a Friday, in a timezone the pilot never considered, somebody hits an error message and has ninety seconds to decide whether to wait, guess, or go back to the spreadsheet. That ninety seconds is your adoption number for the next year.

So here is the frame this lesson rests on. Scaling is not repetition. It is a different engineering problem: making something that worked under supervision work without it. The pilot answered "can this work?" The scale-out answers "can this work when nobody who understands it is in the room?" Treat the second as a logistics exercise and you spend a year producing fragile copies of one good thing.

You have met a smaller version of this. At practitioner scale, the replication playbook taught what transfers from a first process to a second (the discipline: baselines, the verification standard, human gate design) and what must be rebuilt each time (the specifics: data, exception taxonomy, local rules). That still holds. But taking one verified win to many sites or geographies adds four problems that replicating to a single second process does not have: the workflow must survive without its author (hardening), destinations legitimately differ (variant management), somebody must answer questions forever (the support model), and every destination arrives at a different readiness on a different date.

The artifact for that job is the Scale-Out Pack: the five things that must exist in writing before a verified win goes anywhere else.

  • The hardened workflow. The version that survives the absence of its author.
  • The variant policy. The written split between what is fixed everywhere and what each site may set locally, within bounds.
  • The support model. Who answers the question at four o'clock on Friday, within what time, funded from which budget line.
  • The local-readiness gate. The per-site check deciding whether this destination receives the workflow now or after remediation.
  • The rollout sequence. The order of sites, chosen for readiness and learning value, with a pause built in.

Hardening: The Unreachable-Author Test

"Pilot-grade" is not an insult. It is a correct engineering choice: a pilot should be cheap, fast, and only as robust as it needs to be to answer its question. The failure is not building pilot-grade but shipping it to six sites and calling that a rollout. Hardening is hard to plan, though, because the missing parts are being supplied invisibly by a human being. You need a diagnostic that makes the invisible visible, and one question does it better than any checklist:

What would break if the person who built this were unreachable for two weeks? The honest answer is your hardening backlog.

Ask it with the pilot team and one skeptical operator present, and enforce two rules: no answer may include a name, and "we would just call them" is not a mitigation. What comes out lands in five predictable categories, each of which you should hunt deliberately.

  • Error handling for cases escalated to a person who happened to be sitting there. In the pilot, a malformed document or a timeout produced a shrug and a manual fix at the next desk. There is no next desk at Site 4. Every silent human catch must become an explicit path: retry, route to a defined queue, or fail loudly with a message telling the user what to do next. The tell is that nobody logged it as an incident, because it never felt like one.
  • Configuration hard-coded to one site's reality. Date formats, currency, supplier naming conventions, the two document layouts that dominate at the pilot site, a threshold typed in during week three. Anything a second site could legitimately differ on must move into declared configuration, or Site 2's team will do it themselves, badly, in a copy of the code.
  • A runbook for the failure modes the transformer handled personally. Not the happy path (the pilot always has that) but the ugly list: the queue empty when it should not be, accuracy down for two days, a source system changing a field, output that looks confidently wrong. Each entry needs a symptom, a first action, and an escalation point.
  • Monitoring that was a person watching a dashboard. A human noticed drift because they looked every morning. At six sites nobody looks. Monitoring must alert a named recipient against a defined threshold, with someone owning the response. If the answer to "how would we know?" is "we would notice," you have found a hardening item.
  • Documentation that assumed its author was reachable. Pilot documentation is written for people who already understand the workflow; enterprise documentation is written for a supervisor at Site 5 inheriting it in month fourteen with no context. Test it the only way documentation can be tested: give it to someone who was not there and watch.

Then triage. A typical backlog runs eight to twelve items, of which only three or four genuinely block travel. The test: would this item, unresolved, cause a user at another site to lose trust or revert to the old method? Cosmetic gaps can ship and close in flight. Trust-reverting gaps cannot ship, because trust reversion is the failure that does not come back.

Budget it honestly. Five to seven weeks between a verified pilot and the second site is normal, and it is the cheapest five weeks in the program: the alternative is paying for the same discovery at every site, in front of users, on their first impression. It is also the step momentum kills fastest, being the only phase where work happens and no site goes live, so name it on the plan as a phase with a deliverable. An unnamed pause gets compressed; a named phase gets defended.

The Variant Policy: What Is Fixed, What Is Local

If hardening is the most commonly skipped component, the variant policy is the most consequential, because it quietly defines the whole scale-out: how much local adaptation is permitted. Most programs never decide it explicitly. They discover it eighteen months later, as an unpleasant inventory. There are two ways to get it wrong, both with a recognizable shape.

Pole one: rigid standardization

One configuration everywhere, no exceptions, usually announced with the words "consistency" and "one way of working." It is administratively beautiful and it fails wherever local process genuinely differs, which is everywhere. You met the pure form in the process-maturity work: a retail group's board, impatient with pilots, mandates a uniform AI rollout of one procurement workflow across all twelve regions at once. That produces not consistency but two behaviours: workarounds (the site uses the tool for the 60 percent of cases it fits and runs a shadow spreadsheet for the rest, invisible until an audit) and refusal (the site logs in for the compliance metric and works the old way). Both look like adoption on a usage chart. Neither produces value, and both are correct responses by competent operators to a tool that does not fit the work in front of them.

Pole two: unconstrained localization

The opposite instinct, usually born from the first one's scar tissue: let every site configure what it needs. It feels respectful and adaptive, and it produces eleven variants nobody can support, upgrade, or measure comparably. Support cannot answer a question without first establishing which version they are looking at, a fix found at Site 7 cannot be pushed to Site 3 because Site 3 has drifted, and the numbers stop being addable because each site measures a slightly different thing. Unconstrained localization does not scale one pilot to eleven sites; it quietly recreates eleven separate pilots, on one budget with one team.

The working answer: a declared core-and-variable split

The answer is neither pole but a written division, decided before the second site configures anything, governed by one principle that settles most arguments: anything the enterprise measures or governs must be fixed; anything that only affects how local work flows can be local, within stated bounds.

Fixed everywhere (the core)Local, within declared bounds
Exception categories and reason codes (the taxonomy every site reports against)Queue routing and work allocation inside the site
The verification standard: what a human checks, when, against whatNumeric thresholds, within a range set centrally
The data schema and field definitionsLanguage, templates, and user-facing wording
The audit trail: what is logged, retained, reviewableLocal approval chains and named approvers
Escalation rules: what goes to a human, and to which roleWorking hours, shift patterns, local scheduling

Three details make this work. First, bounds are numeric and published: a confidence threshold may be set between 0.82 and 0.94 by the local owner, not "adjusted as needed." A request outside the bounds is a change request against the core, decided centrally, which is exactly the conversation you want: a site pushing hard on a bound is usually telling you something true about the process that the core got wrong.

Second, the split must be written and owned before the second site starts. In the absence of a written policy the split still exists: it is decided by whoever configures first, unilaterally, and inherited by everyone after them as an accident nobody chose and nobody can now change without breaking Site 2's numbers. Third, every variable a site sets goes in one register, so "what is Site 5 running?" is a lookup, not an investigation.

The Support Model, the Readiness Gate, and the Sequence

Who answers at four o'clock on Friday

Every scale-out meets this question late, usually in week three of the second site, as an escalation to someone senior who should not be in the loop. Answer it in advance with a tiered model, each tier carrying a named population and a written response expectation that does not have to be impressive, only true.

TierWhoHandlesResponse expectation
1Local champion at the siteHow do I do this, where is the button, is this normalSame shift
2The function's super-userProcess judgment, exceptions, local configuration within boundsOne working day
3The central capability teamDefects, core changes, cross-site patterns, data issuesTwo to three working days, triaged same day
4The vendor or platform ownerPlatform faults, model behaviour, contractual issuesPer the contract, tracked centrally

Now the part programs get wrong, which is a resourcing truth rather than a design truth. Support demand per site does not fall to zero after go-live. It falls to a plateau. The curve is steep for six to eight weeks, then flattens and stays flat: new joiners, seasonal document types, an upstream system change, a novel exception, a supervisor asking whether a number can be trusted. Call the plateau three questions per site per week (illustrative, but the right order of magnitude). At one site that is background noise. At twelve it is thirty-six a week, permanently: a standing cost with a real name on it.

Here is the mechanism that kills delivery functions, and it is the most predictable consequence of successful scaling. Uncosted, the plateau does not disappear: it is absorbed heroically by the delivery pod, who know the answers and will not leave a site stranded. Six months later the pod built to deliver new use cases spends most of its capacity answering questions about old ones, throughput collapses, and leadership concludes the team has lost its edge. It has not. It has been converted into a help desk by its own success, one helpful answer at a time. Cost the plateau, name its owner, and put it in the base budget, not the project budget, because it outlives the project by years.

Readiness is a property of the site, not the solution

This inverts how rollout plans are usually built. Plans treat readiness as something the solution achieves once (it passed its pilot, so it is ready) and treat sites as interchangeable destinations. The reality is the opposite: the solution's readiness is settled, and every destination has its own. So each site gets a miniature readiness check before it receives the workflow, on five things.

  1. Data quality for this use case. Not the site's data maturity in general, but the fields this workflow reads, at this site, sampled. The pilot's accuracy figure was earned on the pilot site's data.
  2. Process-variant compatibility. Does this site's process fit inside the declared variables, or does it need something the core does not offer? If the latter, that is a core change request with a date, not a go-live surprise.
  3. Champion coverage. A named Tier 1 person with time actually allocated, plus a backup. One champion with no backup is not coverage; it is a single point of failure with a holiday booked.
  4. Training completion. Not scheduled. Completed, by the people who will do the work, including the verification standard, the part users skip and governance depends on.
  5. A captured baseline. The site's own cycle time, volume, error rate, and cost per unit, measured before go-live.

The baseline is non-negotiable per site, and it is the check most often waived under time pressure with "we already know what good looks like from the pilot." You do not. You know what good looked like at one site. Without a local baseline you cannot prove this site improved or compare sites honestly, and you have reproduced, six times over, the original sin that put 95 percent of pilots on the wrong side of the measurement line. A site that cannot produce a baseline is not showing you an administrative gap; it is showing you that it cannot see its own process, which is itself a reason to delay.

And the discipline that makes the gate real: a site that fails gets remediation and a later slot, not an exception. Exceptions granted at gates are how gated programs become ungated programs. Remediation is usually small (two weeks of data cleanup, a second champion, one training cohort) and a later slot costs a fortnight. Going live at an unready site costs that site's trust, a multi-year asset.

Sequencing for learning, and the pause nobody defends

Rollout order is usually set by two bad criteria: geography (it looks orderly on a slide) or politics (a regional director asked loudly and early). Sequence instead by readiness and by learning value, and make one deliberate choice that pays for itself repeatedly.

Choose the second site for difference, not for ease. The instinct is to pick a site resembling the pilot, because it will go smoothly and produce a clean second data point. That instinct is exactly wrong. An identical second site teaches nothing, and worse, appears to validate a variant policy that has never been tested, so the first genuinely different site (Site 5 by then, with momentum, a published plan, and four live sites on the core) becomes where you discover the split was wrong. At Site 2 that costs weeks; at Site 5 it costs a rework of everything behind you. Pick the site that differs in the dimension you are least sure about: a smaller operation, a merged role structure, a different regulatory context, language, or document mix. Its job is to stress the variant policy while stressing it is cheap.

Then pause. After site two or three, build a two to three week absorption window into the plan with no site going live, whose deliverables are: fold the learning into the core, update the variant policy and runbook, re-cost the support plateau against real demand, re-sequence the rest. This is the highest-return decision in most rollouts and the hardest to defend, because by then the program has momentum, the plan has dates, and pausing looks like doubt. It is not doubt. It is the difference between rolling out one workflow six times and rolling out six workflows once each.

Variance, Owners, and Two Rollouts

Cross-site variance as a diagnostic

Each site carries its own baseline and delta, which is the floor. The ceiling is the comparison, the one diagnostic instrument no single-site pilot can own: when the same hardened workflow, on the same core, produces materially different results at different sites, the variance is information. Not noise, not a verdict on the sites, but a signal that something differs in the process, the people, or the data.

So investigate the worst-performing site first, seriously, with the central team on site, treating the visit as research rather than intervention. The finding is very often upstream and general: someone enters a field differently, a supplier sends a different file, a step happens out of order, a cohort was taught an older version. Fix it and you frequently improve three other sites that had a milder version of the same problem. That compounding is what makes scale-out worth doing, and it is the effect McKinsey finds in the high performers: the roughly 6 percent reporting meaningful bottom-line impact are about three times more likely to have fundamentally redesigned workflows rather than layered tools onto them.

Hence the warning. Do not publish cross-site variance as a league table. The moment site performance becomes a ranking with names on it, sites manage the number instead of the process: exceptions handled off-system, awkward cases excluded, problems no longer travelling upward. You will have destroyed your best diagnostic instrument to produce a slide. Report variance with the investigation attached ("here is what Site 4 taught us about intake") and sites bring you problems instead of hiding them.

The owner who prevents six orphans

All of the above is worthless in eighteen months without one structural decision, whose absence is the most common scale-out failure at that distance. The capability transfers to a permanent owner at each site (accountable for that site's numbers and local variables) and to a single global process owner accountable for the core: taxonomy, verification standard, schema, audit trail, bounds, and the change process governing them.

The global owner runs the improvement loop centrally, so a fix found at Site 7 reaches Sites 1 through 6 as a versioned change rather than a rumour. Without that role, drift is not a risk but a certainty with a schedule: each site improves locally, each improvement is reasonable, none is shared, and within a year twelve sites are twelve orphans running twelve versions of one former capability, with no comparable measurement and no upgrade path. The program can still point at twelve live deployments. It cannot point at one capability, and that distinction separates a transformation from a collection of installations. It is also why Gartner's forecast that over 40 percent of agentic AI projects will be cancelled by the end of 2027 reads less as a prediction about technology than about ownership.

Worked example: six sites in nine months

Numbers here are hypothetical, chosen to show the shape of a disciplined scale-out, not to promise results. The exceptions workflow is verified at the pilot site and approved for five more. Instead of a copy exercise with a three-month date, the program runs the Scale-Out Pack.

  • Hardening: 5 weeks. The unreachable-author test produces 9 items; 3 are load-bearing: the silent human catch on malformed scans (now an explicit exception route), a hard-coded supplier-name mapping (now configuration), and monitoring that was the transformer's morning glance (now a threshold alert). The other 6 close in flight.
  • Variant policy: declared in week 3. Core fixed: categories, verification standard, schema, audit trail. Four local variables with published bounds: queue routing, confidence threshold (0.82 to 0.94), user-facing language, approval chain.
  • Site 2, chosen for difference. Not the sister site with the same systems, but the smallest operation, where one person holds a merged clerk-and-supervisor role. In three weeks it surfaces two variant-policy gaps: the escalation rule assumed two distinct roles, and the approval chain needed a self-approval bound with a compensating audit control. Fixable in days at Site 2. At Site 5, with four sites live, they would have meant reworking the core and re-validating everything behind it.
  • Pause: 2 weeks. No site goes live. The core absorbs both changes, the runbook gains four entries from the first real tickets, the sequence is reordered.
  • Readiness gates: 6 assessed, 4 pass. Two go to remediation: one for data quality (a source field populated inconsistently for eight months), one for champion coverage (one champion, no backup, three shifts). Remediation takes 3 and 5 weeks; neither gets an exception, and both go live later without incident.
  • Support: costed, not absorbed. Demand plateaus at roughly 3 questions per site per week by week 12, about 18 a week across six sites and a projected 36 at twelve. It enters the base budget as a named standing cost, which is what lets the pod start two new use cases during the rollout instead of none.
  • Variance investigated, not ranked. At month six the worst site sits well below the others. The central team visits and finds an upstream data-entry practice at intake producing a field variant the extraction step handles poorly. Fixing it lifts that site sharply, and three others with a milder version.
  • Month nine. Six sites live, cycle-time improvement holding within about 15 percent variance across sites (the number that says one capability rather than six installations), one global process owner, six local owners, one variant register.

The discipline cost about 7 weeks of delay plus two remediation slots. It bought a supportable estate, comparable numbers, an improvement loop, and a pod still delivering.

The failure story: the pilot that could not travel

A manufacturer builds a quality-inspection AI at one plant, and it is genuinely excellent. Over eight months it holds its accuracy, inspectors trust it, defect escape rate drops measurably, and the plant manager becomes its loudest advocate. On that record the board approves a rollout to four more plants as a copy exercise on a three-month timeline, and the pilot team, flattered, agrees.

Plant 2's cameras are mounted at a different height and angle; the models underperform immediately and the local team, wanting to help, adjusts thresholds themselves. Plant 3's inspectors use a different escalation practice, informal but real, that the workflow does not accommodate, so they build a parallel logging sheet and use the tool for part of the job. Plant 4's network cannot carry the integration, so a contractor builds a batch upload. With no variant policy, none of these improvisations is wrong exactly: each is a competent person solving a real problem with no guidance about what they were allowed to change.

By month nine there are four incompatible variants, no two plants measure defect escape the same way, and support falls entirely on the original pilot team, who now triage four codebases instead of starting anything new. The improvising plants resist standardization, correctly noting that their version works for them. The program has spent a year and its credibility producing four fragile copies of one good thing.

The autopsy is short. The pilot was excellent and was never hardened: it worked because a plant manager, an engineer, and a fixed camera rig were in the same building. No variant policy, so "what may we change?" was answered four times, locally, on four different Tuesdays. No support model, so the delivery capability was consumed. No readiness gate, so a plant whose network could not carry the integration was scheduled anyway. No global owner, so nothing learned at one plant reached another. The difference between a hothouse and a climate was paid for four times, at full price, in a currency the program could not afford: the belief of four plant managers who now know that "corporate AI" means a tool that half works and a team that does not answer.

Set that against the wider record and the pattern stops being anecdotal. MIT's 95 percent, S&P Global's 42 percent of companies scrapping most AI initiatives in 2025 (up from 17 percent the year before), McKinsey's gap between 88 percent using AI and roughly 39 percent seeing any EBIT impact: none is a story about models failing. They are stories about the 70 percent of BCG's 10-20-70 rule (10 percent algorithms, 20 percent technology and data, 70 percent people and process) arriving unfunded at the scale-out stage.

What to Do Monday Morning

Five moves, in order, each doable in a fortnight.

  1. Run the unreachable-author test on your best pilot. One hour, the pilot team plus one skeptical operator, one question: what breaks if the builder is unreachable for two weeks? Hunt all five categories (error handling, configuration, runbook, monitoring, documentation), then mark each item load-bearing or not with the trust-reversion test. That list is your hardening backlog, and it goes on the plan as a named phase with a deliverable.
  2. Declare the core-and-variable split in writing, before the second site configures anything. Two columns, published, owned by a named person, numeric bounds on every variable. When you argue about a row, apply the principle: anything the enterprise measures or governs is fixed. If you do not write it, it will still be decided, by whoever configures first.
  3. Cost your support plateau and put it in the base budget. Questions per site per week at steady state, times site count at full rollout, converted to a standing role or fraction of one, in the operating model rather than the project. Then say it to your sponsor: unfunded, this gets absorbed by the delivery pod and new-use-case throughput goes to roughly zero within two quarters.
  4. Choose your second site for difference, not for ease. Rank destinations by how much they differ from the pilot in the dimension you are least confident about, pick near the top, and tell the steering committee why: this site's job is to break the variant policy while breaking it is cheap. Then put the two-week absorption pause on the plan right after it, before anyone can argue about dates.
  5. Name a global process owner before the third site goes live. One person accountable for the core, the change process, and the improvement loop across all sites, with local owners underneath. If you cannot get that name agreed, you have learned whether this rollout is producing a capability or a collection of installations.

Key Takeaways

  • Treat scaling as a different engineering problem, not repetition: a pilot runs in a hothouse (motivated team, present author, forgiving scope, daily attention) and an enterprise is a climate, which is where MIT's roughly 5 percent pilot-to-production rate is decided.
  • Build the Scale-Out Pack before a verified win travels: the hardened workflow, the written variant policy, the funded support model, the local-readiness gate, and a sequence with a deliberate pause.
  • Run the unreachable-author test across error handling, hard-coded configuration, the runbook, monitoring, and documentation, then ship only once the trust-reverting items are closed.
  • Declare a core-and-variable split with numeric bounds, fixing anything the enterprise measures or governs, since rigid standardization produces workarounds and refusal while unconstrained localization recreates eleven separate pilots.
  • Cost the support plateau into the base budget instead of absorbing it heroically, because demand per site falls to a plateau rather than to zero, and that plateau converts a delivery pod into a help desk.
  • Gate every destination separately on data, process fit, champion coverage, training, and a local baseline, sending failures to remediation with a later slot rather than granting exceptions.
  • Sequence for learning by choosing the second site for difference rather than ease, then pause two to three weeks to absorb before committing the rest of the estate.
  • Read cross-site variance as a diagnostic and never as a league table, and name a global process owner plus local owners early, because without one twelve sites become twelve orphans within a year.

Scaling multiplies what works. A portfolio only compounds, though, if what does not work is stopped with the same discipline, and stopping is far harder at enterprise scale, because by then the thing has sites, owners, budget lines, and defenders. That is the kill discipline at enterprise scale, and it is next.