←
AI Readiness & Process Transformation
Proficient · M25 · lesson 25 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Why Paving the Cow Path Fails: Redesign, Don't Overlay

15 min

Stand at the corner of a certain street in downtown Boston and you can watch modern life bend around a decision made by a cow. The street curves for no visible reason, narrows where no engineer would narrow it, and meets the next street at an angle that makes delivery drivers swear. Local legend, repeated by city guides for a century, says the oldest downtown streets follow the meandering tracks that cattle wore into the ground in the 1600s, ambling around long-vanished boulders and long-drained bogs. When the city grew, nobody asked whether the routes made sense. The routes were simply there, so the city cobbled them, then paved them, then hung traffic lights over them. The paving did not fix the wrong turns. It made the wrong turns permanent, expensive, and fast. Process people have a name for doing this to a workflow, and this level of the program exists because it is the single most common way a well-assessed, fully chartered, competently piloted AI initiative still dies: paving the cow path.

You Have a Charter. Now Comes the Trap.

Level 2 ended with a signature. You assessed a process, built its baseline pack, scored its readiness, presented without making enemies, and converted approval into a Pilot Charter with pre-committed success criteria and an automatic kill condition. In the running example this program has followed, that process is invoice-exception handling: roughly a $428,000 annual run rate, about 1,150 exceptions a month, a median cycle time of 3.2 days with a p90 of 9 days driven mostly by a two-day approval queue, 12 percent rework, and a charter clause that kills the pilot if verified accuracy sits below 85 percent at week six. That is more discipline than most of the world's AI pilots ever see. It is still not enough.

Here is the uncomfortable truth that opens Level 3: everything you built in Levels 1 and 2 tells you whether a process is ready and whether a pilot is honest. None of it yet tells you what the process should become. And in the gap between "we are allowed to deploy AI here" and "we have redesigned the work," a default answer rushes in, because it is the easiest answer in the building: take the process exactly as it is, and bolt the tool onto it. Keep every step, every handoff, every approval, every queue. Just make one of the steps artificial-intelligence-flavored. That is the overlay. That is the paved cow path. The route the work travels was worn in by history, not designed by anyone, and the tool just made the wrong route faster.

Every mature process is full of cattle tracks. The approval that exists because of a fraud incident in 2019, long since fixed by a system control nobody remembers installing. The double data entry born of a 2017 system migration whose "temporary" bridge became load-bearing. The queue that exists because two teams once had a turf war and the truce was a handoff. None of these steps has an owner who can say what it is for. They persist because processes accumulate; they almost never shed. When you overlay AI onto a process like that, you are not transforming anything. You are pouring asphalt over the cow's opinion.

The evidence for how much this matters is the strongest single finding in the entire failure record. McKinsey's State of AI research draws the line plainly: 88 percent of organizations now use AI regularly, yet only about 39 percent can attribute any impact to it in EBIT (earnings before interest and taxes, the profit line a CFO actually watches), and most of those report the impact at under 5 percent. Meanwhile the high performers, roughly 6 percent of organizations, are about three times more likely to have fundamentally redesigned workflows as they deployed AI, and workflow redesign shows up in that research as among the strongest drivers of bottom-line impact. Read those two findings together and the 88-versus-39 gap stops being mysterious. The gap is overlay at scale: tens of thousands of organizations using AI on top of unchanged work, and a thin layer of organizations changing the work itself. MIT's 95 percent, the number that opened this program, is the same finding read from the morgue: the autopsy causes were no workflow integration, no learning loop, adoption without transformation. Not one of those causes is a model property. All three are descriptions of an overlay.

Why Overlay Feels So Right, Especially to Careful People

If overlay were obviously stupid, it would not be the dominant failure mode of a trillion-dollar technology wave. It survives because, at the moment of decision, overlay is the responsible-looking choice. It deserves a slow, honest anatomy, because you will feel its pull on your own pilot within the month.

Overlay minimizes disruption. Nobody's job changes shape. Nobody's approval authority is questioned. The SOP (standard operating procedure) gets one new paragraph instead of a rewrite. The change-management plan fits on a sticky note. For a leader whose last transformation program left scar tissue, "we are not changing the process, just adding a tool" is a sentence that closes objections before they open.

Overlay requires no process argument. The moment you propose deleting a step, you must find the person who believes that step protects them and win that conversation. The 2019-fraud approval has a defender. The duplicate log has a defender. Overlay lets you skip every one of those conversations, which is precisely why it should worry you: a deployment that nobody objects to is usually a deployment that changes nothing.

Overlay ships fast. The vendor can integrate against the existing sequence in weeks, because the existing sequence is documented and stable. A redesign needs mapping, argument, sign-off, retraining. If the steering committee measures progress by go-live date, overlay wins the race every time.

And here is the cruelest part: overlay demos identically to redesign. In week two, both show the same screen: an invoice exception goes in, a neat categorization comes out, the room applauds. You learned in Level 1 to distrust the pilot that demos well, and this is the mechanism behind that distrust operating at process scale. The demo shows the step. It cannot show the system around the step. By month six, only one of the two is alive, because month six is when the business number either moved or did not, and only the redesign was ever connected to the number. Week two cannot tell them apart. Your baseline pack can, which is why the transformer's first act is never opening the tool. It is rereading the baseline.

The value was in the redesign. The AI made the redesign possible. Neither alone moves the number.

The Three Ways Overlay Kills Value

Overlay does not fail vaguely. It fails in three specific, diagnosable ways, and every dead overlay you will ever autopsy contains at least one of them.

Kill one: automating the waste

Every process contains steps that should not exist. Lean practitioners call this muda, the Japanese term for waste: any activity that consumes resources without creating value the customer would pay for. The overlay cannot tell muda from value, because the overlay inherits the step list as given. So the AI now drafts, beautifully and in seconds, the weekly summary report that nobody has opened since the manager who requested it left in 2022. The step that should have been deleted got accelerated instead. Lean spent seventy years teaching that you never automate a wasteful step, you eliminate it; AI is simply the newest and most expensive way to do the wrong thing efficiently. Worse, acceleration entrenches: once the useless report is generated by an AI integration someone paid for, deleting it means admitting the integration was wasted, so the report becomes more immortal than it was when a human resented writing it.

Kill two: preserving the bottleneck

This is the kill your baseline pack exists to predict. Take the invoice-exception process. A vendor's extraction model can read and categorize an exception in about 4 seconds, a step a clerk currently does in about 6 minutes. Impressive ratio. Now place it in the flow: the categorized exception lands in the same approval queue where it waits, on median, two days for a team lead who approves everything and is booked solid. What happens to cycle time? Nothing you could measure. The 3.2-day median and the 9-day p90 were never about reading speed; the baseline showed the queue wait dwarfing every touch time in the process. Speeding a non-bottleneck step is, in constraint-theory terms, a rounding error: the constraint sets the throughput, and the constraint is untouched. Your own baseline pack predicted this outcome before the vendor ever demoed, which is exactly why a process transformer reads baselines before touching tools. The overlay pilot would have run six honest weeks, produced no movement in the chartered metric, and been killed by its own well-written kill condition, a clean death, but a death that redesign would have avoided entirely.

Kill three: stacking verification on top instead of designing it in

An overlay adds AI output to a flow whose roles, time budgets, and SOPs are unchanged. But you learned in Level 2 that every AI-touched figure must be verified, so verification has to happen somewhere. In an overlay, "somewhere" means: on top of everyone's existing job, as an extra task nobody was given time for, with no entry criteria, no sampling plan, and no line in anyone's performance objectives. For two weeks people check diligently. Then the quarter close arrives, checking quietly stops, and the organization is now running unverified AI output through a process that was never designed to catch it. When the first confident wrong output causes damage, the verification log shows a gap exactly where the incident sits. The fix is not "remind people to check." The fix is structural, and it gets a full lesson later in this level: a human gate is a designed step with entry criteria, a review standard, and a time budget, not a hope stapled to an unchanged workload.

The Redesign Method: Questions in Order, Asked on the Map

So what does the redesigner do instead? Not "be more creative." Redesign is a discipline with an order of operations, and the order is what protects you from both failure modes: overlaying out of caution, or blowing up a process out of enthusiasm. You work on the as-is map you built during assessment, and you ask four questions in strict sequence.

Question one: what is this process for? State its output and its customer freshly, in one sentence, as if nobody had ever run it. "Invoice-exception handling exists to get a supplier paid the correct amount, quickly, while catching the small fraction of exceptions that are errors or fraud." That sentence is your plumb line. Every design choice that follows is measured against it, and you will be surprised how many existing steps cannot explain themselves in its terms.

Question two: which steps create that value? Walk the map step by step and mark the ones that move the work measurably toward the output the customer needs. In most mature processes this is a minority of steps. The rest are inspection, transport, waiting, recording, and ritual.

Question three: which steps exist for historical reasons nobody can name? This is the archaeology pass, and it is the emotional center of the method. Every step must re-justify itself to the plumb-line sentence: not "we have always done it," not "compliance probably wants it," but a named owner and a current reason. Each step leaves the pass with one of five verdicts: keep (it creates value or satisfies a control you can cite), delete (it justifies nothing), combine (it duplicates a neighbor), resequence (it is in the wrong place, often a late check that should be an early one), or, the option classic re-engineering never had, reshape around AI capability (the step survives, but its logic changes because a machine can now do something no human could do at this volume). This is classic business process re-engineering discipline, the 1990s question "if we started from scratch, would we build this?", updated with one genuinely new verb.

Question four, and only now: where does AI capability change what is possible? Notice the sequence. You clean the skeleton first, then ask what AI changes, because AI applied to an uncleaned process automates the archaeology instead of removing it. On a cleaned skeleton, AI capability tends to change the process in three recurring shapes: batch becomes flow (work that waited to be processed in daily lumps can move one piece at a time, because the reading step no longer needs a human to sit down with a stack); sequential review becomes exception-only review (instead of a human inspecting 100 percent of items, the machine handles the routine and routes only the unusual, low-confidence, or high-stakes items to human eyes); and the human moves from doing to verifying (the clerk's judgment gets applied where it discriminates, at the gate, not spread thin across a thousand routine keystrokes).

The output of the four questions is the to-be map: the redesigned flow drawn with its gates and handoffs marked. Drawing that map well is the work of this whole chapter. The next three lessons build its components in order: how to triage which steps are AI-ready and which are human-only, how to design the human gates as real controls with time budgets, and how to engineer the handoffs and failure modes so the to-be map survives contact with a bad day. Consider them the drafting tools; today you are learning why the drawing must happen at all.

One Process, Three Futures: The Invoice-Exception Case

Now the centerpiece. Here is the chartered invoice-exception process taken through all three of its possible futures, with numbers. Every figure is illustrative, continuing the storyline, and the projected figures are exactly that: projections, to be tested against the frozen baseline when the pilot runs.

The as-is: eleven steps and a queue

The as-is map shows 11 steps and 3 handoffs. An exception arrives from the matching system; a clerk logs it in the exception tracker; the clerk also logs it in a spreadsheet (a duplicate born of the 2017 migration); the clerk reads the exception and categorizes it, about 6 minutes each, for all 1,150 monthly exceptions, even though 78 percent of them are structured, routine mismatches (quantity, price tolerance, missing goods receipt) that the baseline shows get resolved by rote; the categorized item joins a first-in-first-out queue (FIFO: processed strictly in arrival order, regardless of urgency); it waits a median of two days; the team lead reviews and approves every single item, routine or not; approved items go to resolution; a weekly summary report is compiled that, interviews revealed, nobody reads; and the cycle closes. Median 3.2 days, p90 9 days, 12 percent rework, mostly re-categorization after the team lead disagrees with the clerk's call.

The overlay: what the vendor proposed

The vendor's proposal, reasonably enough from where the vendor sits, was: our model categorizes the exceptions, everything else stays as is. Clerks stop doing the 6-minute read; the model does it in 4 seconds; the queue, the 100 percent approval, the duplicate log, and the unread report all remain. Before anyone signed, the team ran the proposal against the baseline pack as a simulation, the exercise this lesson's Monday section will assign you. The result: categorization touch time is about 115 hours a month, real but small money; the queue wait it feeds is unchanged; projected median improvement lands somewhere around 0.3 days, inside the baseline's own noise, and the p90 tail does not move at all because the tail was made of queue aging, not reading. Against the license and integration cost, the business case fails before week one. Absorb what almost happened: a process with a signed charter, a clean baseline, and honest measurement would still have joined the 95 percent, because the design was an overlay. Assessment discipline gets you a fair trial. Only redesign gets you a verdict worth having.

The redesign: eight steps and a different shape

The redesign starts with the archaeology pass. Two deletions: the duplicate spreadsheet log (no owner could name a current reason) and the unread weekly report (its one former reader confirmed by email that it could die). One combination: logging and categorization collapse into a single automated intake step. Eleven steps become eight, before any AI does anything clever.

Then the reshaping. The routine 78 percent flows through AI categorization straight into a batch-approval lane: the team lead no longer reviews every item, she reviews exceptions to the AI's confidence, the cases the model flags as unusual or scores below threshold, plus a random verification sample. Her judgment, the thing fifteen years of invoices actually taught her, moves up the chain from inspecting everything to ruling on the genuinely hard cases: this is the people story from the Chapter 2.4 scorecard paying off, the craft promoted rather than displaced. The 22 percent free-text exceptions stay fully human, exactly as the charter's OUT list scoped. Verification is a designed gate, not a stacked hope: a written sampling plan, an explicit time budget in the team lead's week, and entries flowing to the verification log from day one. And the queue itself is restructured from FIFO-by-arrival to priority-by-aging, so the oldest exceptions get pulled first and the p90 tail is attacked directly, a change requiring zero AI and worth more than the model.

The projected numbers, stated as projections and stapled to their assumptions: median cycle from 3.2 days to roughly 1.4; p90 from 9 days to roughly 4; rework from 12 percent to roughly 8, on the logic that categorization disagreements, the main rework driver, now occur only on the reviewed minority. Chapter 3.3 will teach you how to test projections like these against the baseline honestly. For now, notice where the value lives. The AI contributed one capability: reading structured exceptions reliably at volume. Everything else that moves the numbers, the deletions, the exception-only review, the priority queue, the designed gate, is process design that the AI made possible but did not perform. That is the law this case exists to teach, and it is the blockquote above: neither the redesign nor the AI moves the number alone.

DimensionAs-isOverlay (vendor proposal)Redesign
Steps / handoffs11 / 311 / 38 / 2
Human reads every exceptionYes (100%)No, but approves 100%Reviews ~22% + flags + sample
Queue disciplineFIFO, 2-day median waitFIFO, unchangedPriority by aging
VerificationImplicit in total reviewStacked on top, unfundedDesigned gate, sampling plan, time budget
Projected median cycle3.2 days (baseline)~2.9 days (within noise)~1.4 days (projection, assumptions listed)
Projected p909 days~9 days~4 days
Projected rework12%~12%~8%

The Artifact: The Redesign Delta Sheet

Level 3's first named artifact is the one-page instrument that makes overlay impossible to disguise: the Redesign Delta Sheet. It is a before/after comparison with four mandatory columns, and its power is entirely in what it refuses to let you leave blank.

  • REMOVED: which steps were deleted outright, with the archaeology finding that condemned each one. For the invoice case: the duplicate spreadsheet log, the unread weekly report.
  • RESEQUENCED: what now happens in a different order or under a different discipline, and why. For the invoice case: the queue moves from FIFO-by-arrival to priority-by-aging, aimed at the p90 tail.
  • RESHAPED: which surviving steps changed their logic around AI capability. For the invoice case: 100 percent sequential review becomes exception-only review with a batch-approval lane; the team lead moves from doing to verifying.
  • WHAT THE AI ACTUALLY DOES: the machine's real, narrow contribution, stated without marketing. For the invoice case: categorizes structured exceptions with a confidence score; nothing else.

Below the four columns, the sheet carries the projection and its assumptions, listed so the pilot can falsify them: "assumes model accuracy at or above the charter's 85 percent verified threshold on the structured 78 percent; assumes team lead gate takes no more than 45 minutes per day; assumes aging-priority queue clears legacy tail within three weeks." The Delta Sheet is a bluff-check that runs in thirty seconds: if REMOVED and RESEQUENCED are empty and RESHAPED says "n/a," then whatever the deck claims, you are looking at an overlay, and "we added AI" can no longer masquerade as "we transformed the process." Attach it to the charter, bring it to every stage gate, and hand it to the CFO before the CFO asks.

A failure story for range: the summary nobody trusted

One brief story from a different world, so you see the pattern travel. A mid-sized legal team bought a contract-summary AI and overlaid it on intake: same intake form, same review sequence, same partners reading everything, plus now a summary on top. Six months later, total reading time was up, not down. Every partner read the summary and then read the full contract anyway, because nothing in the process defined when the summary could be relied on, for what clause types, verified how. Trust was never designed, so it never arrived. The tool was scrapped as "not accurate enough," and here is the point worth underlining twice: nothing about accuracy was ever the problem. The Delta Sheet for that deployment would have read REMOVED: nothing, RESEQUENCED: nothing, RESHAPED: nothing. Overlay failures get misdiagnosed as model failures, the blame lands on the AI, and the 95 percent grows by one more entry that teaches the organization exactly the wrong lesson. This program's founding correction, that readiness is a property of the organization and not the model, is not a slogan; it is a mechanism, and you have now watched the mechanism run.

The wider arithmetic agrees. BCG's 10-20-70 rule prices transformation at 10 percent algorithms, 20 percent technology and data, and 70 percent people and process: redesign is the 70 percent doing its work, and overlay is the attempt to buy the 10 and skip the 70. S&P Global found 42 percent of companies scrapped most of their AI initiatives in 2025, and a large share of those corpses would show the legal team's signature wound: a healthy model inside an unchanged process, condemned at the post-mortem for a crime the process committed.

What to Do Monday Morning

You hold a chartered process and its as-is map. This week you find out whether your pilot design is a redesign or a paving crew.

  1. Reread the baseline pack before you touch anything else, and write down, in one sentence, where the time and the errors actually live. For most processes it is a queue or a handoff, not a touch step. This sentence is the constraint your design must attack.
  2. Run the archaeology pass on the as-is map. Every step re-justifies itself or dies: demand a named owner and a current reason, and assign one of the five verdicts (keep, delete, combine, resequence, reshape around AI). Budget two hours and one slightly awkward meeting.
  3. Write the Redesign Delta Sheet for your current pilot design, the four columns plus the assumptions block. If REMOVED and RESEQUENCED come out empty, you have documented an overlay in your own handwriting, which is precisely the moment to fix it, before the license is signed.
  4. Check the vendor's proposal against the three overlay kills. Does it accelerate any step the archaeology pass condemned? Does it feed the existing bottleneck untouched? Does it add output without a designed, time-budgeted verification gate? One yes is a redesign conversation; three yeses is a different vendor conversation.
  5. Simulate the overlay against your baseline before anyone buys anything. Take the vendor's claimed step improvement, place it in the as-is flow, and compute the effect on the chartered metric. If the projected movement is inside the baseline's noise, you have just saved the pilot's budget and, per the arithmetic of the 95 percent, probably its life.

Key Takeaways

  • Name the failure mode precisely: paving the cow path means bolting AI onto the process as it is, which makes the historical route permanent and fast, exactly like the old streets paved over cattle tracks.
  • Anchor the stakes in the strongest lever in the failure record: McKinsey's high performers (~6 percent) are ~3x more likely to fundamentally redesign workflows when deploying AI, workflow redesign is among the strongest drivers of EBIT impact, and the 88 percent-use versus 39 percent-impact gap is overlay operating at scale.
  • Diagnose overlays by their three kills: automating the waste (muda accelerated instead of deleted), preserving the bottleneck (a 4-second reader feeding a 2-day queue), and stacking verification on top as an unfunded extra job instead of designing it in.
  • Distrust the week-two demo: overlay and redesign look identical in a demo and diverge only when the chartered business number moves or does not, which is why the transformer reads the baseline before touching the tool.
  • Run the redesign method in strict order: state what the process is for, mark the value-creating steps, force every step through the archaeology pass (keep, delete, combine, resequence, reshape around AI), and only then ask where AI changes what is possible.
  • Build the Redesign Delta Sheet for every pilot: one page, four mandatory columns (REMOVED, RESEQUENCED, RESHAPED, WHAT THE AI ACTUALLY DOES) plus listed assumptions the pilot can falsify, so "we added AI" can never masquerade as transformation.
  • Simulate before you sign: place the vendor's claimed improvement into the as-is flow against your baseline; if the chartered metric barely moves, as in the invoice overlay's ~0.3-day median gain, the business case fails before week one.
  • Remember the law the case teaches: the value was in the redesign, the AI made the redesign possible, and neither alone moves the number, which is why overlay failures get misdiagnosed as model failures and why the 95 percent keeps blaming the wrong suspect.