Process Readiness: Stable, Documented, Measurable
The whiteboard session is forty minutes old when the COO uncaps the marker and draws a circle around the worst process in the building: client intake, the one with the escalation emails, the missed handoffs, the three-week backlog, and the standing agenda slot at every leadership meeting. "This," she says, "is where the AI goes first. It is our biggest pain point." Around the table, heads nod, because the logic feels airtight: aim the expensive new capability at the biggest fire. Six lessons into this program, you already know how this movie ends. Nine months later there will be a stalled pilot, a vendor invoice, and a quiet agreement to stop bringing it up. Not because the AI was bad, but because intake was never one process. It was eleven improvisations wearing a shared inbox, and the AI was asked to automate a thing that did not exist yet. This lesson is about the maturity bar a process must clear before AI lands on it, the three properties that define that bar, and a triage instrument that can sort your entire process inventory in a single afternoon.
The Instinct That Aims at the Fire
Chapter 1 gave you the failure record: 95 percent of enterprise GenAI pilots with no measurable return, and MIT's autopsy showing the failure was organizational, not technical. This chapter has been walking through where, exactly, the organizational failure lives. The 10-20-70 lesson gave you BCG's arithmetic: 10 percent of the effort is algorithms, 20 percent is technology and data, 70 percent is people and process. The previous lesson covered the data layer. Now we arrive at the largest single slice of that 70 percent, and at the most reliable way organizations get it wrong.
The wrong move has a name in this program: aiming at the fire. When leadership teams pick their first AI target, they overwhelmingly point at the most chaotic, most complained-about process, because that is where the pain is, and pain feels like a business case. It is exactly backwards. Chaos is a redesign problem, not an AI problem. A process that runs differently every time, that lives in nobody's documentation, that has no measured baseline, is not a candidate for automation of any kind, AI or otherwise. It is a candidate for the older, less glamorous disciplines: process mapping, standardization, exception analysis. AI does not fix a broken process. AI is an amplifier, and an amplifier is loyal to whatever signal it receives.
Land AI on a stable process and it scales the standard: faster cycles, fewer keystrokes, the same output shape every time. Land it on chaos and it scales the chaos: it automates the inconsistency, executes the improvisations at machine speed, and, this is the specifically AI-shaped danger, it hallucinates the missing rules. Remember the hallucination lesson from Chapter 2: a generative model fills gaps confidently. Ask it to run a process whose rules were never written down, and it will not stop and ask for the rules. It will invent plausible ones, apply them fluently, and produce output that looks like the work of a well-run operation right up until an auditor, a customer, or a regulator discovers otherwise.
There is a repair sequence, and its order is not negotiable: stabilize, then document, then measure, then automate. Every step earns the next. Stabilization makes the process describable; documentation makes it checkable; measurement makes the AI's impact provable; and only then does automation compound value instead of compounding noise. Notice what that sequence is made of: it is Lean, it is BPM (business process management), it is the standard-work discipline process professionals have carried for decades. This is the structural reason, previewed here and argued in full in this level's career chapter, why process people win the AI decade: the scarce skill is not prompting the model, it is getting a process to the state where a model can land on it.
Stabilize, document, measure, then automate. Run those steps in any other order and you are paying to make chaos faster.
One more piece of the instinct deserves a name: the pain paradox. The processes that hurt the most are very often high-pain precisely because they are unstable; the chaos is the cause of the pain, not a separate feature of it. Which means the loudest process in the building is usually the least ready for AI, and the best first target is usually sitting one desk over: the boring, stable, high-volume process nobody complains about because it merely consumes 4,000 quiet hours a year. This is the same shape as MIT's budget finding from Chapter 1: money chased the visible and glamorous while the measurable ROI hid in the unglamorous back office. Pain gets the pilot; boredom delivers the return.
Stable: Would Three People Describe the Same Steps?
The first property is stability. A process is stable when it runs essentially the same way most of the time: the same steps, in the same order, triggered by the same events, with a bounded set of known exceptions. Stability does not mean zero variation; it means the variation is the exception and the standard path is the rule.
Why AI needs it: every AI deployment, whether it drafts, extracts, predicts, or acts, is trained, configured, or prompted against a pattern. If the process has no pattern, there is nothing to configure against. The model will be shown ten examples that are secretly ten different processes, and it will average them into an eleventh process that no one has ever actually run. Instability also destroys verification, and verification, as Chapter 2 taught you, is the job. A reviewer can check output against a standard in seconds; checking output against "whatever Priya would have done in this situation" takes as long as doing the work.
The three-person test. Separately, without letting them compare notes, ask three people who perform the process to describe its steps from trigger to done. Then put the three descriptions side by side. If they match in sequence and substance, with differences only in phrasing, you have a stable process. If you get three genuinely different workflows, you do not have an unstable process so much as three processes sharing a name, and no AI tool can serve three masters that each believe they are the only one. This test costs ninety minutes and predicts pilot outcomes better than most paid assessments.
The exception-rate test. Pull the last 50 or 100 executions and sort them: standard path or exception. As a working threshold, if more than roughly 20 percent of cases leave the standard path, treat the process as unstable for AI purposes. Below that line, exceptions are a routing problem (send them to a human, as you will see in the partial-readiness section). Above it, "exception" has stopped meaning anything; the process is improvisation with a standard-path costume, and the fix is redesign, not routing.
Documented: The SOP That Matches Reality
The second property is documentation, and the bar is specific: a standard operating procedure (SOP, the written step-by-step of how the work is actually performed) that matches current reality. Not a process narrative written for the 2019 quality audit, exported to PDF, and untouched since. An SOP that has drifted from practice is worse than no SOP at all, because it will be handed to the AI team as ground truth, and they will faithfully configure the tool to perform a process the organization quietly abandoned years ago.
Why AI needs it: documentation is where the rules live, and the rules are what separate a model that follows your process from a model that invents one. Prompts, retrieval sources, agent instructions, validation checks, reviewer checklists: every one of these artifacts is written from the SOP. If the SOP does not exist, the pilot team will reconstruct the rules from interviews under deadline pressure, which produces a document that is part memory, part politics, and part guess. And when the tool's output is wrong, documentation is what makes the wrongness detectable; without a written standard, every review collapses into opinion.
The walk-through test. Take the SOP, pull three recent real executions of the process from your systems (tickets, files, email threads, whatever the work leaves behind), and walk the document against each one, step by step, counting divergences: steps performed but not documented, steps documented but skipped, sequences that differ, decision rules applied that appear nowhere on the page. Zero to two minor divergences across three executions: the SOP is live, and the process passes. A material divergence in every execution: you have a fiction with a revision history. The test takes half a day and is, incidentally, the fastest possible education in how the process actually runs, which is why the walk-through notes are never wasted even when the SOP fails.
A warning for the honest scorer: the presence of a document is not the property. In triage sessions, teams routinely award full documentation credit because a file exists. The property is correspondence between the page and the work. Score the correspondence.
Measurable: Four Cells You Can Fill Today
The third property is measurability: the process's basic operating numbers are known, or knowable from existing systems within about a week. Four numbers matter, and together they form the baseline row this program will keep returning to: volume (how many times per month or year the process runs), cycle time (trigger to done, on average and at the slow tail), error or rework rate (what fraction of executions come back, get corrected, or fail downstream), and cost per unit (loaded labor and system cost divided by volume).
Why AI needs it: this is the measurement discipline lesson from Chapter 1 wearing process clothes. A pilot on an unmeasured process cannot succeed, not because it cannot create value but because the value can never be demonstrated, and in a budget review, undemonstrable and nonexistent are the same word. The baseline is also your targeting system: volume tells you whether the prize is large enough to bother, cycle time tells you where the hours are, error rate tells you what quality bar the AI must beat, and cost per unit turns the whole conversation into money, which is the only language a steering committee reliably speaks.
The four-cell test. Draw a one-row table with the four cells and try to fill it today, from data you already have: time-tracking, ticket systems, finance extracts, document logs. Four cells filled with defensible numbers: fully measurable. Two or three cells, with the rest fillable in a week of pulls: measurable enough to proceed while you close the gaps. Zero or one: the process is running on anecdote, and any AI business case built on it will be a work of creative writing. Note the asymmetry with the other two properties: measurability is usually the cheapest gap to close. Stabilizing a process takes weeks of redesign; documenting it takes days of shadowing; a baseline can often be assembled in an afternoon of exports. When the triage grid shows measurability as a process's only gap, that is the best news on the page.
Partial Readiness: The Standard-Path Play
Before the instrument, one refinement that separates a practitioner from a checklist-reader: readiness is not always all-or-nothing across a process. A process can be ready in one segment and hopeless in another, and the most common split is the standard path versus the exception tail. Picture an accounts-payable process where 80 percent of invoices match a purchase order and sail through identical steps, while 20 percent are a swamp of missing POs, disputed amounts, and one-off vendor arrangements. Scored as a whole, the process might land in the middle of every scale. Scored as two segments, the 80 percent standard path is stable, documentable, and measurable, and the 20 percent tail is a redesign project.
The play: scope the pilot to the standard path, and route the tail to humans by explicit rule. The AI handles the matched invoices; anything that fails the match criteria goes to the exception desk exactly as it does today. This does three valuable things at once. It gives the pilot a stable pattern to land on, it converts the exception tail from a pilot-killer into a defined backlog for later process work, and it builds the routing discipline (clear entry criteria for what the AI touches) that Level 4's human-gate design will formalize. When you triage your inventory, therefore, ask of every middling scorer: is this one mediocre process, or a ready segment stapled to an unready one? The second answer is an opportunity the whole-process score hides.
The Artifact: The Afternoon Triage
Here is the lesson's deliverable, and it is built for speed on purpose. Deep process assessment has its place (Level 2 devotes a full chapter to process discovery and selection), but most organizations do not need a six-week study to find their first targets. They need to sort twelve or twenty candidate processes into three piles in one afternoon, with the leadership team in the room, using scores they assigned themselves. That last clause is the political engine of the instrument: numbers the room produced are numbers the room cannot dismiss.
For each process in the inventory, assign four scores of 0 to 2 using the anchors below, and sum to a total out of 8. Score in pencil, score honestly, and when the room argues about a score, write down the argument: the disagreement is itself diagnostic data (a stability argument between two people who both run the process is the three-person test failing live).
| Dimension | 0 points | 1 point | 2 points |
|---|---|---|---|
| Stability | Performers describe different workflows; ad hoc handling is the norm, not the exception | A recognizable standard path exists, but exception rate exceeds roughly 20 percent or steps vary by performer | Runs the same way most of the time; three people describe the same steps; exception rate under roughly 20 percent |
| Documentation | Nothing written, or a document nobody has opened in years | An SOP exists but diverges materially from real executions in the walk-through test | A current SOP survives a walk-through against three real executions with only minor divergences |
| Measurability | Zero or one of the four baseline cells (volume, cycle time, error rate, cost per unit) can be filled today | Two or three cells can be filled today; the rest within a week from existing systems | All four cells can be filled today with defensible numbers |
| Volume and repetition | Rare or bespoke: a handful of executions per year, each substantially unique | Regular but modest volume, or noticeable variation between executions | High volume of near-identical executions: hundreds or more per year of the same shape |
Then read the totals against three bands:
- 7 to 8: AI-ready now. This is your pilot shortlist. Rank within the band by volume times cycle time (the size of the prize) and start where the prize is largest and the risk of a wrong output is most survivable.
- 5 to 6: fixable this quarter. Not ready, but one identifiable gap away. The obligation of this band is to name the gap in writing: "documentation, one week of SOP shadowing" or "measurability, one afternoon of exports." A 5 without a named gap is a wish; a 5 with a named gap and an owner is a pipeline.
- 0 to 4: redesign first. Not an AI candidate this cycle, and saying so is the instrument's most valuable output. These processes go to the process-improvement backlog, where Lean and BPM methods do their work, and they re-enter triage when they can score their way in. Putting AI here is not ambition; it is the 95 percent, scheduled in advance.
Two rules of use. First, score the process as it runs today, not as the improvement project promises it will run by Q3; the grid measures reality, not roadmaps. Second, re-run the triage quarterly. Scores move, and a process that climbs from 4 to 7 because someone stabilized and documented it is the readiness flywheel working exactly as designed.
Corr & Hale: Twelve Processes, One Afternoon
Here is the instrument at full length, in a hypothetical composite built from the patterns this triage reliably surfaces. Corr & Hale is a fictional 300-person professional-services firm: audit-adjacent advisory work, eleven partners, a back office of about 60. The managing partner has read the headlines and wants "AI in the firm by year end," and his personal candidate is new-client conflict checking, the process he describes, accurately, as the one that "takes forever and scares me the most." The COO, who has taken this program, books a three-hour session with the operations leads, lists twelve back-office processes on the wall, and runs the Afternoon Triage. Here is the wall at the end of the session.
| Process | Stab | Doc | Meas | Vol | Total | Call |
|---|---|---|---|---|---|---|
| Engagement-letter drafting | 2 | 2 | 2 | 2 | 8 | Ready now: first pilot |
| Accounts-payable invoice coding | 2 | 2 | 2 | 1 | 7 | Ready now |
| Timesheet chase and correction | 2 | 1 | 2 | 2 | 7 | Ready now |
| IT access provisioning and offboarding | 2 | 2 | 1 | 2 | 7 | Ready now |
| Client-onboarding file assembly | 1 | 1 | 2 | 2 | 6 | Fixable: stabilize intake variants |
| Expense-exception handling | 2 | 0 | 1 | 2 | 5 | Fixable: write the SOP |
| Proposal assembly | 1 | 1 | 1 | 2 | 5 | Fixable: standardize the template path |
| Vendor-contract renewals | 2 | 1 | 1 | 1 | 5 | Fixable: build the baseline |
| Meeting minutes and action tracking | 1 | 0 | 1 | 2 | 4 | Redesign first |
| Recruitment CV screening | 1 | 1 | 1 | 1 | 4 | Redesign first |
| New-client conflict checking | 0 | 1 | 1 | 1 | 3 | Redesign first |
| Client-satisfaction follow-up | 1 | 0 | 0 | 1 | 2 | Redesign first |
Three hours of scoring, and the inventory has sorted itself into four ready, four fixable, four redesign. Now zoom into three rows, because the zoom is where the money is.
The 8: engagement-letter drafting
The firm issues about 900 engagement letters a year, 92 percent of them built from one of four standard templates. The time-tracking system (this is a professional-services firm; hours are the house currency) shows an average of 3 hours 20 minutes per letter across drafting, partner mark-up, and correction cycles, at a loaded cost of roughly 88 dollars per hour: call it 260,000 dollars a year of effort. The document-management system shows a 7 percent rework rate, letters bounced back by partners for scope or fee errors. Stable, documented (the SOP survived a walk-through in the session's coffee break, two minor divergences), fully measurable, high volume: 8 out of 8. This is the first pilot: AI drafts from the template against the engagement record, a senior administrator verifies against a checklist derived from the SOP, and the baseline row already exists, which means in six months the firm will be able to prove whatever happened, in either direction. Notice how boring this process is. Nobody at Corr & Hale has ever complained about engagement letters in a leadership meeting. That is the pain paradox paying out.
The 5: expense-exception handling
About 2,600 expense exceptions a year flow to one coordinator, Dana, who resolves them with rules that are consistent, sensible, and written down nowhere. The three-person test could not even be run: there is no second person. Stability scored 2 (Dana is a very stable system), documentation 0, and the room heard the risk out loud: the process is one resignation away from scoring 0 on stability too. The named gap: one week. An analyst shadows Dana for four days, drafts a nine-page SOP with a fourteen-rule decision table, and validates it against 40 historical cases; total cost about 3,500 dollars in analyst time. Documentation moves 0 to 2, the score moves 5 to 7, and the firm has, incidentally, insured itself against Dana's next job offer. This is what "fixable this quarter" means: not a vague aspiration, but a one-line work order with a price.
The 3: new-client conflict checking
The managing partner's candidate. In the session, three partners describe how a conflict check works, and the room gets to watch the three-person test fail in real time: one starts from the CRM, one emails all eleven partners and waits for silence, one checks a spreadsheet whose last edit was in 2023. There is no agreed definition of what counts as a conflict "hit," no log of checks performed, and one near-miss last year that everyone lowers their voice to mention. Score: 3. And here the grid earns its keep politically, because the COO does not have to tell the managing partner "no." She points at the wall the room scored together and says: "Not yet, and here is the path. Six weeks to agree one workflow and one source of truth, two weeks to document and log it, then it re-enters triage next quarter, and given the volume it will probably pilot in Q2." An AI pointed at today's conflict checking would automate three contradictory processes at once and hallucinate the definition of a conflict, which is not a pilot, it is a malpractice generator with a dashboard. The managing partner, shown the path instead of the refusal, funds the stabilization work on the spot. That is the instrument's political value in one scene: the grid turns "no" into "not yet, and here is the path," and it makes the "not yet" the room's own conclusion rather than your opinion.
What to Do Monday Morning
The triage is designed to run inside one week from a standing start. Here is the sequence.
- List 10 to 15 candidate processes from your own area or the back office you know best: anything repetitive, document-heavy, or queue-shaped. Do not curate for readiness yet; the grid does the sorting.
- Run the three-person test on the two processes you think you know best. Three separate descriptions, side by side. Budget ninety minutes. Expect at least one unpleasant surprise; that surprise is the lesson installing itself.
- Build the four-cell baseline for one process (volume, cycle time, error rate, cost per unit) from systems you already have. Time how long it takes; that duration is your measurability calibration for every other score you assign.
- Book the afternoon and score the grid with the people who run the processes, not just their managers. Use the anchors verbatim, score today's reality, and write down every argument the scoring provokes.
- Publish the three piles with a named gap for every 5 and 6. One page: ready now (ranked by prize size), fixable this quarter (gap, owner, cost), redesign first (routed to the improvement backlog with a re-entry date).
- When someone senior aims at the fire, answer with the wall. Not yet, here is the score, here is the path, here is the re-triage date. File the scored grid in your readiness portfolio next to Chapter 1's failure-mode checklist.
Key Takeaways
- Aim AI only at processes that clear the maturity bar: stable (runs the same way most of the time), documented (an SOP that matches reality), and measurable (volume, cycle time, error rate, cost per unit known or knowable), because AI amplifies whatever it lands on.
- Resist the chaos instinct: the most painful process is usually high-pain because it is unstable, and chaos is a redesign problem for Lean and BPM methods, not an AI problem; automated chaos is just inconsistency at machine speed with hallucinated rules.
- Run the three fast tests before any pilot conversation: three people describing the same steps and an exception rate under roughly 20 percent (stable), an SOP that survives a walk-through against three real executions (documented), four baseline cells fillable today (measurable).
- Follow the repair sequence in its only working order, stabilize then document then measure then automate, and treat measurability as the cheapest gap and stability as the most expensive.
- Exploit partial readiness: when a process splits into an 80 percent standard path and a 20 percent exception tail, scope the pilot to the standard path and route the tail to humans by explicit rule.
- Sort your whole inventory with the Afternoon Triage: score stability, documentation, measurability, and volume 0 to 2 each; 7 to 8 is ready now, 5 to 6 is fixable this quarter with a named gap and owner, 0 to 4 is redesign first.
- Use the grid as a political instrument: scores the room assigned itself turn "no" into "not yet, and here is the path," which redirects executive pet projects without spending your own credibility.
- Re-run the triage quarterly and treat climbing scores as the point: unready processes are where pilots go to die, and moving a process from 4 to 7 before piloting is the discipline that keeps you out of the 95 percent.
Skill.re