Identifying AI-Ready Steps vs Human-Only Steps
The archaeology is done. The invoice-exception process that entered Chapter 1 as a 14-step fossil bed has come out the other side as an eight-step to-be map, and the map is taped to the wall of a conference room where a redesign team is about to make the decision that will determine whether this initiative joins MIT's 95 percent or escapes it. The facilitator uncaps a marker and asks the sorting question: "Okay, which of these steps does the AI do?" And before anyone can think, two answers arrive at once. The operations lead says, "the boring stuff, obviously." The IT lead says, "well, the vendor demo showed it categorizing emails, so, the categorizing step." Both answers sound reasonable. Both are guesses wearing the costume of decisions. And either one, written onto that map, will quietly place a six-figure bet on vibes. This lesson replaces both answers with a test.
The Sorting Question Nobody Tests
Every workflow redesign reaches this moment. The wasteful steps are gone, the surviving steps are verified, and now each surviving step must be assigned to a performer: the model, the human, or some explicit combination of the two. This is the central sorting act of redesign, and it is astonishing how rarely it is done with anything resembling a method. In most organizations the sort happens in one of two ways, and both fail in documented, expensive, opposite directions.
Sorting by vibes. "AI for the boring stuff" feels like a principle but it is a mood. Boring is a property of the human experience of a step, not of the step's structure, and the two come apart constantly. Approving a payment release is boring; it is also the step where fiduciary accountability lives, and handing it to a model because it is tedious is how you end up explaining to an auditor why nobody human decided to send $180,000 to a vendor. Meanwhile, a genuinely automatable extraction step gets kept human because somebody finds it soothing. Vibes sort by feelings. Processes fail by structure.
Sorting by vendor demo. "AI for whatever the tool does" outsources your process design to a sales engineer who has never seen your process. The demo showed categorization, so categorization gets automated, regardless of whether your categorization step is the routine kind a model handles well or the judgment kind where the category determines whether a vendor gets paid or sued. The tool's capability list becomes the redesign, which is the purest form of paving the cow path you met in the previous lesson: except now the cow path is the vendor's, not even your own.
Both methods produce the same underlying failure, which deserves a name because you will spend your career cleaning up after it: misplaced AI. Misplaced AI is the model doing judgment work it cannot own, or humans grinding through pattern work a model does better. MIT's autopsy of the 95 percent of enterprise generative AI pilots that delivered no measurable profit-and-loss return found the causes in exactly this territory: adoption without transformation, tools bolted onto work nobody re-examined. McKinsey's State of AI research completes the picture from the winners' side: the roughly 6 percent of high performers are about three times more likely to have fundamentally redesigned workflows, and workflow redesign is among the strongest drivers of bottom-line impact. Redesign, concretely, means this sort, done step by step, with a test. And BCG's 10-20-70 rule (10 percent of the effort is algorithms, 20 percent technology and data, 70 percent people and process) tells you where the sorting question lives: squarely in the 70.
The good news is that the sort is teachable. It is a step-level test, run on the map, one step at a time, and marked in three colors. The output is a marked map: the artifact that the next two lessons, on human gates and handoffs, will build directly on top of. Let us build the test.
The Step Triage Card and the Three Colors
The lesson's named artifact is the Step Triage Card: five questions applied to every single step on your to-be map, in order, with the answers written down and the evidence cited. The card ends in exactly one of three verdicts, which you mark on the map in three colors:
- AI-READY (green). The step transforms inputs by learnable regularities, its output can be verified at acceptable cost, no high-stakes consequence hangs on it directly, no relationship depends on it, and its errors are cheap and reversible. The model can perform this step, with verification designed in.
- HUMAN-ONLY (red). At least one of the five questions disqualifies the step: it sets precedent, its errors are invisible or unverifiable, a named human owns its consequence, its value is partly the human contact, or one bad output cascades catastrophically. A human performs this step. AI may assist before it, never replace the human at it.
- HYBRID (amber). The step splits internally: part of it is pattern work the model should do, part of it is judgment work a human must own, and the triage's real output is the step redrawn into its two parts with an explicit handoff between them.
Notice what the card is not. It is not a capability assessment of any particular model, and it is not a return-on-investment calculation. It is a structural test of the step itself: what kind of work the step is, what happens when it goes wrong, and who must answer for it. That is why the card survives model upgrades. A better model changes how well AI performs a green step; it does not change whether a red step should have been green. Now the five questions, slowly, because each one earns its place with a different kind of failure it prevents.
Question 1: Is it pattern or precedent?
Pattern work transforms inputs by learnable regularities: categorizing, extracting, drafting, matching, summarizing, reformatting. The step's logic already exists, written in a standard operating procedure (SOP, the documented instructions for how work is done) or embedded in thousands of past examples, and the performer's job is to apply that existing logic to the next instance. Precedent work is different in kind, not degree: it sets or interprets policy, makes exceptions to rules, or decides what the rules mean. The step that creates the pattern others will follow stays human. The step that follows the pattern is a candidate for AI.
Run the contrast on the invoice-exception process. "Assign this exception to one of five categories defined in section 4.2 of the SOP" is pattern: the categories exist, the logic is written, ten thousand past assignments demonstrate it. "Decide whether this vendor's situation justifies a new category, or a tolerance exception the SOP does not contemplate" is precedent: whatever you decide becomes the rule the next hundred cases follow. A model can imitate the surface of precedent work, and that imitation is exactly the trap: it will produce a confident, fluent ruling with no ability to stand behind it, and the organization will inherit a policy nobody decided. Pattern-following is delegation. Precedent-setting is governance. Only one of those can be handed to a machine.
Question 2: Is the output verifiable at acceptable cost?
This is the verification economics you built in Level 2, now applied at design time instead of after deployment. An AI step is deployable when its output can be checked by sample (audit 5 percent of outputs weekly), by reconciliation (the extracted total must match the source system's total), or by downstream detection (the next step fails loudly if this one erred). An AI step whose errors are invisible until a customer, a vendor, or a regulator finds them is not deployable as designed, no matter how pattern-like the work is.
Contrast: extracting the invoice amount is verifiable almost for free, because it reconciles against the purchase-order system automatically. Summarizing "the tone of the vendor relationship" from an email thread is nearly unverifiable: a subtly wrong summary reads exactly like a right one, and the error surfaces months later as a mishandled account. But here is what makes this question generative rather than merely a filter: verifiability is a design property you can add. A step that fails Question 2 today can often pass it after redesign. Require the model to cite the source line for every extracted field, and spot-checking drops from minutes to seconds. Add a reconciliation total. Route low-confidence outputs to a human lane. When a step fails Question 2, the answer is not automatically "human forever"; it is "redesign the verification path first, then re-triage." Question 2 does not just sort steps. It generates redesign work.
Question 3: Who owns the consequence?
This is the accountability question, and it is the program's fourth non-negotiable made operational: accountability stays human. Ask it brutally: if this step goes wrong, who is named in the incident report? Steps that carry legal signature, fiduciary duty, safety consequence, or contractual commitment keep a human owner at the decision, even when AI assists before it. There is no model that can be deposed, sanctioned, or fired, which means there is no model that can own a consequence; there is only a model that can obscure who does.
The distinction to hold crisply is assistance versus delegation. AI assembling the evidence pack that a human approver reads before releasing a payment is assistance: the human decision is better informed and still a human decision. AI releasing the payment when its confidence exceeds a threshold is delegation: the decision has been transferred to something that cannot answer for it, and the human "oversight" of a log nobody reads is a gesture, not a control. Contrast pair: drafting the settlement letter is assistable; signing it is not delegable. Compiling the safety-check evidence is assistable; certifying the check is not delegable. When Question 3 fires, the step is HUMAN-ONLY at the decision point, whatever happens upstream of it.
Question 4: Does it need the relationship?
Some steps produce two outputs: the visible deliverable and an invisible one, the relationship itself. The difficult vendor call, the conversation with a grieving customer, the negotiation where reading the room is the actual work: in these steps, the message is the lesser product. The contact is the asset. Automating the message can be perfectly efficient and still destroy the asset, and the destruction will never appear on a process dashboard, because relationship value is invisible to process metrics until the moment it is gone.
One vivid illustration, hypothetical but drawn from a pattern every procurement veteran recognizes. A distributor's accounts-payable team had, for years, handled disputed invoices from its largest freight vendor with a monthly call between two people who had known each other for a decade. The call resolved disputes, yes; it also surfaced early warnings ("we're changing our fuel surcharge structure next quarter, heads up") that never appeared in any written channel. A redesign automated the dispute correspondence: polite, accurate, instant emails. Dispute cycle time improved 40 percent, and the KPI dashboard (key performance indicators, the metrics leadership steers by) celebrated. Eleven months later, the vendor's surcharge restructure landed as a complete surprise, cost roughly $220,000 in unbudgeted freight before contracts could be reopened, and the post-mortem found the warning had simply had no channel to travel through anymore. Every number in that story is illustrative. The mechanism is not. Contrast pair: the templated "please resubmit with a purchase-order number" note has no relationship in it, automate it freely; the quarterly call with the strategic vendor is the relationship, and the step's true output never appears in the ticket system.
Question 5: How expensive is the exception?
The final question prices the step's worst day. Steps with cheap, reversible errors tolerate AI early: a mis-filed document gets re-filed, a wrong category gets corrected downstream, and the cost of the error is minutes. Steps where one bad output cascades demand gates and staging regardless of how pattern-like they look: the wrong payment released, the wrong dosage transcribed, the wrong legal deadline calendared. The work may be 99 percent mechanical; the 1 percent exception is the whole risk profile.
Contrast: auto-tagging incoming exception emails costs almost nothing when wrong, because the next human touch catches it. Auto-calendaring a statutory response deadline looks like the same extraction work, but a single transposed date can forfeit a legal position worth more than the entire automation program. Question 5 is where you will connect, at the end of this chapter, to FMEA (failure mode and effects analysis, the discipline of cataloguing how each step can fail, how badly, and how detectably, before it does). For now the triage-level version is enough: for each step, write one sentence describing the most expensive single error it could emit, and let that sentence argue with the color you were about to assign.
Sort steps by test, not by vibes: pattern work goes to the machine, and precedent, consequence, relationship, and catastrophic exceptions stay with a named human.
HYBRID: The Category That Does the Real Work
Here is what a first pass with the card reveals on almost every real map: the most valuable steps refuse to be one color. They split internally. The drafting inside them is AI-ready; the deciding inside them is human-only. If you force such a step whole into green, you have delegated judgment; force it whole into red, and you have wasted the model on work it demonstrably does better and faster. The honest verdict is amber, and amber is not a compromise. It is an instruction: redraw this step as two steps, an AI part and a human part, with an explicit handoff between them. "AI drafts, human owns" is the default hybrid pattern, and the handoff it creates is a designed object in its own right, which is why this chapter gives handoffs their own lesson two stops from now.
Watch one full split on the page, because this move is the level's signature skill. Take the exception-categorization step from the invoice process: "categorize the exception into one of five SOP-defined types." Run the card: pattern work (Question 1: green), verifiable by sampling and by downstream correction (Question 2: green), but the category drives routing, and one of the five types triggers a contractual dispute clause (Question 3: flickers red), and roughly a fifth of exceptions arrive as free-text vendor complaints that fit no clean type (Question 1 again: those instances are closer to interpretation than pattern). The step is amber. So redraw it:
- Step 3a (AI): propose. The model reads the assembled case file and outputs three things, never fewer: a proposed category, a confidence score, and cited evidence (the specific invoice lines and SOP passage that justify the proposal). No naked answers; the citation requirement is Question 2's verifiability, designed in.
- Step 3b (human): decide. For the routine lane (high confidence, standard types), the analyst confirms or overrides the proposal, a seconds-long review with full authority to reject. For the free-text lane (low confidence, or the model flags no clean fit), the analyst categorizes from scratch: this lane is human-performed, not human-reviewed, because interpreting an angry three-paragraph vendor email is precedent-adjacent work.
- The handoff between them specifies what 3a must deliver for 3b to start (proposal, confidence, evidence, or an explicit "no fit" flag) and what 3b's override feeds back into the record. That specification is a placeholder for now; the handoff lesson will give it a full grammar.
One step became two steps and a seam. Multiply that by every amber verdict on your map and you can see what the triage actually produces: not a labeling exercise, but the next round of redesign work, generated step by step. Expect HYBRID to be your most common verdict on any process worth transforming; a map with no amber on it usually means the triage was run too coarsely.
The Marked Map: Invoice Exceptions, Triaged
Now the whole eight-step to-be map, triaged in one table. The format below is the deliverable: one row per step, the five answers in shorthand, the color, one line of reasoning. Shorthand key: Q1 P = pattern, PR = precedent; Q2 V = verifiable at acceptable cost, X = not as designed; Q3 LO = low-stakes consequence, HI = named-owner consequence; Q4 N = no relationship, Y = relationship-bearing; Q5 CHP = cheap reversible errors, EXP = expensive cascading exception. All volumes and figures are illustrative: assume roughly 1,100 exceptions a month and a team of four analysts.
| # | Step | Q1 | Q2 | Q3 | Q4 | Q5 | Color | Reasoning |
|---|---|---|---|---|---|---|---|---|
| 1 | Intake and case assembly (log exception, pull invoice, PO, receipt into one case file) | P | V | LO | N | CHP | AI-READY | Pure matching and collation; reconciles against source systems; a bad assembly fails loudly at the next step. |
| 2 | Variance extraction and computation (line-level fields, variance amount and percent) | P | V | LO | N | CHP | AI-READY | Learnable extraction with arithmetic reconciliation built in; cited source lines make sampling cheap. |
| 3 | Exception categorization (five SOP-defined types) | P/PR | V | LO/HI | N | CHP | HYBRID | Routine lane is pattern; free-text lane is interpretation; one category triggers a dispute clause. Split as 3a/3b above. |
| 4 | Resolution recommendation against tolerance policy | P/PR | V | HI | N | EXP | HYBRID | Applying written tolerances is pattern; granting exceptions to them is precedent. AI computes and proposes; human decides anything outside tolerance. |
| 5 | Routine vendor correspondence (templated resubmission and information requests) | P | V | LO | N | CHP | AI-READY | Templated, sampled weekly, no relationship content, worst error is a confusing email that triggers a reply. |
| 6 | Vendor dispute call for strategic accounts | PR | X | HI | Y | EXP | HUMAN-ONLY | Question 4 territory: the call is the relationship; its second output (early warnings, goodwill) is invisible to metrics and unautomatable. |
| 7 | Final payment release or write-off approval | PR | V | HI | N | EXP | HUMAN-ONLY | Question 3 territory: fiduciary consequence with a named owner in every incident report. AI assists before the decision; never makes it. |
| 8 | Monthly pattern review and tolerance-policy update | PR | V | HI | N | EXP | HYBRID | AI summarizes exception patterns and drafts the analysis; setting the new tolerance is precedent and stays with the process owner. |
Read the shape of the result: three green, two red, three amber. The greens cluster where work is extraction, matching, and templated drafting with built-in reconciliation. The reds are exactly where Questions 3 and 4 fire: the relationship-bearing call and the fiduciary release. The ambers are the valuable middle, each one now owing you a split and a handoff. Priced illustratively: steps 1, 2, and 5 currently consume about 60 percent of the four analysts' time, roughly 2,700 hours a year; if AI performs them with a 5 percent human sampling overhead, on the order of 2,300 hours move up the judgment chain into the amber and red steps, which is where the backlog, the vendor escalations, and the write-off leakage actually live. That reallocation, not headcount theater, is what the marked map is for. And the map itself, colors and all, is the input to the next lesson, where every red and amber step gets its gate designed as a real control with entry criteria and review standards, not a checkbox.
The Two Classic Mis-Triages
Two failure patterns account for most triage wreckage, and they fail in opposite directions. Learn them as named patterns and you will spot them in the first ten minutes of any redesign review.
The Competence Trap: automating judgment because the model sounds confident
A mid-market manufacturer, in an illustrative story assembled from real failure patterns, triaged its own invoice-exception process and marked the rejection decision green: the model was so consistently, fluently right about "obviously invalid" exceptions in testing that the team let it auto-reject them, no human in the lane. For three months the dashboard glowed. Then a controller noticed a cluster: exceptions from one recently acquired subsidiary were being rejected at four times the base rate. The model had learned the parent company's invoice formats and vocabulary; the subsidiary's legacy documents, legitimately structured but differently shaped, pattern-matched to "invalid." Roughly $48,000 in valid vendor credits had been wrongly refused, two suppliers had put shipments on hold, and unwinding the damage took a quarter. The autopsy line belongs on your wall: sounding right is not owning the consequence. The model's confidence was a property of its fluency, not of its correctness, and the rejection decision was always a Question 3 and Question 5 step wearing a Question 1 costume. The fix was never a better model. It was the triage the team skipped: auto-rejection is delegation of a consequence-bearing judgment, and the card would have marked it amber at best, with the reject decision human-owned.
The Martyr Trap: keeping pattern work human out of misplaced respect
The opposite error is quieter and costs more, because it produces no incident to autopsy. A regional insurer's operations director, in a second illustrative story, blocked automation of claims-data re-entry with a sentence that sounded like loyalty: "our people are too skilled for this to be automated." It felt protective. It functioned as a sentence: eleven experienced clerks spent two more years retyping fields between systems, work a model performs with a lower error rate, while a competitor triaged honestly, automated the retyping, and redeployed the same skilled people up the judgment chain into complex-claims review, the work the clerks had actually been hired for their judgment to do. When attrition finally forced the insurer's hand, the two lost years had cost, on illustrative arithmetic, over 20,000 hours of skilled labor spent on pattern work, and the best two clerks had left for the competitor, who let them do judgment work. Sentiment is not a triage criterion, and it serves nobody, least of all the people it claims to protect. Respect for skilled workers is redeploying their skill to where only humans can spend it, not preserving their keystrokes.
Notice the symmetry: the Competence Trap fails Question 3 by pretending fluency is ownership; the Martyr Trap fails Question 1 by pretending sentiment is structure. The card catches both, because the card asks about the step, not about the feelings anyone has about the step.
AI's Role in the Triage Itself (and Its Hard Limit)
Should AI help run the triage? Yes, in two specific seats, and never in the third. First seat: first-pass drafter. Feed the model your to-be map and the underlying SOP and prompt it in the pattern you learned in Level 2: "For each of the eight steps, answer the five triage questions. Cite the SOP passage or map element that justifies each answer. Where the evidence is insufficient to answer, say so explicitly rather than guessing." You will get a draft card for every step in minutes, with the citation requirement making each claim checkable, and the drafts will be usefully wrong in instructive places: the model routinely misses relationship value (Question 4 lives in territory no SOP documents) and routinely underprices exceptions (Question 5 requires knowing what a bad day costs, which is institutional memory, not text).
Second seat: adversarial reviewer. For every verdict you feel most certain about, prompt: "Argue that step 4 belongs in the opposite category. Steelman the case." This is cheap red-teaming for your own confirmation bias, and once or twice per map it will find something real: a verification path you had not considered that turns a red step amber, or a cascade you had not priced that turns a green step amber.
The third seat, the deciding seat, is structurally closed to the model, and the reason is not caution but logic: Question 3 is literally the question of who owns consequences, and it cannot be answered by the thing being triaged. A model asserting "this step's consequence can be safely delegated to a model" is a conflict of interest wearing a confidence score. The human runs the card, signs the colors, and answers for the map. AI drafts the triage; it does not get a vote in it.
What to Do Monday Morning
The marked map is buildable in one focused week on the process you already carry from the previous lesson. Here is the sequence.
- Run the Step Triage Card on every step of your to-be map. Five questions per step, answers in shorthand, one line of reasoning, evidence cited. Do it in a table like the worked example; the table is the deliverable, not notes toward one.
- Mark the three colors on the map itself. Green, red, amber, physically on the wall or in the diagram file. The colored map is what the human-gate and handoff lessons consume next.
- Split one hybrid step explicitly. Pick your most valuable amber step and redraw it as its AI part and its human part: what the model produces (proposal, confidence, cited evidence), what the human decides, and one sentence describing the handoff between them.
- Check the map against both traps. For every green step, ask the Competence Trap question: if this step's output were confidently wrong for a subpopulation, when would we find out, and who pays? For every red step, ask the Martyr Trap question: is this human because the structure demands it, or because automating it feels disrespectful?
- Use AI as drafter and adversary, not judge. Generate first-pass cards from the SOP with citations required, then have the model argue the opposite of your two most confident verdicts. Keep the deciding seat human, in writing.
- Have the process owner challenge two of your markings with evidence. Not opinions: evidence. If a color survives a challenge from the person who owns the process, it is probably right; if it flips, the card just did its job before deployment instead of after.
Key Takeaways
- Reject the two default sorting methods, vibes ("AI for the boring stuff") and vendor demo ("AI for whatever the tool does"); both produce misplaced AI, the failure of models owning judgment they cannot answer for and humans grinding through pattern work models do better.
- Run the Step Triage Card on every surviving step of the to-be map: pattern or precedent, verifiable at acceptable cost, who owns the consequence, does it need the relationship, how expensive is the exception; verdicts are AI-READY, HUMAN-ONLY, or HYBRID.
- Treat verifiability as a design property you can add, not just a filter: citations, reconciliation totals, and confidence-routed lanes can move a step from unverifiable to deployable, which makes Question 2 a generator of redesign work.
- Keep a named human at every decision carrying legal signature, fiduciary duty, safety consequence, or contractual commitment; distinguish assistance (AI informs the decision) from delegation (AI makes it) crisply, because only humans can own consequences.
- Expect HYBRID to be the workhorse verdict: valuable steps split internally into "AI drafts, human owns," and each amber step must be redrawn as an AI part, a human part, and an explicit handoff, like the categorization step's propose-and-decide split.
- Name the two classic mis-triages and check every map against them: the Competence Trap (auto-rejecting exceptions because the model sounds confident, until the acquired subsidiary's formats surface as a $48,000 illustrative cleanup) and the Martyr Trap (two years of skilled clerks retyping fields in the name of respect).
- Use AI to draft first-pass triage answers with SOP citations and to argue the opposite of your confident verdicts, but keep the deciding seat human, because Question 3 cannot be delegated to the thing being triaged.
- Deliver the marked map as this lesson's output: three colors on eight steps in the invoice example, with the red and amber steps now queued for the next lesson, where their human gates get designed as real controls.
Skill.re