Catching the Invented Step: Verifying AI Process Maps
The review meeting is going beautifully, which should have been the warning. Twelve people around a table, the AI-drafted map of the accounts payable exception process on the screen, and heads nodding at every box. There it is, the three-way match between purchase order, receipt, and invoice. There is the duplicate-vendor check. There is the manager sign-off on exceptions over $5,000. The controller says "yes, that looks right," and it does look right. It looks exactly like a well-run accounts payable process. The only problem, discovered four months and one mis-aimed automation pilot later, is that two of those boxes describe a process this company has never run. Nobody lied. Nobody was careless. The AI drew the textbook, the room recognized the textbook, and the textbook was approved as reality. This lesson is about the specific kind of wrongness that survives a room full of smart, attentive reviewers, and the structured session that kills it: the Walkthrough Verification Protocol.
Why Plausible-Wrong Beats Absurd-Wrong Every Time
If the AI had drawn something absurd into that map, a box labeled "invoice reviewed by marketing intern" or a flow arrow running backwards into the mailroom, the meeting would have caught it in four seconds. Absurd errors are cheap to catch because they trip the pattern-matcher every reviewer carries: does this look like a sensible process? Absurdity fails that test instantly and dies in review.
The invented step that matters is the opposite creature. It passes the sensibleness test with honors, because it was generated from the sensibleness test. A large language model has read thousands of descriptions of accounts payable processes, and in that training corpus, good AP processes have a three-way match, a duplicate-vendor check, and a sign-off threshold. When your interview notes are silent on a detail, the model does not leave a hole. It fills the hole with the most statistically typical content for a process of that shape. The result is a step that belongs in the textbook version of your process, drawn confidently into the map of your actual one.
Here is the psychology of the failure, stated in operator terms rather than academic ones. When a reviewer looks at a process map, the question their brain actually answers is "does this look like a sensible process of this type?" That question is fast, automatic, and feels exactly like scrutiny. But the only question that matters for a map you will make decisions with is "is this OUR process, as it ran last Tuesday?" Those are different questions, and the invented step is engineered, by the very mechanics of how the model works, to pass the first while failing the second. Your reviewers are not being lazy. They are answering the wrong question with full diligence, because the wrong question is the one a conference room can answer. The right question can only be answered by evidence, and the evidence is not in the room.
This matters at the scale of your whole program because of where maps sit in the chain of work. The previous lesson gave you the drafting method: narrative to flow, every step cited to the Source Pack, inferences flagged. This lesson is the deep verification that method promised, and it comes before everything downstream. The baseline you will build in the next lesson hangs numbers on this map. Pilot scoping chooses its target step from this map. Redesign redraws this map. A fabricated step does not stay a drawing error; it becomes a baseline measuring something that does not happen, a pilot aimed at a step nobody performs, a redesign that "removes" work that was never done while missing the work that was. MIT's autopsy of the 95 percent of GenAI pilots with no measurable return found missing workflow integration near the center of the failure pattern, and you cannot integrate into a workflow you have mapped wrong. The map is the aiming device for everything that follows. This is the lesson where you calibrate it.
The Three Species of Map Fabrication
Fabricated steps are not random. In practice they come in three recognizable species, each produced by a different mechanic inside the drafting process, and each carrying a tell you can learn to spot on paper before you ever book a verification session. Learn the species and half your catches happen at your desk.
Species one: the imported step
The imported step is best practice smuggled in as fact. It is the three-way match, the quality assurance (QA) review before release, the segregation-of-duties check, the manager approval gate: steps that appear in every well-written standard operating procedure (SOP) template for that process type, and therefore in the model's sense of what such a process contains. Your sources never mentioned it. The model added it because processes like yours usually have it, in the same way an artist asked to draw "a kitchen" adds a sink without being told.
The tell: no citation traces to the Source Pack. If you followed the previous lesson's discipline, every step in your draft map carries either a citation to a specific source or an explicit [INFERRED] label. An imported step either carries a vague citation that does not survive a re-read ("per interview 2," but interview 2 says nothing of the kind) or arrived without any label at all in a moment of drafting haste. The desk check is mechanical: walk the map step by step, open the cited source for each, and confirm the source actually says what the box says. Any box whose citation dissolves under that check is a suspect until the walkthrough clears it.
Species two: the bridged gap
The bridged gap is a connector invented to make the flow logical. Your interviewee described step A ("the invoice arrives and gets coded") and step C ("then AP resolves the discrepancy"), and never explained how work travels from one to the other, because to them the connection is too obvious to say out loud. The model, needing an unbroken flow, built a bridge: "coded invoice is routed to the AP specialist for review." Plausible. Grammatical. Invented. And dangerous in a specific way: the bridge the model builds is usually a clean, direct handoff, while the real connection is often the messiest part of the process, a shared queue, a weekly batch, a spreadsheet somebody emails on Fridays. The invented bridge does not just add a false step; it papers over the exact spot where the real process hides its delays.
The tell: connector steps with vaguer wording than their neighbors. Steps drawn from real testimony inherit the texture of testimony: system names, role names, specific verbs ("Priya posts it in Coupa and tags the buyer"). Steps invented to bridge a gap have no testimony to inherit texture from, so they come out smooth and generic ("the request is routed for review," "the item is forwarded to the appropriate team"). Read your map like an editor: wherever the prose suddenly goes abstract between two concrete neighbors, mark the box. The vagueness is the watermark of invention.
Species three: the promoted exception
The promoted exception is real, which is what makes it treacherous. One interviewee, once, mentioned a workaround: "if the vendor record is locked we just override the hold and fix it later." That sentence entered the Source Pack legitimately. But somewhere in synthesis, the model promoted it: the workaround got drawn as a main-line step on the standard path, as if every transaction takes the override. Nothing was fabricated, yet the map now describes an exception as the rule, and any baseline or pilot built on it will size the process around behavior that happens three days a quarter.
The tell: single-source steps carrying main-line weight. Check the citation density of your happy path. Steps on the standard flow should be corroborated by multiple sources: several interviewees, the SOP, system data. A main-line step resting on one sentence from one person is either a promoted exception or a genuinely under-evidenced claim, and either way it has no business carrying the weight of the standard path until the walkthrough decides where it really belongs.
| Species | Where it comes from | The tell on paper | What kills it in the walkthrough |
|---|---|---|---|
| Imported step | Textbook best practice filling a source gap | No citation traces to the Source Pack | The negative pass: "when did you last actually do this?" |
| Bridged gap | Model invents a connector between A and C | Vaguer wording than neighboring steps | The gap question: "what happens between these two boxes?" |
| Promoted exception | Real workaround drawn as the standard path | Single-source step carrying main-line weight | The exception census, checked against system data |
The Artifact: The Walkthrough Verification Protocol
Desk checks catch what the paper can confess. For everything else you need this lesson's named artifact: the Walkthrough Verification Protocol, a structured 90-minute session that tests the map against reality instead of against plausibility. It has five movements, and the order is not decorative. Run them in sequence.
Movement one: preparation
Print the map. Physically, on paper, large enough to write on, because the session works through pointing, crossing out, and drawing arrows, and a screen turns participants into an audience. Bring the Source Pack, because disputes get settled by opening the source, not by whoever speaks most confidently. And before anyone arrives, pre-label every step with its citation or its [INFERRED] flag, in writing, on the printed map. This pre-labeling is the single highest-leverage hour of preparation you can spend: it converts the session from a vague "does this look right?" review into a targeted interrogation, because everyone can see which boxes stand on evidence and which stand on inference. Invite the people who actually touch the process, the doers, not only their managers. A manager knows the process as it was designed; the protocol needs the people who know it as it runs.
Movement two: the transaction replay
This is the heart of the protocol, and it is the one test plausibility cannot pass. Take two or three real, recent, completed transactions: an actual invoice with its number, an actual ticket, an actual order. Not hypothetical ones, not "a typical invoice," because a hypothetical transaction gets walked through the map using the same pattern-matching that approved the fabrication in the first place. A real transaction has a history that exists independently of anyone's beliefs about the process.
Now trace each transaction through the map, step by physical step, with the people who touched it, and hold one standard: "show me in the system where this happened." Not "does this step happen?" but "show me where it happened to invoice 48213." For each box: who did this, on what date, in what system, and can we see the record? An imported step fails this test in the most clarifying way possible: the room goes looking for the system evidence of the duplicate-vendor check running against this invoice, and there is nothing to find, because the step does not exist. No amount of reasonableness survives the absence of a timestamp.
Notice what this is, in the language of the Verification Habit lesson from the previous chapter: transaction replay is spot-audit sampling applied to flow logic. You are not checking every transaction the process ever ran, just as the spot audit never re-reads every transcript. You are pulling a small sample and inspecting it completely. And the sample is astonishingly efficient here, for a structural reason: a single replayed transaction exercises every step on its path, which is typically dozens of boxes. Two or three transactions, chosen to include at least one clean case and one exception case, will traverse most of the map and collide with most structural fabrications. This is why the protocol needs 90 minutes and not a week. The sampling math you already learned is doing the heavy lifting.
Movement three: the negative pass
The replay tests the map's steps against the transactions that happened to flow through them. The negative pass tests every remaining step directly, with one question per box, asked to the person who would perform it: "when did you last actually do this?"
The phrasing is load-bearing. "Do you do this?" invites a yes, because the step sounds like something a diligent professional should be doing, and no one enjoys confessing to a room that they skip the duplicate-vendor check. "When did you LAST actually do this?" demands an episode: a date, a specific memory, an instance. Imported steps die here, usually gently: a pause, a glance around the table, and then someone says "I think the system does that automatically?" and someone else says "I don't think it does," and now you know. For every gap between boxes, ask the mirror question: "what happens between these two boxes?" Ask it even where the map shows a clean arrow, especially where the connecting step's wording is suspiciously smooth. Bridged gaps die here, and what replaces them is usually the most operationally interesting content in the whole session: the queue, the batch, the Friday spreadsheet, the two days of waiting that no interview mentioned because waiting does not feel like a step to the people doing it.
Movement four: the exception census
Promoted exceptions cannot be caught by asking whether a step happens, because they do happen. They are caught by asking how often. Put the question to the room: "what percentage of items leave the happy path, and where do they go?" Then, and this is the part that separates the protocol from a conversation, check the answers against data, not memory. Pull the ticket counts, the exception queue volumes, the override logs from the system. Human memory systematically misestimates frequency: painful exceptions feel common because they are memorable, and routine workarounds feel rare because they have become invisible. The census does two jobs at once: it demotes promoted exceptions back to the exception lane where they belong, and it sometimes reveals the opposite and more expensive surprise, an "exception" running at 30 percent of volume that is functionally a second standard path your map does not show at all.
Movement five: sign-off
The session ends with the map earning its papers. Add a verification block to the map document itself: the date of the walkthrough, the transaction identifiers that were replayed, the names and roles of the participants, and the list of corrections made, by species if you want the practice. Then the map enters version control: the corrected map becomes v1.1, the draft it replaces is retired but kept, and any future change goes through the same discipline. The verification block is what makes this map different in kind from every unverified diagram in your organization's shared drive. Anyone who picks it up can see exactly what standard of evidence it met and when. And log the session in your Verification Log from the previous chapter: deliverable, passes run, errors found and fixed, sign-off. A verified map with a visible pedigree is the unit of credibility your whole readiness program is built from.
A process map is verified when real transactions have been replayed through it, not when a room full of smart people has nodded at it.
Worked Example: The Invoice-Exception Map Meets Reality
Here is the protocol run end to end on the map you have been building since the clean-inputs lesson: the invoice-exception process. All numbers are illustrative, chosen to be realistic rather than reported from any specific company.
The session: 90 minutes, six participants (two AP specialists, one procurement analyst, the AP team lead, the process assessor facilitating, one note-taker). Two replayed transactions: invoice 48213, a clean price-mismatch exception resolved in four days, and invoice 47881, a quantity-short exception that escalated to procurement. The printed map has 23 steps, 19 carrying citations, 4 pre-labeled [INFERRED]. Findings, one of each species, which is roughly what a first walkthrough on an AI-drafted map of this size tends to yield:
- One imported step. The map shows a "duplicate-vendor check" the AP system supposedly runs before an exception is opened. During the replay of invoice 48213, the standard is applied: show me in the system where this ran. Nobody can. The AP specialists eventually agree the enterprise resource planning (ERP) system has a duplicate-vendor module, unlicensed and switched off since implementation. The step existed in the vendor's brochure and in the model's training data, and nowhere else. Crossed out. Time to catch: about six minutes.
- One bridged gap. The map shows the coding clerk "emailing the exception to the AP specialist for review." The negative pass asks what happens between the coding box and the review box, and the answer redraws the map: there is no email. Exceptions land in a shared queue, and specialists pick them up when they clear their current items. Average pickup wait, confirmed later against queue timestamps: about two days. The invented email was not just wrong, it was expensively wrong, because a direct handoff hides exactly the delay that the shared queue creates. The map gains a queue symbol and a wait annotation.
- One promoted exception. The map draws a "manual price override" as a standard step on the main path. The exception census asks how often it actually runs, and the override log answers: 11 uses in the last quarter, 9 of them in the final week, all by the team lead, all during quarter-end close. It is a quarter-end pressure valve, not a process step. Demoted to an exception lane with a frequency note.
The corrected map ships as v1.1, verification block attached: date, invoices 48213 and 47881, six named participants, three corrections logged by species. Total cost of the exercise: one 90-minute session plus about two hours of preparation and correction, call it nine person-hours.
Now price the alternative, because this is where the protocol pays for itself a hundred times over. Suppose the unverified map had shipped and the automation pilot had been scoped from it. On the drafted map, the visible villain was the review step, and the natural pilot is an AI-assisted review tool: perhaps $120,000 in licenses and integration and a quarter of effort, all illustrative figures. But the replay revealed that the actual cycle-time villain was the two-day shared-queue wait that the invented email handoff had hidden. An automation pilot aimed at the review step would have accelerated a step that was never the constraint, while every invoice continued to sit in the queue for two days first. Measured cycle-time improvement: approximately nothing, discovered only after the money and the quarter were gone. That is not a hypothetical failure mode; it is the mechanism. Pilots aimed with wrong maps are how workflow-integration failure happens in practice, and workflow-integration failure is what MIT found at the center of the 95 percent of pilots with no measurable return. Nine person-hours of walkthrough against six figures and a quarter of misdirected effort is not a close call. It is the best trade in this program.
The Failure Story: Twelve Maps and a Ruined Word
The success path has a mirror, and it is worth sixty seconds of discomfort. A transformation office at a mid-sized firm, excited by how fast AI drafting made mapping, publishes a pack of twelve process maps in a single quarter. The pack is beautiful. Leadership applauds the velocity; the office presents it as proof of what AI-enabled process work can do. No walkthroughs were run. The maps had been reviewed, in conference rooms, by people who nodded at plausible boxes.
Over the following quarter, three separate teams flag steps in their processes that do not exist. A compliance reviewer finds an approval gate the map promises and the system has never enforced. A team lead discovers her group's map shows a QA review her team has never performed and now looks negligent for skipping. The office quietly recalls the pack for revision. But the recall is not the real damage. The real damage is linguistic: around the company, "the AI maps" becomes shorthand for unreliable, said with a small smile in meetings the transformation office is not in. When the office publishes its next deliverable, verified or not, it prices in that smile. Credibility, once spent, prices every future deliverable, and the office spends the next year buying back at a premium what one quarter of unverified velocity sold at a discount. Every readiness program runs on exactly one currency, and it is not tooling budget.
Set the two stories side by side and notice what actually separated them. Not talent, not tools, not even diligence in the ordinary sense. The difference was 90 minutes per map and the refusal to let plausibility stand in for evidence. This is the program rule you have now met three lessons running, and it bears its full weight here: verify every AI-touched figure, step, and claim. The previous lessons taught you to verify figures and claims. This one closes the loop on steps. And it hands the baton directly to the next lesson, where you will hang cycle times, volumes, and costs on this map to build the process baseline. Numbers hung on a wrong map are wrong numbers, no matter how carefully you measure them. Verify the skeleton before you weigh the body.
What to Do Monday Morning
One map, one session, one week. Here is the sequence.
- Pick one AI-drafted map that matters: one that a baseline, a pilot decision, or a redesign will actually stand on. If you built a map in the previous lesson, use that one.
- Run the desk check and pre-label every step with its citation to the Source Pack or an explicit [INFERRED] flag. Mark the three tells as you go: citations that do not trace, connector boxes vaguer than their neighbors, main-line steps resting on a single source.
- Book the 90-minute walkthrough with the people who touch the process, the doers and not only the managers, and pull two real, recent, completed transactions before the meeting: identifiers, system records, dates. One clean case, one exception case.
- Run the replay, the negative pass, and the exception census in order. Hold the replay standard without flinching: "show me in the system where this happened." Ask "when did you last actually do this?" of every uncorroborated step, and "what happens between these two boxes?" at every gap. Check the exception percentages against system data before the session ends, or assign an owner and a date for pulling them.
- Publish the corrected map with its verification block: date, replayed transaction identifiers, participants, corrections made. Version it as v1.1 and retire the draft visibly, so the unverified version cannot circulate as a ghost.
- Log the session in your Verification Log from the Verification Habit lesson: deliverable, passes run, errors found and fixed by species, sign-off. Three walkthroughs from now, that log is your evidence that verification is a working control in your program, not an aspiration on a slide.
Key Takeaways
- Fear the plausible error, not the absurd one: the invented step survives review precisely because it matches what the process should look like, and reviewers pattern-match "is this sensible?" when the only valid question is "is this OUR process?"
- Learn the three species of map fabrication and their tells: the imported step (no citation traces to the Source Pack), the bridged gap (connector boxes vaguer than their neighbors), and the promoted exception (single-source steps carrying main-line weight).
- Run the Walkthrough Verification Protocol on every AI-drafted map that decisions will stand on: preparation with pre-labeled citations, transaction replay, negative pass, exception census, and sign-off, in 90 minutes.
- Replay two or three real transactions through the map with the people who touched them, holding one standard: "show me in the system where this happened." Replaying reality is the one test plausibility cannot pass.
- Ask "when did you last actually do this?" of every step and "what happens between these two boxes?" at every gap; imported steps die on the first question and bridged gaps die on the second.
- Check exception frequencies against ticket and system data, never against memory, because memory promotes memorable workarounds and hides routine ones.
- Attach a verification block (date, replayed transactions, participants, corrections) and version the corrected map, so anyone who picks it up can see the standard of evidence it met.
- Remember the stakes on both sides: nine person-hours of walkthrough versus a six-figure pilot aimed at the wrong step, and a credibility loss that prices every future deliverable once "the AI maps" becomes shorthand for unreliable.
Skill.re