Incident Playbook: When the AI Step Breaks
At 9:40 on a Tuesday morning, the distribution watch on the invoice-exception dashboard turns amber: the confidence mix in site 2's lane has shifted, and two tripwire cases have hit in three days. Nothing is down. No pager has fired. Every system reports healthy, every queue is moving, and the AI step is producing output that looks exactly as polished as it did last week. In the operations center of a company that runs servers, this is a non-event. In the operations center of a company that runs AI-assisted decisions, this is minute zero of an incident, and the next ninety minutes will reveal whether this organization wrote its playbook in peacetime or plans to improvise one in a forty-person email thread. This lesson is about writing it in peacetime: the five-phase AI Incident Playbook that turns everything this chapter has built (the failure-mode map, the fallback SOP, the escalation ladder, the monitors, the audit trail) into a rehearsed response that runs on numbers instead of adrenaline.
The Failure That Does Not Page Anyone
You almost certainly already know incident management. Systems go down, severity levels page people, a bridge call opens, someone becomes incident commander, and afterward there is a postmortem with a timeline. Operations has run this play for decades, and if your organization is mature, the play works beautifully for outages. That maturity is exactly the trap, because an AI-step incident breaks the outage mold in three specific ways, and a playbook written for outages mishandles all three.
First, AI failures are usually partial and plausible. An outage announces itself: the page does not load, the queue backs up, the alarm fires. A degraded AI step does the opposite. The system stays up, latency stays normal, and the step keeps producing confident, well-formatted, wrong output, often in only one lane of the work: one document type, one site, one vendor's format. No infrastructure alarm will ever fire, because from the infrastructure's point of view nothing is wrong. Detection is entirely the job of the monitors you built in the monitoring lesson, and the playbook's first discipline is believing them. An outage playbook waits for certainty. An AI playbook declares on suspicion, because suspicion is the only signal this failure mode emits.
Second, the damage is retroactive. When a server goes down at 2:00, the damage starts at 2:00. When your monitors catch an AI defect on Tuesday, the defect did not start on Tuesday. It started whenever the input mix shifted or the vendor's model silently changed, and by detection day, days or weeks of decisions may already carry it. An outage playbook is entirely fix-forward: restore service, resume flow. An AI playbook must also look backward: it needs a phase, missing from every outage runbook you have ever used, that reconstructs the damage window and sweeps the decisions made inside it. That phase runs on the audit trail from the previous lesson, and it is the reason that lesson insisted on version stamps and queryable records.
Third, the fix is often not a rollback. Outage response ends with a satisfying reversal: roll back the deploy, restart the service, fail over to the replica. You cannot un-update a vendor's model. You cannot roll back the world when the defect came from a shift in the inputs themselves. The AI response toolkit is different: containment through routing, rule patches around the defect, lane suspensions to manual, and sometimes a renegotiation with the vendor whose silent update caused the whole thing. If your responders walk in expecting a rollback, they will spend the worst hours of the incident discovering that the lever they reach for by habit does not exist.
The answer to all three is the lesson's named artifact: the AI Incident Playbook. It is a written document, owned by a named person, containing five phases, a severity ladder with response clocks, a role roster, communication templates, and a drill schedule. Every element is decided before anything breaks, because the one resource an incident never grants you is time to think clearly. The chapter you have been building is the parts list: the Failure Mode and Effects Analysis, FMEA, the structured exercise that enumerated how your AI step can fail, supplies the scenarios; the standard operating procedure, SOP, contains the manual-fallback section the playbook will execute; the escalation ladder supplies the pre-committed rungs; the monitors detect; the trail records. The playbook is the choreography that makes the parts dance together.
Playbooks are choreography, and choreography is decided before the music starts.
Phases One and Two: Detect, Declare, Contain
Phase One: detect and declare
Inputs to this phase arrive from five directions: a monitor alert, a tripwire hit, a gate escape discovered downstream, a complaint from a customer or counterparty, and, never to be dismissed, a reviewer's bad feeling. That last one matters enough to design for. In lean manufacturing the andon cord is the cord any worker can pull to stop the line, and its power depends on one cultural rule: pulling it is always honored and never punished. Your playbook adopts the same rule in incident form: anyone who touches the workflow can declare, and a declared non-incident is a free drill, never an embarrassment. The moment a false declaration earns someone a raised eyebrow, your detection latency doubles, because the next person with a bad feeling will wait for proof, and proof arrives weeks late in this failure mode.
Declaration thresholds are pre-written as severity definitions, each with a response clock and a roster, so that nobody negotiates severity in the heat of the moment. A workable ladder for an AI-assisted process looks like this (calibrate the specifics to your own risk appetite):
| Severity | Definition | Response clock | Who is in the room |
|---|---|---|---|
| SEV-3 | Quality or error-budget breach contained inside one lane; nothing defective has left the building | Response begins within one business day | Process owner, senior reviewer on the rota |
| SEV-2 | Defective output has escaped with external visibility: a customer, vendor, or counterparty saw or acted on it | Containment within four hours | Add the transformation lead, the sponsor, and a comms owner |
| SEV-1 | Money moved wrongly, a regulatory obligation is touched, or safety is implicated | Immediate; containment within one hour | Add counsel and the privacy owner, by name, with phone numbers |
Notice what this ladder connects to: the escalation ladder you built in the SOP lesson already has a red rung, the pre-committed condition under which the AI step's autonomy is pulled. The playbook does not replace that ladder; it recognizes that the red rung is a declaration. When the ladder says stop, an incident exists by definition, and the playbook takes over from there.
Now the AI twist on detection. Because AI failures whisper instead of shouting, the playbook scripts cheap confirmation steps so that declaring on suspicion costs almost nothing to verify. The single best one: run the fast canary deck now. You built this deck in the drift lesson: twenty representative cases with known correct answers, runnable in about twenty minutes. In peacetime it runs weekly as a drift check. In an incident it becomes a thermometer: a reviewer with a bad feeling plus a twenty-minute canary run either produces "deck is green, log it as a drill, thank the puller" or "three of twenty degraded, we have a real one." The deck converts the vaguest input your detection system accepts (unease) into the crispest (a score), for the price of one coffee break.
Phase Two: contain
Containment means stopping the bleeding without breaking the business, and the way you achieve both at once is a pre-priced containment menu, ordered from smallest intervention to largest:
- Tighten routing thresholds. Lower the confidence bar for automatic processing so more cases route to human review. Smallest lever, fastest to apply, fully reversible.
- Suspend one lane to manual. The affected document type or site drops back to the manual process. This is the SOP's manual-fallback section, executed for real, and it is precisely why the SOP lesson mandated a quarterly fire drill: a fallback that has never been rehearsed takes a day to stand up, and a day is what you do not have.
- Suspend the AI step entirely. Everything goes manual. Expensive, unambiguous, sometimes right.
- Freeze agent autonomy to approve-before-act. For agentic steps, this is the Oversight Matrix's emergency setting: the agent may still draft and propose, but nothing executes without a human click. Gartner's projection that over 40 percent of agentic AI projects will be canceled by the end of 2027 is, in part, a projection about organizations that never installed this switch and then met their first agent incident without it.
The craft is in the phrase pre-priced. Each menu item carries, written in the playbook in peacetime, its cost in throughput and staffing, so the incident commander chooses with numbers rather than dread. Illustrative arithmetic for the invoice-exception workflow: the process handles roughly 1,150 exceptions per month, about 53 per working day. Manual handling averages 12 minutes per exception. Full manual fallback therefore costs about 10.6 clerk-hours per day, call it 11: roughly a person and a half, findable for a week if you know the number, panic-inducing if you are deriving it at midnight. Tightening the routing threshold in one lane, by contrast, might add 40 gate reviews per day at 5 minutes each, about 3.3 reviewer-hours: a much cheaper first move. The commander who has this table picks option one in ten minutes. The commander who does not convenes a meeting.
Phase Three: Scope the Damage Window
This is the phase outage playbooks do not have, the centerpiece of AI incident response, and the payoff for every hour you spent on the audit trail. The question it answers: when did the defect start, and which decisions carry it?
Bounding the window uses three instruments in sequence. The version stamps and the change log give you the outer bound: if a model update, prompt change, or rule change deployed on the 14th, the window cannot start before the 14th (and if the change log shows no internal change at all, your suspect is external: the input mix or a silent vendor update, which is itself diagnostic). The canary history refines it: the fast deck was green on its run of the 12th and degraded on the 15th, so the defect entered in that gap. And then the audit trail enumerates the affected population: every decision in the window, in the affected lane, matching the defect's signature. Read that sentence again as a database operation, because that is what it is: a query. Filter by date range, lane, and characteristics; return the list. Twenty minutes of work, producing a number and a set of case IDs, because the trail was designed in the previous lesson to make exactly this query possible.
Now contrast the firm without the trail, because the contrast is the whole argument. Same defect, same detection day. When did it start? Nobody knows; the vendor is asked, the vendor is vague. Which decisions are affected? There is no per-decision record of which version processed what, so the answer is archaeology: pull every case from a guessed window, re-review all of them by hand, and pray the guess was generous enough. The trail-less firm does not have a damage window; it has a damage fog, and fog is expensive in exactly the currency incidents charge: reviewer hours, counterparty patience, and legal exposure if the guess turns out short.
With the affected population enumerated, the remediation sweep begins, and here the severity logic runs backward. You triaged incoming work by consequence when you designed the human gate; now triage the look-back the same way. High-consequence decisions first: anything that moved money, touched a regulated obligation, or went to an external party gets re-reviewed immediately. Routine internal decisions follow in batches. Every corrected case is logged to the trail (the incident is itself an auditable event), and any counterparty affected is made whole through the communications phase, which is where we go next.
One design note that saves future pain: write the damage-window query as a runnable, saved procedure, not a description. The middle of an incident is the wrong time to remember which field holds the model version. The playbook should say, in effect, run query Q3 with these two dates and this lane code, and a drill should have proven that Q3 actually runs.
Phases Four and Five: Communicate, Then Learn
Phase Four: communicate from templates written in peacetime
Incident communications drafted under panic are reliably terrible: too defensive, too vague, or too candid in the wrong places. The playbook therefore carries three templates written on a calm afternoon.
- The internal incident brief. One page, four headings: what broke, what is contained, what is being swept, when normal operations resume. This is the governance-visibility discipline from the 40-percent-curve lesson applied in incident form: sponsors who hear about problems from you, early, with a plan attached, fund your program; sponsors who hear about problems from the hallway cancel it.
- The counterparty note. The vendor whose invoices were misrouted, the customer whose case was mishandled, gets a factual account: what happened to their items, what has been corrected, what prevents recurrence. Candor sized to consequence: a two-day internal-catch defect earns a different note than a six-week external escape, but both notes exist as templates so that neither is composed at midnight.
- The regulatory-touch escalation. For SEV-1, the playbook names the counsel and the privacy owner in the roster, with the decision pre-made that they are contacted before any external statement goes out. The playbook knows who to wake so the commander does not have to decide.
And one trap named in advance, because it will tempt every drafter: do not write "the AI made an error" as a causal account, internally or externally. Say what the system did and what the design missed: "the workflow misclassified 61 exceptions after an unannounced vendor model update that our monitoring caught within two days; we have added contract-notification checks." Blaming the model is blaming the hammer. It reads as evasion to outsiders, it teaches the organization nothing, and it quietly erodes the accountability principle this whole program is built on: a human owner made design choices, and the design, not the tool, is what gets fixed. Language discipline in the templates becomes language discipline in the culture.
Phase Five: the postmortem that feeds the machine
Blameless postmortem is a term of art from operations, and it means something sharper than being nice: the timeline is reconstructed from the audit trail (not from memory), the five-whys analysis lands on design causes rather than people, and the output is change, not catharsis. MIT's autopsy of the 95 percent of stalled GenAI pilots found the missing learning loop to be a defining failure: systems and organizations that do not improve from feedback decay. An incident is the most expensive feedback your workflow will ever generate. It is tuition, and tuition either compounds into capability or evaporates into a document nobody reads.
The playbook makes it compound with an output contract: every postmortem must produce, as a checklist with named owners and dates,
- one FMEA row added or rescored (the incident either revealed a failure mode you missed or repriced one you had),
- one or more canary cases added from the actual trigger cases (the defect that fooled you joins the deck that screens for it forever),
- one monitor threshold, rule, or routing patch shipped,
- one edit to the playbook itself (something in the response was slower or vaguer than it should have been), and
- the drill scenario for next quarter, taken from what just happened.
Five artifacts, every incident, no exceptions. This is BCG's 10-20-70 arithmetic (10 percent algorithms, 20 percent technology and data, 70 percent people and process) showing up one more time: notice that four of the five contract outputs are process artifacts, not model changes.
The drill program and the roles
A playbook that has never been rehearsed is a hypothesis. The drill program is quarterly, one hour, alternating scenarios drawn from the FMEA's top-ranked rows, exactly as that lesson promised its rows would someday be used. Format: tabletop walkthrough of the five phases plus one live element, and the live element is the point. Actually run the manual fallback on real volume for 30 minutes. Actually execute the damage-window query against a hypothetical two-day window. The rehearsal is what finds the expired system access, the clerk who transferred teams, the query that references a renamed field. Small drills prevent large improvisations. Drill findings are logged and fixed exactly like incident findings, on the same output contract.
Roles stay deliberately simple, because AI incidents at process scale do not need a war room bureaucracy: an incident commander (the senior reviewer on the rota, by default), a scribe (light duty, since the audit trail records most of the timeline automatically), a comms owner (activates at SEV-2), and one pre-made decision rule for when the transformation lead and sponsor are pulled in: severity, not vibes. SEV-3 is handled and reported; SEV-2 and above, they are in the room. Nobody should ever spend incident minutes deciding whether the situation is "big enough to bother the boss." The ladder already decided.
One Incident Start to Finish, and One Improvisation
The boring triumph (illustrative numbers throughout)
Here is the full play run once, continuing the vendor-portal thread from the pilot lessons, now at production scale in the invoice-exception workflow.
Day 0, 9:40. The distribution watch flags a confidence-mix shift in site 2's lane; the log shows two tripwire hits in three days. The reviewer on rota pulls the cord. 10:05: the fast canary deck runs. 10:30: three of twenty canaries degraded, all in site 2's document format. 11:10: SEV-3 declared: budget breach, contained in lane, nothing external. Ninety minutes from first amber to declaration, and the playbook's culture treats that latency as a stat to celebrate, not a lapse to explain: the whole point of the monitors is that this moment happens in hours, not weeks.
Day 0, afternoon: containment. The commander reads the pre-priced menu and picks the smallest sufficient lever: routing thresholds tightened for site 2's lane. Known cost, printed in the playbook months ago: about 40 extra gate reviews per day, 3.3 reviewer-hours, absorbable for two weeks. Containment is live the same hour. No meeting was required.
Days 1 to 2: scoping. The change log shows no internal deployment in three weeks, so suspicion turns outward. The vendor is queried through the escalation contact; by day 2 the vendor confirms a model update deployed two days before day 0, unannounced. Window bounded: the canary history was green four days before day 0, so the defect entered with the update. The saved trail query runs against the two-day window, site 2's lane: 61 affected decisions. Triage re-reviews them high-consequence first: 4 required correction, all caught internally at the gate before anything left the building. Zero external escapes. The sweep closes in two days.
Days 2 to 4: comms. The one-page internal brief goes to the sponsor on day 2 (what broke, contained same day, 61 swept, 4 corrected, normal by day 6). The vendor escalation goes out citing the change-notification clause that this contract, unlike the one in the drift lesson's failure story, actually contains, because someone asked the Level 1 due-diligence question about update notification before signing. The clause converts an apology into a commitment: advance notice of model updates, in writing, henceforth enforced with a monitoring check.
Day 6: postmortem, on the output contract. FMEA row "vendor silent update" rescored upward on likelihood; the three degraded canary cases join the deck permanently, making the deck sharper against exactly this defect; a contract-notification monitor is added (a weekly check of the vendor's release feed against received notices); the playbook's vendor-escalation step gets a tightened turnaround demand; and next quarter's drill is set: silent-update scenario. Total elapsed: six business days. External damage: zero. Cost: some reviewer hours, all pre-priced. Nobody outside the team will ever tell this story, which is precisely the triumph. Well-run AI incidents are boring, and boring is the deliverable.
The improvised version
Now the same maturity of tooling with none of the choreography, compressed from a pattern you will recognize. A financial-services firm has decent monitors and no playbook. The monitors catch a real SEV-2: fee waivers misapplied by an AI-assisted step, customer-visible. Then the improvisation begins. Three days go to deciding who owns the response, because no roster exists. Containment options are debated in a forty-person email thread while the defect keeps running, because no one has the authority the playbook would have granted or the pre-priced menu that would have made the choice obvious. The damage window is guessed, not queried, because decision records lack version stamps: they sweep six weeks to be safe. The actual window was eleven weeks, a fact eventually established by a customer's lawyer. And the external statement, drafted at midnight under panic, contains the sentence "the AI malfunctioned," a phrase that anchors every subsequent press mention and reads, to the regulator who later asks questions, as an admission that nobody was accountable for the system's design.
Run the diagnosis: every individual capability existed. Monitors detected. People were competent. Records existed, if raggedly. What did not exist was the choreography: the roster, the clocks, the menu, the query, the templates. That is the whole case for this lesson's artifact in one contrast. Capabilities without choreography produce a slow, loud, expensive incident; choreography turns the same capabilities into six boring days.
What to Do Monday Morning
- Write your severity definitions with response clocks. Three levels, one paragraph each, using the SEV-3/2/1 pattern from this lesson: contained breach, external escape, money-regulatory-safety. Attach a response clock and a roster to each. One page, and the hardest decisions of your next incident are already made.
- Pre-price your containment menu. For each of the four levers (tighten routing, suspend a lane, suspend the step, freeze autonomy), compute the throughput and staffing cost using your real volumes, including the manual-fallback arithmetic: monthly volume, minutes per case, clerk-hours per day. Put the numbers in the playbook, not in someone's head.
- Run the damage-window query drill. Pick a hypothetical two-day window and one lane, and actually execute the query against your audit trail: every decision, in the window, in the lane. Time it. If it takes more than an hour or cannot be done at all, you have found this quarter's most important gap while it is still free to fix.
- Draft the three comms templates in peacetime language. Internal brief, counterparty note, regulatory escalation. Apply the language discipline: what the system did and what the design missed, never "the AI made an error" as an explanation.
- Book the first quarterly drill. Take the top row of your FMEA as the scenario, one hour on the calendar, tabletop plus one live element: 30 minutes of real manual fallback or one live query execution. Log the findings like incident findings.
Key Takeaways
- Recognize the three ways AI incidents break the outage mold: failures are partial and plausible (no alarm fires by default), damage is retroactive (weeks of decisions may already carry the defect), and the fix is often containment and redesign, not a rollback.
- Build the AI Incident Playbook in peacetime: five phases, a severity ladder with response clocks and rosters, pre-priced containment options, comms templates, and a quarterly drill schedule, all owned by a named person.
- Declare on suspicion and honor the andon cord: a declared non-incident is a free drill, never a punishment, and the twenty-minute fast canary deck converts a reviewer's unease into a score before lunch.
- Pre-price every containment lever, from tightened routing thresholds to full manual fallback, so the incident commander chooses with staffing arithmetic (11 clerk-hours per day at 1,150 exceptions monthly, in the illustration) instead of dread.
- Treat the damage window as the AI-specific phase: bound it with version stamps and canary history, enumerate affected decisions with a saved audit-trail query, and sweep them by consequence, high-stakes first.
- Enforce language discipline in every communication: describe what the system did and what the design missed, because "the AI made an error" blames the hammer, anchors hostile coverage, and teaches the organization nothing.
- Bind every postmortem to the output contract: one FMEA row, new canaries from the trigger cases, one monitor or rule patch, one playbook edit, and next quarter's drill scenario, so the incident's tuition compounds instead of evaporating.
- Drill quarterly with one live element, because rehearsal is what finds the expired access and the untested query; capabilities without choreography improvise, and choreography is decided before the music starts.
Skill.re