Crisis Management for AI Failures
The alert came at 6:47 on a Tuesday morning. Ernesto Fuentes, the Chief Information Officer at a regional transit authority serving 1.1 million daily riders, was in his car when his deputy called. The AI-based scheduling algorithm that managed bus operator assignments had generated a shift schedule the night before that was facially legal but operationally impossible. It had assigned 23 operators to overlapping routes, created four routes with no assigned operator, and violated the rest-time provisions of the collective bargaining agreement for 11 drivers. The algorithm had been running for four months without incident. The failure had no obvious cause. Forty-one routes were scheduled to begin service in three hours and forty minutes. Ernesto had a crisis playbook for bus fires. He had one for cyberattacks. He had nothing for the failure mode his own AI system had just demonstrated. This lesson is the playbook he built afterward.
Why AI Failures Need Their Own Crisis Framework
Traditional government crisis management is designed for physical events: natural disasters, infrastructure failures, security incidents. AI failures share some characteristics with those events but differ in three critical ways. They are often invisible until a downstream consequence surfaces. They can affect thousands of individual decisions simultaneously before anyone notices the pattern. And the responsible party, whether vendor, agency or the model itself, is rarely clear without investigation. A framework built for a burning bus assumes the event announces itself. An AI failure does not.
That third difference is the one that damages agencies most. In a physical crisis, the question of who acts is settled by the nature of the event: the fire department fights the fire. In an AI failure, the first hour is frequently consumed by an argument about whether the problem belongs to the program office that owns the outcome, the technology office that owns the system, or the vendor that owns the model. Every minute of that argument is a minute the failure continues to operate. The framework's real job is to answer the ownership question in advance so the response can start before the diagnosis does.
Detection Is the Harder Half
The transit scheduling failure was invisible until a supervisor ran a manual cross-check at 6:30 a.m. By that point the bad schedule had already been transmitted to operator devices. The failure had not triggered any system alert, because the algorithm's output looked syntactically correct: it produced a complete schedule with no missing fields. The problem was semantic. The schedule was internally inconsistent in ways that only became apparent when checked against operational reality.
This is the general shape of AI failure detection in government. Format validation tells you the system produced something. It does not tell you the system produced something true, lawful or workable. A benefits eligibility model returns a well-formed determination for every applicant, including the ones it got wrong. A routing model returns a complete route, including the one that crosses a closed bridge. If your monitoring only asks whether output arrived and whether it parsed, your first detector is a member of the public, a union representative, a caseworker or a journalist, and by then the incident already has an audience.
A crisis framework for AI failures must therefore address two distinct problems: detection, meaning how you come to know something has gone wrong, and response, meaning what you do in the window between detection and public impact. Most agencies invest in the second and neglect the first, which is backward, because the size of that window is set almost entirely by how quickly you detect. Ernesto's window was whatever a single sceptical supervisor's manual cross-check had bought him, on a schedule that had already left the building.
Deciding What Counts as a Crisis
Before the playbook can run, somebody has to say the word. That sounds trivial and it is where a surprising number of AI incidents go wrong, because the first person to see the anomaly is rarely senior enough to declare anything and is usually being asked, implicitly, to justify the disruption they are about to cause. If the only way to trigger the framework is to convince a leader that this is serious, the framework triggers late, and it triggers least often for exactly the ambiguous cases where early action would have helped most.
Ernesto's algorithm ran for four months without incident, which is a claim worth interrogating. It ran for four months without a detected incident, on a system whose only detector turned out to be a human being with a habit of checking. Model drift was degrading the output during that period, so the honest statement is that the agency did not know. Agencies routinely describe systems as clean when what they mean is quiet, and the gap between those two words is where a crisis accumulates.
The practical fix is to lower the cost of raising a hand and to make the declaration a defined act with a defined trigger, so nobody has to argue their way into it. Write down, per system, the observations that automatically start the playbook, and make explicit that starting it is not an accusation and carries no penalty if the anomaly turns out to be nothing. The transit authority's staff will only surface the next semantic failure if surfacing the last one was uneventful for the person who did it.
The AI Crisis Playbook: Five Phases
The playbook Ernesto built afterward runs in five phases. The phase boundaries matter more than the phase names, because each boundary marks a decision that gets much harder to make once the next phase has started. The timings below are the ones his authority adopted. Treat them as one agency's commitments, calibrated to a service that begins at a fixed hour, rather than as a standard to copy unexamined into a program with a different clock.
Phase one: immediate containment, first 30 minutes. The first priority is limiting the blast radius. For Ernesto this meant suspending the AI scheduling system's authority to transmit schedules and reverting to the last known-good schedule while a manual assessment began. Containment does not mean fixing the problem. It means stopping the problem from getting worse. Every AI system with operational authority, whether scheduling, routing, case assignment or eligibility determination, needs a documented containment procedure. Who has authority to halt it? What is the fallback state? How long does fallback operation remain viable? Those questions need written answers before the crisis, not during it.
That third question is the one agencies forget, and it is the one that turns a contained incident back into an uncontained one. Falling back to manual scheduling is survivable for a morning and unsustainable for a month. Write down the horizon of your fallback in the same document that authorises it, because the person invoking the fallback at dawn is not the person who will discover, weeks later, that nobody planned for what comes after it.
Phase two: impact assessment, first two hours. Once the system is contained, the agency needs to understand the scope. How many decisions, records or operational outputs were affected? What is the downstream consequence of each category? Are any affected outputs irreversible, for example a benefit denial already communicated to a constituent? For the transit authority the assessment took 90 minutes and identified 41 routes as operationally affected, with 23 operator conflicts requiring manual resolution before service start. No route had yet begun service, which meant the window for remediation was still open.
Sort affected outputs by reversibility rather than by count, because reversibility, not volume, determines what the agency owes the people on the other end. A large batch of wrong outputs that were never acted on is a data cleanup. A handful that reached constituents and changed what those constituents did is a remediation obligation, and possibly a legal one. Ernesto's 41 routes were fully reversible only because the failure was caught before the first departure, which was luck rather than design, and he says so when he tells the story.
Phase three: remediation and communications, two to four hours. Remediation is the technical fix. Communications is the parallel track that runs simultaneously and is equally important. For Ernesto, remediation meant manual rescheduling of the 41 affected routes. Communications meant notification to operators before they arrived at their assigned posts, notice to the union representative, required under the collective bargaining agreement for any system affecting scheduling, and an internal briefing to the transit authority's board chair. No press statement was issued in this phase because service was restored before the first route's scheduled departure. Had service gaps occurred, a public-facing communication would have been required within the first two hours.
Phase four: oversight notification. Federal transit agencies with FTA (Federal Transit Administration) funding have reporting requirements for significant operational incidents. State-level transit authorities typically report to a state transportation board or oversight commission. AI-related failures are a new category that most notification templates do not address explicitly, but they should be treated as operational incidents for notification purposes. Ernesto's agency submitted an incident report to the state transportation board within 72 hours, describing the AI scheduling failure as an operational incident and noting that an internal technical review was underway.
Phase five: root cause analysis and policy response, within 30 days. The root cause of the transit scheduling failure turned out to be model drift. The algorithm had been trained on pre-pandemic ridership patterns and had not been retrained after route changes implemented eighteen months earlier. The model's performance had degraded gradually, and Tuesday morning was the first time the degradation crossed a threshold visible to operational staff. The policy response included a mandatory quarterly retraining schedule, a new automated cross-check verifying schedule consistency against the rest-time provisions before transmission, and a formal incident report shared with the AI vendor as required under the contract.
Be honest about what that cross-check buys. It catches the failure mode you have already seen. It does not catch the next one, because the next one will be a different kind of inconsistency and the check was written against this one. Every post-incident control is a fossil of the last incident. That is a good reason to build them and a bad reason to relax, and the agencies that get into trouble twice are usually the ones that mistook a specific fix for general coverage.
Running the Full Crisis Simulation
A playbook that has never been run is a hypothesis. Ernesto now runs the transit scenario as a full simulation, and the value is not in confirming that the phases work. It is in finding the small, unglamorous obstacles that only appear when people try to execute under time pressure: the containment authority who is on leave with no delegated alternate, the fallback schedule that lives on a share nobody in operations can reach, the union contact number that changed last year.
Build the simulation from a real failure mode of a real deployed system rather than from a generic outage. Give the participants the same partial information the real event would give them, which means starting with a symptom and no cause, and hold back the root cause entirely, because in the real incident nobody had it either. Then run the clock. The point is to make people execute phase one before they understand phase five, which is exactly the discomfort the real morning produces and exactly the skill the playbook is meant to build.
What a completed simulation proves is narrow and worth stating plainly. It proves that this team, on this scenario, on this day, could execute these steps. It is evidence, not assurance. A tabletop that finishes cleanly has told you nothing about the failure modes you did not script, and a crisis plan sitting in a binder has never contained an incident on its own. What the plan actually does is remove decisions from the moment they would be made worst: under time pressure, with partial information, by whoever happened to answer the phone. Design and rehearse it for that purpose and it earns its keep. Sell it as containment and it will disappoint you at the worst possible time. This lesson pairs with AI Incident Response Planning, which develops severity classification and the standing readiness infrastructure in detail; treat that lesson as the plan and this one as the crisis.
Communications During an AI Failure
Communications during an AI failure is complicated by a fact that does not apply to most crises: the cause is often technically opaque. Ernesto could not tell his board chair on Tuesday morning exactly why the algorithm had produced a bad schedule. He could tell her what had happened, what the immediate impact was, what remediation was underway, and what the next steps were. That is enough for an initial communication. It is also the honest answer, and honesty in the first communication is the foundation of credibility in every subsequent one.
The failure mode to avoid is the confident wrong explanation. Under pressure, a technically fluent official will reach for a plausible cause because plausible sounds better than unknown. If that cause turns out to be wrong, and early causes frequently are, the correction becomes the story and every later statement is discounted. Say what you know, say what you do not know, say when you will next say something, and then meet that commitment. A named next update at a stated hour buys more patience than any amount of reassurance.
Media response to AI failures tends to follow a specific pattern: initial coverage of the incident, followed by a secondary story about who is accountable and whether the system should continue operating. Agencies that communicate proactively, acknowledging the failure, describing the impact, explaining the response and committing to a root cause review, control the frame of the secondary story. Agencies that communicate reactively or defensively hand that frame to the reporter, the union or the advocacy group with the most pointed critique.
Internal communications deserve the same discipline. Frontline staff who learn about the failure from a rider, a claimant or a news alert stop trusting the chain, and the next time one of them notices something odd they will hesitate before escalating. That hesitation is expensive, because frontline staff are your most sensitive detector for the semantic failures your monitoring misses. Tell your own people first, in plain terms, including the part where you do not yet know why.
Remediation and Recovery Are Different Jobs
Remediation restores the output. Recovery restores the standing of the system, and the two are routinely confused because remediation is the one with a visible finish line. Ernesto's 41 routes were rescheduled by hand and service ran. That was remediation. Recovery was the harder question that arrived the following week: on what basis does the AI scheduler resume operating, and who decides?
Answer that question before you need it, in the same document that grants containment authority. A defensible return-to-service decision names the person who makes it, states what evidence they require, and specifies what the system is permitted to do first. Staged resumption is usually the right shape: the model returns in an advisory mode where its output is reviewed before it takes effect, runs that way for a defined period against real conditions, and regains operational authority only on the strength of that record. Resuming straight to full authority because the vendor has shipped a patch is how an agency turns one incident into two.
Recovery also has a human component that no runbook captures. The staff who spent a morning cleaning up after an algorithm now have a view about that algorithm, and if the return to service is announced rather than explained, the shadow process they built during the outage will quietly persist. Show them the evidence you used. The people who caught the failure are the people whose confidence you most need back.
Notification Clocks and Whose They Are
The 72 hour figure in Ernesto's playbook, and every other clock in it, is a commitment his agency made to itself. It is not a statutory deadline, it is not a legal test, and meeting it does not establish that the agency has discharged any duty imposed by law. Write that sentence into your own plan in those terms, so that nobody reads an internal target as a safe harbour. Where a statute, a regulation or a grant condition sets a shorter or more specific reporting obligation, that obligation governs and the internal commitment is simply the floor for everything the obligation does not reach.
Do not originate the clock yourself. The work is to map the reporting duties that already apply to this system, this funding stream and this data, list them by name with the specific timing beside each, and have counsel confirm the list before an incident rather than during one. Ernesto's authority reports to a state transportation board; a federally funded transit agency answers to a federal one as well; an agency handling health or tax information carries obligations that neither of those examples touches. The pattern that transfers is proactive posture, not the number. Filing before you are asked, with an honest statement that the technical review is still underway, is what maintains credibility with an oversight body. Waiting for an inquiry converts a technical failure into a governance failure, and oversight bodies punish the second far more reliably than the first.
Legislative and inspector general reporting has its own conventions, covered in Congressional and IG Reporting on AI. What matters for crisis management is that the first notification is usually written under the worst conditions the whole episode will produce, by people who are also doing the remediation. Draft the skeleton in advance. A template with the fields already decided, incident description, systems affected, population affected, actions taken, current status, next review date, converts a long drafting exercise into a short one at the point when drafting time does not exist.
Preparing Before the Crisis
Ernesto's transit authority now maintains a two-page AI incident response card for every deployed AI system. The card names the system, its operational scope, the containment authority meaning who can halt it, the fallback procedure, the notification chain in sequence, and the impact assessment protocol. The cards are reviewed every six months. Every senior operations staff member has read them. The union was briefed on the card for the scheduling system. It took four hours to produce the first draft of the first card and about forty minutes for each subsequent one.
The two-page limit is doing real work and is worth defending against the natural drift toward comprehensiveness. A card that fits on two pages gets read at 6:47 in the morning by somebody standing up. A long annex does not, which means the long annex is a compliance artifact rather than an operational one. Everything that does not fit belongs in the supporting procedure the card points to, and the card should point by name to a document somebody can actually find.
Two habits keep the cards honest. Review them when the system changes rather than only on the calendar, because a model retrained on new data or handed to a new vendor is not the system the card describes. And confirm the names on them: a notification chain full of people who have all moved on is worse than no chain at all, because it produces false confidence and then a scramble. Four hours is not a large investment. It does not prevent a 6:47 a.m. phone call, and nothing in this lesson does. It prevents that phone call from finding an agency with no idea who is allowed to press stop.
Anti-Patterns to Avoid
- Treating the written plan as containment. A crisis playbook that exists has not stopped anything. The plan's function is to move decisions out of the moment they would be made worst, which only pays off if the authorities in it are real, current and exercised. Agencies that report having a plan and cannot say who is permitted to halt a specific system have the artifact without the capability.
- Treating a clean tabletop as proof of readiness. A simulation that finishes on time tells you this team handled this scripted scenario. It is evidence, not assurance, and the scenarios you did not write remain untested by definition. The productive response to a clean run is another scenario, not a slide reporting that the exercise passed.
- Waiting for root cause before containing. The instinct to understand before acting is a good engineering instinct and a bad crisis instinct. Ernesto's containment happened before anybody knew about model drift, and it had to, because the alternative was letting a known-bad system keep transmitting while the investigation ran.
- Monitoring format instead of meaning. A validator that confirms every field is populated will pass the schedule that assigned 23 operators to overlapping routes. If nothing checks output against operational reality, your detection layer is whoever is unlucky enough to be downstream.
- Offering a confident cause you do not have. The plausible early explanation that turns out to be wrong costs more credibility than the honest admission of uncertainty ever would, because the correction becomes the second story and discounts everything you say afterward.
- Declaring victory at remediation. Restoring the output is not restoring the system. Without a stated return-to-service standard and a named decision maker, resumption happens by default, usually on the strength of a vendor assurance, and the shadow processes staff built during the outage never go away.
- Reading an internal clock as a legal one. An agency that meets its own 72 hour commitment and misses a shorter statutory obligation has not complied with anything. Map the duties that actually apply, by name, before the incident.
Practice Prompts
- Containment inventory. List every AI system in your agency that has operational authority, meaning it produces an output that takes effect without a human approving it individually. For each one, write down who may halt it, what the fallback state is, and how long the fallback can run. Any row you cannot complete from existing documents is your first card.
- Semantic detector design. Pick one deployed system and write several checks that would catch an output that is correctly formatted and operationally wrong. Then identify what each check would have missed in your last known incident.
- Reversibility sort. Take a hypothetical failure of one system and sort the affected outputs into reversible, reversible with effort, and irreversible. Decide now what the agency owes the people in the third category, because deciding it during an incident produces inconsistent answers.
- Clock mapping. With counsel, list every reporting obligation that attaches to one specific system: statutory, regulatory, grant condition, contractual, and internal. Put the timing beside each and mark which is shortest. Do not add a clock that no source imposes.
- Cold-start simulation. Run your worst realistic failure with the participants given only the symptom, no cause, and the real clock. Record every obstacle that was about a person, a permission or a phone number rather than about the technology. Those are the findings worth fixing this month.
- First notification skeleton. Draft the oversight notification for an incident that has not happened, filling every field except the specifics. Time how long it takes. That time is what you have just removed from your worst morning.
Reflection Questions
- If your highest-consequence AI system produced a well-formed, entirely wrong output tonight, what would detect it, and how many hours would pass first?
- Who in your agency is currently authorised to halt that system, and would they believe they were authorised at 6:47 in the morning without calling anyone?
- What would your first communication say if you genuinely did not know the cause, and who would have to approve it before it went out?
- On what evidence would you let a suspended AI system resume, and has anyone ever written that standard down?
- Which of the deadlines in your incident procedures came from a statute, a grant condition or a contract, and which came from somebody's judgement about what sounded responsible?
Glossary
- Containment. The act of stopping a failing system from producing further consequential output, by suspending its authority and reverting to a fallback state. Distinct from remediation, and always earlier.
- Blast radius. The set of decisions, records, people and downstream systems that a failure has already touched or could still touch if the system continues operating.
- Fallback state. The defined mode of operation the service runs in while the AI system is suspended, together with the period for which that mode is sustainable.
- Semantic failure. An output that is structurally valid and substantively wrong. It passes format validation and fails contact with operational reality.
- Model drift. Gradual degradation of a model's performance as the conditions it was trained on diverge from current conditions, typically invisible until the degradation crosses a threshold somebody notices.
- Impact assessment. The structured determination of how many outputs were affected, what each category of affected output caused downstream, and which effects are irreversible.
- Return to service. The decision to restore a suspended system's operational authority, ideally staged through an advisory period and made against stated evidence by a named person.
- Incident response card. A short standing document, two pages in Ernesto's authority, recording one system's scope, containment authority, fallback, notification chain and assessment protocol.
Related Lessons
- AI Incident Response Planning builds the standing plan this lesson executes against, including severity classification and readiness infrastructure.
- AI Incident Response: What to Do covers the individual responder's actions once an incident is under way.
- AI Incident Documentation and Response develops the post-incident record that phase five depends on.
- Congressional and IG Reporting on AI covers the conventions of legislative and inspector general notification.
- Enterprise Risk Frameworks for AI places crisis capability inside the agency's wider risk apparatus.
- Continuous Monitoring Fundamentals supplies the detection signals that determine how large your response window is.
- When Government AI Goes Wrong examines failures that were handled badly and what the handling cost.
- Building and Maintaining Public Trust covers the credibility that crisis communications either preserve or spend.
Closing Thoughts
Ernesto built his playbook after the morning that needed it, which is the ordinary way agencies acquire one and the expensive way. The four hours the first card cost him were available at any point in the four months the algorithm ran cleanly. They were not available at 6:47 on a Tuesday, and that is the whole argument.
What a crisis framework for AI actually delivers is narrower than the word crisis suggests and more valuable. It does not prevent the failure. It does not contain the failure. It decides, in advance and calmly, who may act, what they may do, what falls back to what, who gets told, and in what order, so that the people living through the failure spend their attention on the situation instead of on the org chart. Everything else in this lesson is detail attached to that one idea.
Key Takeaways
- AI failures are often invisible until downstream. Syntactically correct outputs can be operationally disastrous. Detection has to check whether the output makes operational sense, not merely whether it parsed, or your first detector will be a member of the public.
- Containment comes first, before root cause. In the first 30 minutes, stop the problem getting worse: suspend operational authority, revert to the last known-good state, begin the impact assessment. Understanding can wait; a running failure cannot.
- Write down how long the fallback lasts. A containment procedure that names the authority and the fallback but not its horizon leaves the agency to discover the limit by hitting it.
- Sort impact by reversibility, not volume. What the agency owes affected people is determined by whether the wrong output reached them and changed what they did, not by how many wrong outputs there were.
- A plan does not contain an incident and a clean tabletop does not prove readiness. Both are evidence about scenarios you chose to consider. Their real value is removing decisions from the moment they would be made worst.
- Communications must be honest, not complete. You will not know the cause in the first two hours. Say what happened, what the impact is, what you are doing, and when you will next speak, then meet that commitment. Tell your own staff first.
- Remediation is not recovery. Restoring output is separate from restoring the system's standing. Decide in advance who authorises return to service, on what evidence, and stage the resumption through an advisory period.
- Notification clocks are commitments unless a law says otherwise. Ernesto's agency filed within 72 hours by its own policy. Map the statutory, regulatory, grant and contractual duties that actually apply, by name, and let the shortest one govern.
- Proactive oversight notification protects credibility. Filing before you are asked, while stating openly that the technical review is unfinished, keeps a technical failure from becoming a governance failure.
- Model drift is a maintenance problem, not a deployment problem. The transit model degraded for eighteen months after route changes it was never retrained on. Retraining schedules and consistency checks are infrastructure, not enhancements.
- Every post-incident control is a fossil of the last incident. The new cross-check catches the failure you have seen. Build it anyway, and do not mistake it for coverage of the next one.
- Keep the card to two pages and keep the names current. A short card gets read while somebody is standing up at dawn. A notification chain full of people who have left is worse than none, because it fails confidently.
Frequently Asked Questions
Is 72 hours the deadline for reporting an AI failure?
No. In this lesson 72 hours is the commitment one transit authority made to itself for filing with its state oversight board, and it is presented as that. It is not a statutory clock and meeting it discharges no legal duty by itself. Your actual obligations come from the statutes, regulations, grant conditions and contracts that attach to your specific system and data, several of which may be considerably shorter. Have counsel list them by name in advance and let the shortest one govern. AI Incident Response Planning treats notification within 72 hours of confirmation as an emerging practice norm, explicitly distinct from any statutory clock, and the same distinction applies here.
How is this different from our existing continuity and emergency procedures?
Existing procedures generally assume the event announces itself and that responsibility follows obviously from its nature. AI failures often do neither. They surface downstream, they touch many decisions before anyone sees the pattern, and ownership is genuinely contested between the program office, the technology office and the vendor until somebody has decided it in writing. Reuse your existing escalation and communications machinery, but add the detection layer and the pre-assigned containment authority, because those are the two places the general procedure does not reach.
Who should hold containment authority?
Someone reachable at any hour who can act without convening a committee, with a named alternate who is not the same person's deputy on the same leave schedule. The specific role varies by agency and by how much operational consequence the system carries. What matters more than the title is that the authority is written down, that the holder knows they have it, and that they believe they can use it without seeking permission first. An authority that requires a phone call to confirm is not an authority at 6:47 in the morning.
Should we suspend the system for every suspected failure?
That is a judgement your severity framework should make for you rather than one to improvise. Suspension has its own costs, and a service that falls back to manual operation is degraded rather than safe. The workable pattern is to decide the suspension question separately from and faster than the notification question, on a documented threshold, so that containment never waits on a notification chain. AI Incident Response Planning develops severity tiers and the different tempos attached to them.
Our AI system is run by a vendor. Does that change the playbook?
It changes who executes some steps and none of who is accountable for them. The agency still owns the outcome, still owes the notifications, and still has to be able to halt the service. Confirm in the contract that the agency can suspend without vendor agreement, that incident information flows to the agency on a defined timeline, and that the agency receives what it needs for its own root cause analysis. Ernesto's authority shared its formal incident report with the vendor as its contract required, which is a different thing from waiting for the vendor's version of events.
How often should the incident response cards be reviewed?
Ernesto's authority reviews them every six months. The more important trigger is change: a retrained model, a new vendor, a new data source or a reorganisation of the notification chain makes the card describe a system that no longer exists. Calendar review catches drift slowly. Change-triggered review catches the specific failures that make a card dangerous, which are almost always a stale name and a stale authority.
Skill.re