AI Incident Response Planning
Priya Anand is the program director for digital services at a mid-sized state department of labor. On a Tuesday, her AI-powered chatbot, the one that answers unemployment-insurance questions for about 9,000 residents a day, started telling claimants they were eligible for a back payment they were not owed. By the time someone noticed, 1,400 people had been told to expect money that did not exist. The phones lit up. A reporter called. Priya's first instinct was to ask her vendor to just fix it. Her second, better instinct was to ask whether they had a plan for this. They did not. What followed was four days of improvisation: guessing who to notify, arguing about whether to take the bot offline, and discovering afterwards that nobody had preserved the logs. The technical failure took an afternoon to fix. The absence of a plan turned it into a two-week ordeal and a hearing before the legislature.
This lesson builds the plan Priya wished she had. AI incidents will happen in any agency running production AI at scale; that is a certainty rather than a probability. The question is not whether you will have one but whether your agency can respond fast, honestly and lawfully when it does. An agency that has prepared contains the damage, preserves its credibility and learns something. An agency that has not produces exactly the kind of crisis that becomes a GAO report, an inspector general finding, a congressional inquiry and, eventually, a class action.
Planning is not the same as responding
There is a difference between knowing what to do during an incident and having decided, in advance, who decides. Incident response is the action you take when something breaks. Incident response planning is the work you do on a calm afternoon so that the action is fast, lawful and defensible instead of improvised at eleven at night by whoever happens to be online. The distinction sounds pedantic until you watch a team spend its first two hours arguing about authority rather than about the system.
The controlling idea is the fire drill. A building does not survive a fire because people are brave. It survives because the exits are marked, the alarm is wired, somebody is assigned to count heads at the assembly point, and everyone has walked the route before there was smoke. An AI incident response plan is that drill: roles assigned, thresholds defined, authorities granted in advance, and the route walked at least once before the real emergency. Be clear about what this does and does not deliver. A written plan does not contain an incident. It removes a set of decisions from the moment when they would be made worst, which is a large benefit and a narrow one.
What counts as an incident
Define it in writing, or your staff will freeze on the threshold question while the clock runs. A workable definition covers any event where an AI system produces incorrect, biased, unsafe or unauthorised output that affects the public, exposes data, or could draw legal or media scrutiny. Priya's chatbot crossed that line the moment it told one person they were owed money they were not, which is worth saying explicitly, because teams routinely wait for a headcount before conceding that anything has happened.
The harder problem is that many AI incidents do not look like the incidents your existing process was built for. A model that begins producing biased outputs after a data drift is not a network intrusion, but it is an incident under OMB Memorandum M-24-10 if the use is rights-impacting. A model that processes data correctly and produces harmful recommendations because of flaws in its training data is not a technical failure in the FISMA sense, and it still affects citizens. A prompt injection attack that causes a chatbot to leak privileged information is simultaneously a security incident and a data breach, and it activates two playbooks at once. Your plan has to be broad enough to hold all three shapes, with a distinct escalation path and a named coordinating office for each.
The AI incident taxonomy
A defensible taxonomy has three top-level categories, and the value of writing them out is that each subcategory tells you what to instrument and who to call. The source describes its own list as a dozen subcategories while enumerating thirteen, so the grid below reproduces the subcategories it actually names and omits the count. Do not treat the boundaries as clean: an incident can occupy more than one cell, and a prompt injection that leaks personal information is both a security incident and a privacy breach requiring both responses to run simultaneously.
| Category | Subcategories |
|---|---|
| Security | Model theft, where attackers exfiltrate weights or architecture; adversarial examples, meaning crafted inputs that force wrong outputs; prompt injection, meaning crafted prompts that override instructions; model inversion, where attackers extract training data from outputs; data poisoning, where attackers corrupt training data; and supply chain compromise, where a model or dataset used downstream is tampered with |
| Performance | Accuracy drift below committed levels; fairness degradation, where disaggregated metrics show new disparities; calibration failure, where confidence estimates are no longer reliable; and data pipeline failure, where input data has become corrupt or stale |
| Safety | Harmful outputs, including defamatory, dangerous or unlawful content; physical harm, where AI-directed actions cause real-world injury; and rights violations, where due process, equal protection or statutory rights are infringed |
For each subcategory the plan should document six things: the detection signal, the initial triage criteria, the coordinating offices, the external notification requirements, the containment options and the typical recovery sequence. That is the work that turns a taxonomy from a diagram into a runbook. It is also where you discover which subcategories you currently have no way of detecting at all, which is the most useful finding a planning exercise produces.
Severity classification, and whose clock it is
Severity is separate from category, and it drives the tempo of everything else. A Severity 1 incident affects mission-critical operations or produces significant citizen harm. A Severity 2 affects important operations or produces moderate harm. A Severity 3 is localised or low-impact. Publish the classification framework so that staff and vendors know in advance what response an incident will get, because ambiguous severity produces slow response and slow response makes incidents worse.
The clocks attached to each tier are commitments the agency makes to itself. They are not statutory deadlines, they are not a legal test, and nothing about meeting them establishes that the agency has discharged a duty imposed by law. Say that in your plan, in those terms, so that nobody reads an internal target as a safe harbour. The examples below come from two different agency plans and are reproduced as examples of the shape such commitments take, not as a standard to adopt unexamined.
- Severity 1, critical. Active harm to people, a rights violation, a data breach of sensitive records, or 100 or more people affected. Suspension is presumed. In one plan, leadership and legal counsel are notified within 2 hours; in another, the Chief AI Officer and Chief Information Security Officer are paged within an hour and the Secretary is briefed within a day.
- Severity 2, major. Significant errors affecting a smaller group, no immediate physical harm, but real financial or legal exposure. Notify within one business day; decide on suspension within four hours. The suspension clock deliberately runs faster than the notification clock, because containment cannot wait on a notification chain.
- Severity 3, minor. Isolated errors, easily corrected, no public harm. Log it, fix it, and review it at the next cycle rather than paging anybody.
The chatbot incident was Severity 1 on headcount alone, since 1,400 people is well past the hundred-person line in that plan, and it would also have qualified on the nature of the harm. Notice how much easier that call is when the threshold exists on paper. Priya's team spent part of their first day debating whether a wrong answer about money counted as harm, which is a question that should have been answered months earlier by somebody who was not under pressure.
The legal and policy frame
Federal incident response does not start from a blank page, and the single most important discipline here is that you map the obligations that already apply rather than inventing timings that feel right. The source states that FISMA and OMB Memorandum M-17-25 require agencies to maintain an incident response capability and report to CISA; confirm the current governing instrument with your Chief Information Security Officer rather than quoting a memorandum number from a training deck. M-24-10 layers AI-specific obligations on top, including timely disclosure of rights-impacting incidents and coordination with the agency's Chief AI Officer.
CISA coordinates federal civilian cyber incident response under its authorities, which the source identifies as the Cybersecurity Information Sharing Act of 2015 and the Cyber Incident Reporting for Critical Infrastructure Act. Reportable incidents go to CISA within the statutorily defined windows, and AI incidents often follow the same timing because most carry a security dimension. Treat CISA as a partner rather than an overhead: the cross-agency intelligence they return is generally worth more than the notification costs. For rights-impacting incidents the Department of Justice Civil Rights Division may engage, and affected beneficiaries hold due process and equal protection rights that shape what the agency must do next. Sector rules stack on top of all of it, including HIPAA breach notification where health information is involved and the disclosure rules governing federal tax information.
Everything in the previous paragraph is a pointer, not a deadline. The plan's job is to write down which of these apply to this system and this data, with the specific clock beside each, verified by counsel before the incident rather than read for the first time during one. Many states and federal programs also carry breach notification statutes with hard deadlines measured in days. Map them into the plan by name. An agency that has to determine its notification duty while the phones are ringing will get it wrong in the direction that is hardest to defend.
The response lifecycle
NIST Special Publication 800-61 supplies the lifecycle that federal AI incident response adapts: preparation; detection and analysis; containment, eradication and recovery; and post-incident activity. The source labels this lifecycle with a phase count that does not match the phases it then lists, so the count is omitted here and the phase names are what matter. Preparation means standing up the AI incident response team, which the source staffs with the Chief AI Officer, the Chief Information Security Officer, the Senior Agency Official for Privacy, the General Counsel, the program owner and communications. It also means writing runbooks for each subcategory, exercising them, and maintaining the evidence infrastructure that an incident will demand: lineage, logs, a model registry and audit trails.
Detection and analysis brings the monitoring stack to bear on discovery. Alerts from drift detection, fairness monitoring, security monitoring and user reports all feed a single on-call queue that the response team reviews, because signals scattered across four dashboards owned by three teams are signals nobody correlates. Triage determines whether an alert is a real incident, which subcategory it falls into and its initial severity. Analysis then expands the picture: which systems, users and data are affected; what timeline the incident followed; which controls failed and which held. Analysis produces the facts every later phase relies on, which is why doing it in parallel with containment rather than after it is worth the staffing.
Containment, eradication and recovery is the operational heart. Containment stops the bleeding, typically through model rollback to a known-good version, traffic diversion to a fallback system, disablement of specific features, or in extreme cases full system takedown. Eradication removes the root cause: patching the vulnerability, retraining on clean data, retiring the compromised artifact. Recovery restores service, verifies that the restoration worked, and monitors closely for recurrence. For rights-impacting incidents, recovery also includes remediation for affected beneficiaries, which may mean re-adjudicating their cases, and that obligation should appear in the plan rather than being negotiated afterwards.
Where this sits in the frameworks you already owe
Priya does not have to invent governance from scratch. The NIST AI Risk Management Framework organises AI risk into the Govern, Map, Measure and Manage functions, and incident response lives in Manage: the framework expects agencies to plan for, respond to and recover from incidents, and to feed what they learn back into the system. The framework is voluntary guidance rather than binding law, which is exactly why naming its functions in your plan is useful; it converts a general expectation into something a reviewer can check. M-24-10 separately requires agencies to monitor rights-impacting and safety-impacting AI in production and to be able to stop using a system that fails its protections. An incident plan is how an agency delivers on both. You are not adding bureaucracy; you are operationalising a duty you already carry.
Preserve before you patch
Two decisions are easy to get wrong under pressure, and both are cheap to decide in advance. Preserve the logs, the inputs and outputs, the model version, the configuration and the internal communications, before anyone fixes the system and overwrites the evidence. Priya's team lost the original logs because the vendor redeployed a patch over them, which turned a knowable question about what the model had actually said into a matter of reconstruction and argument. Write the preservation step into the runbook ahead of the remediation step, and write it into the vendor's obligations too, since in Priya's case the person who destroyed the evidence was not an employee.
Suspension is the other decision, and it is a genuine balancing act rather than an obvious call. Taking the bot offline stopped the false promises and also cut off 9,000 legitimate users a day, most of whom needed an answer about their benefits that afternoon. Decide the default in advance: for Severity 1, suspension is presumed unless an executive accepts the risk in writing. That formulation does two things. It puts the burden of argument on continuing rather than on stopping, and it produces a signed record of who chose to keep the system running and on what reasoning, which is the document an inspector general will ask for first.
Communicating with affected people and the public
Public communication in an incident is a leadership discipline, not a public relations exercise. The principle is that affected citizens deserve to know what happened, what the agency is doing about it, and how their own situation will be addressed. For rights-impacting incidents that typically means individual notification to affected people, a public statement, a path for beneficiaries to request re-adjudication, and updates as remediation progresses. The FTC's breach notification model and the HIPAA breach rules supply usable templates for the structure of such a notice, which saves you drafting one under deadline pressure.
On timing, the source is explicit: notification within 72 hours of confirmation is emerging as a norm for significant incidents, and waiting longer typically worsens the outcome. Carry that as written, and understand what it is. It is a fast-moving expectation about practice, not a statutory clock, and it does not displace any specific statutory deadline that applies to your data, which may well be shorter. Where the two differ, the statute governs and the norm sets the floor for everything the statute does not reach. Agencies that communicate clearly and quickly build trust; agencies that hide, minimise or delay lose it and frequently face worse legal exposure for the delay than for the incident.
The instinct to minimise is strong and it is worth naming. Mature agencies treat incidents as learning opportunities rather than reputational threats to be managed down. Research from the Partnership for Public Service on high-performing federal programs consistently finds that candour about incidents produces greater trust than denial, and that accelerated remediation produces stronger long-run outcomes than bureaucratic delay. The leadership practice reduces to four moves: when something goes wrong, say so, fix it fast, learn from it, and tell the public what you learned.
External coordination
AI incidents rarely stay inside one agency, and the relationships that make coordination work cannot be built during the incident. For rights-impacting incidents, loop in the Department of Justice Civil Rights Division early. If your system has produced disparate impact across protected classes, they are a natural interlocutor, and engaging them before a class action is filed produces a materially better outcome than engaging them after. Inform your Inspector General early for significant incidents as well, under the CIGIE protocols and the Inspector General Act as amended. Early engagement makes an adversarial audit less likely and a collaborative remediation more likely; it does not guarantee either, and it should be done because it is right rather than because it buys protection.
Name the individuals now, not the offices. A plan that says "notify the Inspector General" is a plan that begins with somebody searching a directory. A plan that names a person, a deputy, a phone number and an after-hours route is a plan that begins with a phone call. Priya's tabletop exercise found that nobody knew the privacy officer's after-hours number, which is exactly the kind of gap that costs two hours at the worst possible moment and costs nothing to fix in advance.
Post-incident learning
After every significant incident the agency runs a blameless post-incident review: what happened, why, and what will change. Blameless is deliberate. If staff fear punishment they hide the next incident, and hidden incidents are the ones that become scandals. A rigorous root cause analysis names contributing factors without scapegoating, and an agency that punishes the people who surface root causes produces a culture that buries them next time.
Five artifacts anchor the phase. The timeline is a reconstruction from first signal through full recovery, including every detection, decision and communication event. The root cause analysis identifies contributors at multiple levels: the immediate technical cause, the governance gap that allowed it, and the institutional condition that allowed the governance gap. The impact assessment quantifies the harm, covering how many people were affected, how severely, for how long and how reversibly. The remediation plan names specific actions, owners and deadlines, with the AI governance board as the accountability body. The public report, for significant incidents, presents all of it in a form a citizen can read.
Each artifact then feeds something downstream, which is the argument for producing them even when the incident feels closed. The timeline supports any subsequent inspector general or GAO audit. The root cause analysis drives changes in the system and often in the governance program. The impact assessment supports remediation budgeting and litigation exposure analysis. The remediation plan gets tracked to closure with the same rigour as any other program commitment. The public report builds trust when written candidly and damages it when written defensively, which means the drafting posture matters as much as the content.
The incident register
Continuous improvement extends learning across incidents rather than within them. Maintain an incident register so that patterns become visible over time: a single incident is a data point, three similar incidents is a signal, and ten is a structural issue. Chief AI Officers who review the register quarterly find structural problems their teams are too close to see. Trends in the register feed the annual Chief AI Officer report, the M-24-10 inventory update and the agency's AI governance roadmap. Agencies that skip the post-incident phase relearn the same lessons repeatedly, and usually learn them publicly the third time.
A usable artifact: the one-page plan
This template fits on a single page and lives where staff can reach it in 30 seconds, not buried in a shared drive behind a login nobody remembers. Fill it in per system, because the answers differ by system even inside one program.
| Field | Entry for this AI system |
|---|---|
| System name and owner | Named individual plus backup, with phone numbers |
| What counts as an incident | Plain-language trigger definition |
| Severity 1, 2 and 3 criteria | Headcount, harm type and data exposure thresholds, marked as agency-set |
| Notification list by tier | Role, name, contact and the internal deadline in hours |
| Legal and breach deadlines | The statutory clocks that apply to this data, verified by counsel |
| Suspension default | Who can take it offline, and the default by severity |
| Evidence to preserve | Logs, inputs and outputs, model version, configuration, communications, and where they are held |
| Vendor obligations | Preservation duty, notification duty and support commitment during an incident |
| Public communication owner | Public affairs lead plus the location of the holding statement |
| Post-incident review owner | Who convenes it, and within how many days |
| Last drill date | Walk it at least annually |
Exercises and standing readiness
No incident response plan is real until it has been exercised. The source recommends at least one tabletop exercise and one live exercise per year, with scenarios drawn from the taxonomy. A tabletop walks a team through a realistic scenario without touching production. A live exercise simulates detection and response in a controlled environment, sometimes with a red team playing the adversary. Exercises surface gaps in runbooks, escalation paths, tools and relationships that a paper review cannot find, and the Department of Homeland Security's Cyber Storm exercises offer a federal template that a growing number of agencies now extend with AI-specific scenarios.
Six weeks after the chatbot failure, Priya ran her first tabletop. She invented a bias complaint against an AI resume-screening tool and walked her team through the template in real time. They found the gaps in 40 minutes: nobody knew the privacy officer's after-hours number, and the holding statement template did not exist. Both were fixed that week. The next real incident, a minor data-formatting bug, was closed in three hours with the logs intact. That is evidence the plan helped, on one incident of one type. It is not proof that the plan works, and the honest conclusion Priya drew was that she needed more scenarios rather than fewer.
Readiness also needs standing infrastructure rather than a document. An on-call rotation staffed around the clock for Severity 1. Runbooks kept current rather than kept. A war room with pre-provisioned communication channels. Named relationships with CISA, DOJ, the Inspector General and communications. Executive briefing templates ready to customise. Agencies that build this during the quiet periods can activate it quickly; agencies that build it during an incident lose the hours they can least afford. Building it in advance shortens the response; it does not make the response good, which is what the exercises are for.
The cultural foundation
Agencies with strong AI incident response exhibit three traits, and no amount of documentation substitutes for them. Dissent is safe, so early signals reach somebody with authority instead of dying in a team channel. Learning is prioritised over blame, so root causes get named accurately. And transparency is the default, so citizens and oversight bodies can tell the agency is being straight with them. Agencies that punish early warning, scapegoat individuals or conceal incidents produce weak response regardless of how polished the runbooks are.
The Chief AI Officer and Chief Information Security Officer carry particular responsibility here, because culture is established at the top and felt everywhere else. How those two behave in their own first difficult incident sets what everyone below them believes is safe. That makes serious incident response as much a leadership practice as a technical one, and it is the part of the plan you cannot delegate to a template.
Anti-Patterns to Avoid
- Treating the written plan as containment. A documented plan, a completed runbook and a finished tabletop do not stop an incident. They remove specific decisions from the moment when they would be made worst. Agencies that describe a plan as assurance of containment stop investing in the detection that actually finds the next incident.
- Reading one clean response as proof the plan works. Priya's three-hour close on a formatting bug is evidence of improvement on one incident of one type. The security and safety branches of her taxonomy remained entirely unexercised.
- Patching before preserving. The fix that overwrites the logs destroys the only record of what the system actually did, and it usually comes from a well-meaning engineer or a vendor acting fast. Put preservation ahead of remediation in the runbook and in the contract.
- Confusing an internal clock with a legal duty. A two-hour internal notification target is a management commitment. Meeting it says nothing about whether a statutory notification deadline has been met, and staff who conflate the two will miss the one that matters.
- Reading a statutory deadline for the first time during the incident. Notification duties differ by data type, program and jurisdiction. Map them into the plan by name with counsel's confirmation, because the version improvised under pressure is the one that gets litigated.
- Defaulting to keep the system running. If continuing requires no signature and suspending requires an argument, the system will keep running through incidents it should not. Reverse the default for critical severity and require the risk acceptance in writing.
- Notifying offices rather than people. A plan naming the Inspector General starts with a directory search. A plan naming a person, a deputy and an after-hours number starts with a phone call.
- Running a review that assigns blame. One punitive post-incident review teaches an organisation to route the next signal away from the people who could act on it, and the incidents do not stop, they only stop being reported.
- Treating candour as a legal risk to be minimised. Delay and minimisation typically produce worse legal exposure than the underlying incident, and they cost the trust that makes the next disclosure survivable.
- Exercising only the scenario you already understand. A tabletop that reruns last year's incident confirms you fixed last year's incident. Draw scenarios from the categories you have never tested, particularly the ones you currently could not detect.
Practice Prompts
Use these to build and stress-test the plan itself. Do not use a general-purpose tool to determine a legal notification deadline; that answer comes from counsel and from the statute, and a generated deadline that sounds right is exactly the failure mode this lesson exists to prevent.
- "Here is a description of one AI system we operate. For each incident subcategory in this taxonomy, tell me what detection signal would reveal it, and mark the ones for which we have described no detection capability at all."
- "Draft a tabletop scenario for a fairness degradation incident in this system, including the first alert, three complicating facts that emerge over the first day, and two moments where an authority question has to be answered."
- "Review this incident response plan and list every step that names an office rather than a person, and every step that assumes information will be available without saying who preserves it."
- "Given this incident timeline, produce a root cause analysis at three levels: the immediate technical cause, the governance gap that allowed it, and the institutional condition that allowed the governance gap."
- "Rewrite this internal incident notice as a public statement a citizen affected by the incident could read, stating what happened, what we are doing, and how they can get their own case addressed. Flag anything I have asserted that we have not established."
Reflection Questions
- If your highest-risk AI system produced a harmful output right now, who would learn about it first, and how long before somebody with authority to suspend it knew?
- Which of the three incident categories could your current monitoring not detect at all? What would it cost to change that?
- Does anybody in your organisation have written, pre-granted authority to take a production AI system offline? If the answer is unclear, that is the answer.
- Which statutory notification clocks apply to the data your AI systems touch? Who verified that, and when?
- Think about the last time something went wrong in your organisation. Was the person who raised it treated as helpful? What would the honest answer predict about the next incident?
Glossary
- Incident. Any event where an AI system produces incorrect, biased, unsafe or unauthorised output that affects the public, exposes data, or could draw legal or media scrutiny.
- Severity classification. The agency-set tiering, separate from incident category, that determines response tempo and executive engagement.
- Prompt injection. Crafted input that overrides a system's instructions, which can be simultaneously a security incident and a data breach.
- Model inversion. An attack that extracts training data from a model's outputs.
- Data poisoning. Corruption of training data by an attacker, so that the resulting model behaves as the attacker intends.
- Calibration failure. A state in which stated confidence no longer tracks observed accuracy, so confidence-based review gates stop working.
- Containment. Stopping ongoing harm, typically through model rollback, traffic diversion, feature disablement or full takedown.
- Eradication. Removing the root cause, by patching, retraining on clean data or retiring the compromised artifact.
- Re-adjudication. Reconsidering the decisions made about affected beneficiaries during a rights-impacting incident, as part of recovery.
- Blameless review. A post-incident review that names contributing factors without scapegoating individuals, so that the next incident is reported rather than hidden.
- Incident register. The standing record of incidents over time, reviewed periodically so that patterns invisible in any single incident become visible.
Related Lessons
- AI Incident Response: What to Do covers the execution side once an incident is under way.
- AI Incident Documentation and Response develops the five post-incident artifacts in detail.
- Enterprise AI Risk Management places the incident register inside the wider risk apparatus.
- Cybersecurity for AI Systems covers the security branch of the taxonomy at depth.
- Privacy Engineering for AI addresses the breach dimension when an incident exposes personal data.
- Continuous Monitoring Fundamentals supplies the detection signals this plan depends on.
- When Government AI Goes Wrong examines incidents that were handled badly and what the handling cost.
Closing Thoughts
What made Priya's incident a two-week ordeal was not the model. The model was fixed in an afternoon. It was that six separate questions with obvious answers had to be settled from scratch while people were watching: who decides, whether to suspend, what to preserve, who to call, what to tell the public, and who runs the review. Each of those has a defensible answer that takes fifteen minutes to write on a quiet day and hours to argue about on a loud one.
Write the plan for one system this month, name real people in it, and then exercise it against a scenario you have not seen before. Expect the exercise to embarrass the plan; that is what it is for, and a tabletop that finds nothing has usually tested nothing. Above all, keep the two categories of clock separate in your own head and in the document: the commitments your agency makes to itself, which you set and can improve, and the deadlines the law imposes, which you verify with counsel and never estimate.
Key Takeaways
- Plan on the calm day. The value of a plan is that the hardest decisions, who decides, what to suspend, what to preserve, were made before anyone was panicking. A plan is not containment.
- Define the trigger in writing. Without a written definition of an incident, staff freeze on the threshold question and lose the hours that matter most.
- Use a taxonomy with three branches. Security, performance and safety incidents look different, are detected differently, and activate different coordinating offices, and one event can occupy more than one branch.
- Severity drives tempo, and severity clocks are self-set. Internal notification targets are management commitments, not legal tests; statutory deadlines are separate, verified with counsel, and never estimated.
- Preserve before you patch. Logs, inputs, outputs, model version, configuration and communications are evidence, and a well-meaning fix by staff or a vendor can destroy them.
- Decide suspension defaults in advance. For critical incidents presume the system comes offline unless an executive accepts the risk in writing, which also produces the record of who decided.
- Communicate early and specifically. Individual notification, a public statement, a re-adjudication path and progress updates; the source treats notification within 72 hours of confirmation as an emerging norm for significant incidents, distinct from any statutory clock.
- Run blameless reviews and keep the register. Five artifacts per incident, and a register reviewed periodically where one incident is a data point, three is a signal and ten is a structural issue.
- Exercise the plan against scenarios you have not seen. At least one tabletop and one live exercise a year, drawn from the categories you currently cannot detect.
- Culture decides the outcome. Safe dissent, learning over blame and default transparency determine whether the runbooks ever get used as written.
Frequently Asked Questions
Is there a legal deadline for notifying the people affected by an AI incident?
That depends entirely on the data, the program and the jurisdiction, and it is a question for your counsel rather than for this lesson or for an AI assistant. Many federal programs and many states carry breach notification statutes with hard deadlines measured in days, and sector rules such as HIPAA breach notification stack on top where health information is involved. The source describes notification within 72 hours of confirmation as an emerging norm for significant incidents, which is a practice expectation and not a substitute for whatever statutory clock applies to you. Map the applicable clocks into the plan by name, in advance.
Our incident is a fairness problem, not a security breach. Does the plan still apply?
Yes, and this is the case most agencies are least prepared for. A model producing biased outputs after a data drift is not a network intrusion, and under M-24-10 it is still an incident if the use is rights-impacting. The performance branch of the taxonomy exists precisely for this, and it needs its own detection signals, its own coordinating offices including the civil rights function, and its own containment options, because none of them are the ones your security operations centre uses.
Who has the authority to take a production AI system offline?
Whoever your plan says, granted in writing before the incident. The design principle that matters more than the specific choice is where the default sits: if continuing requires no signature and suspending requires winning an argument, the system will keep running through incidents it should not. Setting presumed suspension for critical severity, overridable only by a written executive risk acceptance, reverses that and produces a record of the decision at the same time.
What exactly should be preserved, and who is responsible?
The logs, the inputs and outputs, the model version, the configuration and the internal communications, held somewhere the remediation cannot overwrite. Name the responsible person in the plan and put the same duty in the vendor's contract, because in Priya's case the evidence was destroyed by a vendor deploying a patch, not by anyone on her staff. Preservation belongs ahead of remediation in the runbook sequence for exactly this reason.
How often should we exercise the plan?
The source recommends at least one tabletop exercise and one live exercise per year, with scenarios drawn from the incident taxonomy. Frequency matters less than coverage: three tabletops of the same familiar scenario test less than one exercise aimed at a category you currently have no way of detecting. Priya's first tabletop found two gaps in 40 minutes, which is the normal result, and the correct response to that is to schedule the next one rather than to conclude you are ready.
Should we tell the Inspector General before we have all the facts?
For significant incidents, early engagement is the source's recommendation, under the CIGIE protocols and the Inspector General Act as amended, and the same logic applies to the Department of Justice Civil Rights Division for rights-impacting incidents. Early disclosure makes an adversarial audit less likely and a collaborative remediation more likely, though it guarantees neither. You can report what is known and what is still being established; a factual interim report is more useful to an oversight body than a complete one that arrives after the decisions have been made without them.
Skill.re