←
AI for Government
Capable · M7 · lesson 7 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Incident Documentation and Response
📖
now learning

AI Incident Documentation and Response

15 min

On a Wednesday morning in March, Yolanda Ferreira pulled up her dashboard and saw the number: 340. That was how many Supplemental Nutrition Assistance Program (SNAP) renewal denials her county's automated eligibility AI had issued overnight, in a single batch run. Her agency had deployed the tool four months earlier to triage renewals and flag likely-ineligible cases for expedited denial. The system had been running quietly and nobody had been watching the output closely. By 9:15 a.m., Yolanda had seven voicemails from caseworkers reporting confused clients whose food benefits had stopped. One was a diabetic single mother of three. Another was a 74-year-old man on a fixed income. A software configuration error had caused the AI to misread income verification fields, and the denials were wrong. Yolanda had never managed an AI incident before, and the decisions she made in the next thirty minutes would shape every investigation, lawsuit and public records request that followed.

Why Government AI Incidents Are Different

AI systems fail. They make serious errors, they discriminate, and they do both at scale and quietly. Private companies experience this too, but when a company's recommendation engine misfires the stakes are usually revenue or reputation. When a government AI system misfires, the stakes can include civil rights violations, wrongful deprivation of benefits, harm to vulnerable populations, failures in critical infrastructure, and legal liability that follows the agency for years. Without procedures, the response to all of that is improvised by whoever happens to notice first. With procedures, the response is systematic and the agency actually learns something.

Three features make government AI incidents uniquely high-stakes. The first is FOIA exposure. Every email Yolanda sent that morning and every internal report she filed is a potential public record. Courts and oversight bodies subpoena AI system logs, training data documentation and internal incident communications. If the agency's response looks chaotic or indifferent, those records will eventually show it, and they will be read by people who were not there and who have no obligation to be generous about what they find.

The second is civil rights liability. Government benefits decisions must comply with the Fourteenth Amendment's equal protection clause and with federal statutes including the Civil Rights Act of 1964, the ADA and Title VI. An AI system that produces disparate outcomes for protected classes, even unintentionally, creates legal exposure for both the agency and individual officials. The 340 wrongful denials were not merely administrative errors. They were potential due process violations, and the difference between those two descriptions is the difference between a ticket and a lawsuit.

The third is Inspector General oversight. Most federal and many state agencies have an IG office with independent authority to investigate AI system failures. If the IG investigates before the agency has documented its own response, the agency loses control of the narrative. Proactive, well-documented incident response is the single best protection against an IG finding that turns a technical error into a governance scandal. In government, how you respond to an AI failure matters almost as much as preventing it. A well-documented response protects the public; a poorly documented one becomes the story.

The Controlling Analogy: The Medical Chart

Think of AI incident documentation the way hospital staff think about a patient's chart. The chart exists because decisions made under time pressure, by multiple people, need to be reconstructable by anyone who picks up the case later: a specialist, a lawyer, an auditor. Every entry is timestamped. Every decision is recorded with its rationale. The chart is the defense against "we do not know what happened." Classification is triage. Documentation is charting. Notification chains are consult orders. Root cause analysis is the differential diagnosis. Post-incident review is the morbidity and mortality conference. Each step exists because the next person who needs to understand what happened, including a federal auditor two years from now, will not have been there when it unfolded.

Severity Classification: Triage Before You Act

Before Yolanda can do anything else, she needs to classify the incident. Classification drives everything downstream: who gets notified, how fast, what gets suspended, and who leads the response. Classify in writing within thirty minutes of identifying the incident. If you are unsure between two levels, classify higher. You can always downgrade later with an explanation. Upgrading after the fact looks like you were minimizing the problem, and that is exactly what an investigator will conclude.

Severity 1, critical

The AI caused or is actively causing harm to individuals, is producing legally prohibited outcomes, or has affected 100 or more people. This tier covers potential loss of life or serious injury, widespread discrimination affecting many people, a system making systematically wrong decisions, and major civil rights violations. Immediate shutdown is presumed, followed by a full investigation. Examples: wrongful denial of emergency benefits to 50 or more households; an AI-generated output used to make an unconstitutional arrest; a child welfare AI that failed to flag an at-risk case and a child was harmed. Yolanda's incident, 340 wrongful SNAP denials affecting potentially more than a thousand household members, is Severity 1.

Severity 2, high

Significant harm to a meaningful population, significant disparate impact on a protected class, a failure affecting critical decisions, or clear potential legal liability. In headcount terms this tier covers incorrect outputs affecting roughly 10 to 99 individuals, and it also covers smaller numbers when a protected class, a sensitive benefit category or a constitutionally protected right is involved. The response is immediate containment followed by urgent investigation: suspension evaluated within four hours and leadership notified within four hours.

Severity 3, medium

Localized impact: isolated errors affecting a small number of people, roughly one to nine individuals, or outputs caught by human review before any action was taken. Fairness concerns exist but are limited in scope and look manageable through process improvement. The response is prioritized investigation with enhanced monitoring. Notify the relevant teams within 24 hours, and require supervisor notification with logging and review inside 48 hours.

Severity 4, low

An isolated error in an edge case, where the system works correctly overall and there is no pattern of problems. Unexpected output was caught before any downstream action and no individual was affected. The response is investigation and documentation rather than escalation: notify through normal channels within a few days, log the event, and schedule a technical review within two weeks. The point of logging Severity 4 events at all is that patterns only become visible once you have a few of them written down.

What to Document and When

Start a log the moment you suspect something is wrong. A timestamped note reading "9:12 a.m., caseworker reports unusual volume of client calls about SNAP denials, investigating" is more valuable than a comprehensive report written at 2 p.m. Timestamps prove when you knew what you knew, which is the question every subsequent reviewer will ask first. A messy accurate log beats a clean reconstructed one, and reconstructed logs have a way of being recognizable as such.

Your running log should capture when and how the incident was first noticed and by whom; how many individuals appear affected, marked as preliminary; whether the AI is still running and, if suspended, when and by whom; what types of data were processed; every decision with its owner and its time, written as "9:47 a.m., Yolanda Ferreira suspended overnight batch run" rather than "around 10 a.m. the system was paused"; and any external communications and who made them. That log is the raw material for the formal record, which has five parts.

Record sectionWhat it must answer
Incident reportDate and time discovered, who discovered it, what was wrong, severity classification, initial assessment of scope
Investigation findingsRoot cause, how many cases affected, which populations were affected, how long the problem had been occurring, and why it was not caught earlier
Immediate responseActions taken such as shutdown, isolation and notification; communications to affected parties; the interim process (manual review, extended timelines) while the system is down
ResolutionFix implemented, testing and validation of the fix, broader system improvements made, changes to monitoring and alerts
PreventionWhat will prevent recurrence, what changes to system, process or monitoring are required, and what procedures or alert rules were updated

Two of those fields carry more weight than the rest. "How long was the problem occurring?" is the question that converts an incident from an event into a population, because a configuration error that ran for one night affects one batch and the same error running for three weeks affects everyone processed in three weeks. "Why was it not caught earlier?" is the question that names your monitoring gap, and it is the field agencies are most tempted to leave vague. Answer both in writing even when the answers are unflattering, because an incomplete record is read as a concealed one.

One field in that table deserves more attention than it usually gets: the interim process. Suspending an AI system does not suspend the work it was doing. Yolanda's tool was triaging renewal applications, and the moment it stopped, those applications began accumulating against statutory processing timelines that the outage does not pause. The record has to say what replaced the system, who staffed it, and what happened to the timelines people were relying on, including any extensions granted and how applicants were told about them. Agencies that plan the manual fallback in advance can suspend a system without hesitating. Agencies that have not planned it do hesitate, and the hesitation is what lets an incident keep running while the fallback is invented.

The communications entry matters for the same reason. Recording that affected parties were contacted is not sufficient; the record should show who was contacted, when, through what channel, and what they were told, because the remedy available to a wrongly denied applicant usually depends on them acting within a window. If a notice went out by mail to an address the system already had reason to doubt, that is a fact the record should contain rather than one an investigator should have to discover.

Notification Chains: Who Learns What, and When

Notification means the right people have enough information to make their own decisions. It is not an invitation to panic, and it is not a request for permission. It is a professional obligation. The severity classification sets the pace: a Severity 1 incident goes to leadership, legal and public affairs immediately; Severity 2 reaches key stakeholders within hours; Severity 3 goes to the relevant teams within 24 hours; Severity 4 travels through normal channels within a few days. The specific clocks below are the outer bounds one county agency worked to for a Severity 1 event, not a universal standard, and your agency's own policy may be tighter. Confirm the clocks that bind you before you need them.

Internal leadership, within two hours. Give the severity classification, the estimated number affected, whether the system is suspended, what is known about the cause, and what immediate steps have been taken. One page or less. Leadership does not need the diagnosis yet, only the scope and the fact that someone competent has hold of it.

Legal counsel, within two hours, concurrent. Legal needs to evaluate litigation risk and advise on communications. Their role is not to stop the response; it is to ensure the response does not inadvertently create additional liability. Send this notification in parallel with the leadership notification rather than waiting for leadership to approve it, because the sequential version costs hours you do not have.

Inspector General, within 24 hours. Proactive IG notification is counterintuitive but almost always the right move. An IG who hears from the agency before hearing from a whistleblower is far more likely to treat the agency as a cooperative subject. For federally funded programs, IG notification may also be a regulatory requirement, so check what your program actually obliges you to report and when.

Public affairs, within four hours. Affected clients will call advocacy organizations, who will call reporters, who will call the agency. Giving public affairs a head start is the difference between a prepared statement and "we cannot comment at this time." If the issue is systemic rather than isolated, assume a public communication will be needed and start drafting it before you are asked for it.

Affected individuals, within 24 hours. People whose benefits were wrongfully denied have a right to know, a right to appeal, and, depending on state law, a right to emergency reinstatement. Use plain language: what happened, what the agency is doing about it, and how to get help immediately.

What you say matters as much as when you say it, and the content is the same across audiences even though the framing differs. State what happened as facts rather than speculation. State the impact: how many people, and what the harm is. State the immediate response, what you are doing right now. State the next steps, including the investigation and a realistic fix timeline. And tell people how to report related problems they discover, because the affected population is often larger than the first count and the people who find the rest of it are usually the ones on the phone.

Root Cause Analysis

Root cause analysis is not about assigning blame. It is about understanding the chain of failures well enough to break it. Four questions frame the investigation. What went wrong technically, whether a model failure, a data issue or a logic error? Why was it not caught, whether through a testing failure, a monitoring gap or a process failure? What conditions allowed it, such as bad data, an edge case or an unusual scenario? And is this specific to one system or systemic, meaning the same weakness probably exists in others you have not checked yet? That last question is the one that turns a single incident into an agency-wide improvement.

The 5-Why technique is the standard tool for getting past the first plausible answer. You ask why, accept the answer, and then ask why about the answer, five times over. A generic chain runs: the system failed because of a data quality issue; the data quality issue existed because there was no validation; there was no validation because everyone assumed the data was clean; that assumption held because the requirement never specified validation; and the requirement never specified it because nobody understood how important it was. Applied to Yolanda's incident, the chain is concrete.

  1. Why were 340 SNAP applications wrongfully denied? The AI classified them as income-ineligible.
  2. Why? The income verification field contained incorrect values.
  3. Why? A data pipeline update three weeks earlier changed the field format, and the AI model was not updated to match.
  4. Why was the model not updated? No documented change management process required AI model review when upstream data sources changed.
  5. Why not? The agency's change management policy was written before AI systems were deployed and does not include AI model dependencies as a required review category.

The root cause is not a software bug. It is a policy gap, and a software patch alone would not prevent the next incident. Document the 5-Why analysis in writing and include it in the incident record, because it is the evidence that the agency understood what went wrong rather than only what broke. Note also what the third answer implies about scope: the pipeline change was three weeks old, so the investigation has to establish whether the overnight batch was the first affected run or merely the first one anybody noticed.

Post-Incident Review and Continuous Improvement

The post-incident review happens after the immediate crisis is resolved, usually within two to four weeks. It is not a blame session. It is a structured examination of what the incident revealed about agency processes, technology governance and oversight gaps. A review of a Severity 1 AI incident produces four outputs.

  • A written incident report covering detection, response timeline, root cause analysis, affected population and remediation steps. This is a likely FOIA target. Write it as if it will be published: clearly, factually, without spin.
  • A corrective action plan with specific owners, deadlines and metrics. "We will improve our change management process" is not a corrective action. "By April 15, the IT Director will revise Policy CMP-07 to include AI model dependency review in all upstream data source changes" is.
  • A monitoring update identifying what checks should have caught this earlier and how they will be implemented.
  • A lessons-learned summary suitable for sharing with peer agencies, because the next agency to hit this failure mode should not have to discover it the way you did.

Track completion at 30, 60 and 90 days. Assign a named individual, not a committee, to each item. Committees do not get corrective actions done; people do. And do not stop at the immediate fix, because the value of an incident is almost entirely in what you change afterwards. The improvements fall into three categories, and a serious review produces something in each.

System improvements make the technology harder to break: better monitoring to catch similar issues earlier, stronger validation and testing, architecture that fails more gracefully, and better human oversight at the points where an error becomes an action. Process improvements change how people work: updated procedures, new training, a revised incident response process, and different escalation triggers where the existing ones proved too slow. Monitoring improvements change what you watch: new metrics, different alert thresholds, more frequent checks, and new dashboard indicators. Yolanda's incident produced all three, and the one that mattered most was the least technical. Nobody had been watching the output of an automated system that was issuing benefit denials overnight.

Anti-Patterns

  • Hiding the incident. The instinct to contain a problem quietly is strong and it is always wrong in government, where the record eventually surfaces through FOIA, an IG referral or litigation. Full transparency is required, and it is also the cheaper path. An agency that self-reports is treated differently from one that is discovered.
  • Skipping the investigation. Restoring service and moving on leaves you with a system whose failure mode you never characterized. Root cause analysis is what distinguishes a fixed problem from a recurring one.
  • Fixing the symptom. Patching the code that misread the income field and calling it done ignores the change management gap that let the mismatch reach production. If your corrective action would not have prevented the incident had it been in place beforehand, it is not addressing the root cause.
  • No systemic learning. Treating the incident as a property of one system wastes it. Ask whether the same weakness exists elsewhere in your AI footprint, and check rather than assume.
  • Silence toward stakeholders. Leadership, legal, oversight bodies and above all the affected individuals need to hear from you while you still control the sequence. People who learn from a reporter that their benefits were wrongly cut off do not become reachable again easily.
  • Classifying down to avoid the paperwork. Choosing Severity 2 because Severity 1 triggers an immediate shutdown and a set of notifications is the version of this failure that shows up in audits. The rule is to classify up when uncertain, and to record the reasoning either way.
  • Reconstructing the log afterwards. A tidy narrative written at the end of the day, in one voice, with round times, is visibly not a contemporaneous record. Keep the rough timestamped notes and hand those over as well.
  • Corrective actions owned by a committee. An action item assigned to "the governance board" with a deadline of "next quarter" is an item nobody will complete. Name the person and the date, then check at 30, 60 and 90 days.

Practice Prompts

  • Draft the incident response procedure. Write the procedure for one AI system your agency operates. Who declares an incident, who can suspend the system, who classifies it, and what happens in the first hour?
  • Build the classification scheme. Adapt the four severity levels to your program's actual harms. What counts as Severity 1 in a benefits agency differs from a permitting office, and the thresholds should reflect the population you serve.
  • Write the notification template. Produce a one-page Severity 1 notification covering classification, estimated scope, suspension status, known cause and immediate steps, and a plain-language version for affected individuals.
  • Run a 5-Why on a real incident. Take an incident your organization has already had, AI-related or not, and run the chain until you reach a policy or process answer rather than a technical one.
  • Design the interim process. If the system in your first prompt were suspended tomorrow morning, what is the manual fallback, who staffs it, and what happens to the timelines people are relying on?
  • Audit your own detection. For each AI system in your area, write down how you would find out it had failed overnight, and who would see that signal. Where the honest answer is "a client would call," you have found your monitoring gap.

Reflection

Yolanda's system had been running quietly for four months and nobody was watching its output closely. That is the detail worth sitting with, because it is the most common condition in government AI right now and it is rarely anybody's fault in particular. Deployment ends with a launch, monitoring belongs to nobody specific, and the system does useful work for long enough that it stops being interesting. Ask yourself which of your agency's automated decisions currently has no named human reading its output, and how many days of wrong results would accumulate before anyone noticed.

Then consider the thirty-minute question. If you took that call, would you know who to notify and whether you have the authority to suspend a running system? Incident response fails most often not because people decided badly under pressure, but because nobody had decided in advance who was allowed to decide.

Glossary

  • Severity classification. The written determination of how serious an incident is, made early and on the basis of harm and scope. It sets the notification clocks, the suspension decision and the level of investigation.
  • Containment. Stopping the incident from affecting more people, usually by suspending or isolating the system, before the cause is understood. Containment precedes diagnosis.
  • Incident log. The contemporaneous timestamped record of what was noticed, decided and done, kept from the first suspicion onward. Distinct from the formal incident report written later.
  • Root cause analysis. Structured investigation into the chain of technical and organizational failures behind an incident, aimed at the condition that allowed it rather than the component that broke.
  • 5-Why technique. Asking why repeatedly, each time about the previous answer, until the chain reaches a process or policy cause. In government AI incidents it usually terminates in a governance gap.
  • Post-incident review. The structured examination conducted after resolution, typically within two to four weeks, producing an incident report, a corrective action plan, a monitoring update and a lessons-learned summary.
  • Corrective action plan. The list of changes that will prevent recurrence, each with a named owner, a date and a way to tell whether it happened.
  • Interim process. The manual or degraded procedure that carries the workload while a suspended AI system is out of service, including what happens to the timelines people depend on.

Closing

Yolanda's thirty minutes went reasonably well, and the reason is not that she knew the right answers. It is that she wrote things down as she went, suspended the batch run before she understood the cause, and told legal and leadership at the same time rather than in sequence. Those three habits are learnable in an afternoon and they are most of what separates a documented incident from a scandal. The rest of the work, the root cause analysis and the corrective actions, happens on a calendar rather than a clock.

What makes AI incidents feel different from ordinary system failures is scale and silence. An automated process makes the same mistake 340 times before breakfast, and unless somebody is watching the output, the first signal is a person whose food benefits stopped. Build the procedures now, while nothing is on fire, and decide in advance who classifies, who suspends and who calls the Inspector General. Then the incident becomes a bad morning with a good record, rather than a governance failure with a paper trail that proves it.

Key Takeaways

  • Classify before you act. Severity classification drives every subsequent decision. Classify within thirty minutes of detection, default to the higher level when unsure, and put the classification and its reasoning in writing.
  • Start the log immediately. Timestamped notes taken from the moment you suspect something is wrong are worth more than a polished report written hours later. Timestamps prove when you knew what you knew.
  • Notify legal and leadership together, not sequentially. For a Severity 1 incident the notification obligation is immediate, with a two-hour outer bound in the county example used here, and it runs for both audiences at once.
  • Proactive IG notification is almost always the right move. An agency that self-reports is treated differently from one that is discovered. For federally funded programs, notification may also be a regulatory requirement.
  • Affected individuals must hear from you first. People harmed by an AI error have a right to prompt, plain-language notification, to appeal, and depending on state law to emergency reinstatement. Tell them what happened, what you are doing, and how to get help now.
  • Document all five parts of the record. Incident report, investigation findings, immediate response, resolution and prevention. The two fields people skip, how long the problem ran and why it was not caught earlier, are the two an investigator reads first.
  • The 5-Why technique finds the governance gap, not just the software bug. Root cause analysis in government AI incidents almost always terminates in a policy or process failure upstream of the technical error. Fix the gap, not only the code.
  • Ask whether the failure is systemic. The same weakness usually exists in systems nobody has checked, so turn each incident into system, process and monitoring improvements rather than a single patch.
  • Government incidents are different because FOIA, civil rights law and public trust are always in the room. Every email, log entry and decision memo is a potential public record. Write accordingly.
  • Post-incident reviews need named owners and deadlines, not committees and intentions. Track completion at 30, 60 and 90 days.

Frequently Asked Questions

How fast do I actually have to report an AI incident? That depends on your agency's own policy, your program's funding conditions and, for some benefit programs, state law, so establish the clocks that bind you before you need them. The times used in this lesson come from one county agency's Severity 1 practice: leadership and legal within two hours, public affairs within four, the Inspector General and affected individuals within 24. Treat them as illustrative outer bounds rather than a standard, and note that the underlying expectation for a Severity 1 event is immediate notification, with those hours as the point past which you have a second problem.

Should I suspend the system before I know the cause? For Severity 1, yes; immediate shutdown is the presumption, and containment precedes diagnosis. The asymmetry is what drives it: if you suspend and the system turns out to be fine, you lose some throughput and explain a cautious call. If you leave it running and it is not fine, every additional batch adds people to the affected population and to the eventual liability. Have the interim manual process ready so that suspension is an option you can actually exercise.

Who decides the severity level? Decide this in advance in your incident response procedure, and name a role rather than a person so it survives turnover. Whoever first identifies the incident makes the initial classification within thirty minutes; it can be revised upward immediately or downward later with documented reasoning. Guard against a classification negotiated between people who each want a lower number.

Does proactive IG notification invite an investigation we would otherwise avoid? Sometimes, and it is still the right call. An IG who learns of an incident from the agency is dealing with a cooperative subject and a documented response. An IG who learns of it from a whistleblower, a reporter or a complainant is investigating both the incident and why nobody told them. The second investigation is broader, longer and far more damaging to the people running the program.

What if we cannot determine how many people were affected? Document the number you have, say explicitly that it is preliminary, and record how you got it. Then treat establishing true scope as an investigation task with an owner, driven by the question of how long the problem was occurring. Understating scope early is forgivable if you flagged the uncertainty and corrected it; presenting an unverified number as final is not.