AI Incident Response: What to Do
Dwayne Carter processes vehicle registrations at a state department of motor vehicles. He is not a technologist. He uses an AI assistant his agency rolled out last month to summarize policy memos and draft replies to resident emails. One Thursday afternoon he pasted a resident's email into the assistant to draft a reply, and that email contained the resident's full Social Security number and date of birth. The moment he hit enter, a small voice said: was I supposed to do that? He froze. Close the window? Tell his supervisor? Pretend it never happened? Most government employees who meet an AI problem are exactly like Dwayne, which is to say frontline, busy, and unsure who to call. You do not need to be an expert to do the right thing. You need a sequence, and the nerve to run it.
Why the first minutes matter
An AI incident is any event where an AI system has failed, been compromised, or caused harm. In practice that means one of four situations: the system exposed data it should not have, it made a decision that was badly wrong, it was compromised in an attack, or you simply suspect something is off and cannot say why. That last one counts. Suspicion is a legitimate reason to run this sequence, and waiting until you are certain is how small incidents become large ones.
When an AI system does something wrong, the worst damage often comes not from the mistake itself but from what happens next. If the person who noticed stays silent, the problem keeps running. If they panic and start deleting things, they destroy the evidence needed to understand it. The first hours decide three separate outcomes: whether the incident is contained or spreads, whether evidence is preserved or lost, and whether the agency's liability is limited or compounded. None of those three depend on technical skill. They depend on what a non-technical person does in the first few minutes.
Three questions decide everything, and it is worth having answers to them before you need them. What do you do immediately, in the minute you notice? What do you do in the first hours, once the immediate danger is stopped? And what is the playbook after that, when other people are involved and you are no longer the one driving? Most staff have a rough answer to the first and no answer at all to the second and third, which is why incidents that were caught early still end badly.
Think of it as spotting smoke in a building. You do not need to know what is burning, how hot the fire is, or how the sprinklers work. You need to stop what you are doing, sound the alarm, and not run back in for your coffee mug. AI incident response for frontline staff has the same shape, in five steps: stop, document, report, cooperate, follow up. The first three are yours and they are urgent. The last two are how you stay useful once someone else takes over.
You are not expected to fix the AI. You are expected to notice, preserve, and tell someone. That is the whole job in the first five minutes.
Step one: stop
The instant something feels wrong, stop using the system for that task. Do not keep going to see if it corrects itself. Do not send the output to a resident, post it, or act on it. For Dwayne, stopping meant not sending the drafted reply and not pasting anything else into the assistant. The sensitive data was already submitted and he could not undo that, but he could stop making it worse. Stopping limits the harm, freezes the situation so it can be examined, and buys you a moment to think instead of react.
Three more things belong in this first step. Notify your supervisor immediately rather than waiting for an official report, and a verbal notification is entirely fine to begin with. Isolate the system if you can, meaning limit its access or shut it down, but only if you actually have that authority, and never in a way that destroys evidence. And do not try to fix it yourself. Reconfiguring a misbehaving system, retrying it with different settings, or quietly correcting its output all destroy the record of what it did.
What stopping looks like in the situations you are most likely to meet:
- The AI gave a resident wrong information. Do not send it. Flag it.
- You pasted sensitive data into a tool that should not have it. Stop entering more, and write down exactly what you entered.
- The AI produced something offensive or biased. Do not use it. Capture it.
- You suspect the tool has been tampered with or is behaving strangely. Stop using it entirely and report.
Step two: document, and do not delete
This is the step people get wrong. The instinct when something looks bad is to tidy it away: close the window, delete the chat, clear the history. Do the exact opposite. The record of what happened is what lets your agency understand the problem, notify anyone affected, and protect you by showing precisely what occurred and when. Deleting a chat does not remove the data from the vendor's systems. It only removes your ability to prove what you sent.
Write down five things while they are fresh, because memory decays faster than you expect. What happened, described as plainly as you can. When it happened, including the date, the time, and how long it may have been going on. Who is affected, meaning how many people and what information was involved. How you know, meaning what you actually observed that made you aware. And what the impact is, meaning what follows from this incident if nothing is done. Save that note somewhere you will find it again.
Alongside the written note, preserve the artifacts. Capture what you typed in, meaning the exact prompt or data you entered rather than a paraphrase. Capture what the AI gave back, the full output rather than your summary of it. Capture when and where, meaning the date, the time, and which tool or system. And capture what happened next, meaning whether the output was sent to anyone, acted on, or caught in time. A screenshot taken before you touch anything else is the cheapest and most useful of these.
Dwayne took a screenshot showing the resident's email, the assistant's draft, and the timestamp. That screenshot told the privacy team which record had been pasted and at what moment, which let them scope their inquiry to one resident rather than starting from nothing.
What a screenshot does not prove
Be precise about what your evidence establishes, because overstating it causes real harm. A screenshot proves what was on your screen at a moment in time. It does not prove what the vendor received, how long they retained it, whether it was logged or cached elsewhere, whether it reached a training pipeline, or whether anyone else saw it. Those questions are answered by the vendor contract, the system logs, and the security team, not by your screenshot.
This matters because a well-meaning employee who says "the screenshot shows it was only one record" can shut down an inquiry that should have kept going. Say what you saw. Let the people with access to the logs establish the scope. Your documentation is the starting point of the investigation, and treating it as the conclusion is one of the quieter ways an incident gets under-reported.
Step three: report
Now tell someone. Speed matters more than polish. You do not need a written report, and a phone call or message to the right person within the hour beats a perfect email tomorrow. There is a script, and using it is easier than improvising: say "I am reporting a potential AI system incident, and here is what I observed." Then describe what you observed factually. They will determine the next steps.
Report in this order of priority, escalating past anyone you cannot reach rather than waiting:
- Your supervisor, if you have not already told them in step one.
- Your agency's security team, usually IT Security or the Chief Information Security Officer, who triage technical incidents.
- Your agency's incident response team, if one exists.
- Your agency's privacy office, whenever any personal or sensitive information was involved.
- Your agency's legal office, where there are legal implications.
If you do not know who the privacy officer is, tell your supervisor that the incident involved someone's personal information and let them route it. You are not responsible for knowing the org chart. You are responsible for raising your hand, and for saying the words "personal information" out loud so the right office gets pulled in rather than finding out three weeks later.
That order is a priority list, not a queue you must complete in sequence. If your supervisor is unreachable, go to the security team and tell them you have not yet been able to reach your supervisor. If you cannot tell whether the privacy office needs to be involved, involve them and let them decide. The point of an escalation path is that no single unanswered phone can stop a report, and the most common way an incident goes unreported is that one person was in a meeting and everyone else assumed it had been handled.
One thing stops people reporting, and it is fear of getting in trouble. Agencies in line with federal AI guidance are expected to encourage reporting rather than punish it, because staff hiding mistakes is by far the more dangerous failure. That expectation is worth knowing, and it is not a personal indemnity. Reporting will not always be consequence-free, particularly where a rule was knowingly broken. What is reliably true is the comparison: a reported incident is a containable problem, while a concealed one becomes a second and more serious problem about the concealment. Dwayne told his supervisor within ten minutes, and his supervisor's first words were thank you for catching that.
Step four: cooperate
Once you have reported, your agency's incident response team takes over and your role changes from actor to witness. Answer their questions about what you observed, and stick to what you actually saw rather than what you have since inferred. Hand over any documentation you have, including the messy notes. Follow their instructions, including instructions to stop doing things you thought were helpful. And maintain confidentiality, because incidents under investigation are usually confidential and discussing one in the break room can compromise both the inquiry and the people affected.
Confidentiality during an investigation is not secrecy for its own sake. An incident under investigation usually involves someone's personal information, an unresolved technical weakness, or both, and discussing either one outside the response team can spread the exposure further than the original failure did. It also contaminates the account: once several people have compared recollections, no one can say cleanly what they personally observed. Keep your account to yourself and to the responders, and let the agency decide what gets communicated and when.
Step five: follow up
After the incident is resolved, close the loop for yourself. Ask what actually happened, because the root cause, meaning the underlying reason the incident occurred, is often different from what it looked like on day one. Ask what is being done to prevent a recurrence. Ask how you will be notified of the outcome. If the incident was serious there may be a lessons-learned session, and going to it is the single cheapest way to stop the same failure reaching the next person. In Dwayne's case the agency added a warning banner to the tool reminding staff never to paste personal identifiers, and notified the affected resident.
Four kinds of incident, and where each one goes
The five steps are the same every time, but who you call first depends on what kind of incident you are looking at. These four cover most of what frontline staff actually encounter.
| Type | What it looks like | Immediate action | Report to |
|---|---|---|---|
| Data exposure | You believe the system exposed personal information or other sensitive data. | Stop using the system. Preserve any evidence of what was exposed. | Your security team and privacy office, immediately. |
| Wrong decision | The system made a decision that harms someone, such as denying a benefit to a person who qualified. | Stop using the system for that purpose. Document the harm. | Your supervisor, the system owner, and any relevant legal or compliance office. |
| System compromise | You suspect the system was hacked or manipulated. | Stop using the system. Do not touch anything. Preserve evidence. | Your security team, immediately. |
| Unexpected behavior | The system is producing outputs that are wrong, biased, or simply strange. | Document examples of the behavior. | The system owner or team, and your supervisor. |
The routing differs because the clock differs. Exposure and compromise are time-critical for people outside your chain of command, so they go to security and privacy first and your supervisor learns about it in parallel. A wrong decision is time-critical for the person it harmed, so the system owner and the legal or compliance office matter as much as the technical team. Unexpected behavior is the one case where the system owner is the right first call, because what you are reporting is a defect rather than an event. When you cannot tell which category you are in, treat it as exposure and let security downgrade it.
A usable artifact: the quick-response card
Print this. Tape it inside a desk drawer or pin it near your screen. When something feels wrong you should not have to remember the steps, you should be able to read them.
| Step | Do this | Do NOT do this |
|---|---|---|
| 1. STOP | Stop using the tool for this task. Do not send or act on the output. | Do not keep going hoping it fixes itself. Do not try to fix it yourself. |
| 2. DOCUMENT | Screenshot it. Note what you entered, what it returned, the time, and the tool. | Do not delete the chat, close without capturing, or clear history. |
| 3. REPORT | Call or message your supervisor now. Say if personal data was involved. | Do not wait to write a perfect report. Do not stay silent. |
| 4. COOPERATE | Answer questions, hand over your notes, follow instructions, keep it confidential. | Do not speculate on the record or discuss it outside the response team. |
| 5. FOLLOW UP | Ask what happened, what changes, and how you will hear the outcome. | Do not assume silence means it was handled. |
| My contacts. Supervisor: __________ IT and Security: __________ Privacy Officer: __________ | ||
Fill in those three numbers today, while nothing is on fire. The whole point of the card is that the next Dwayne, who may well be you, does not have to work out who to call while a clock is running.
A worked scenario: the benefit approvals
You are using an approved AI system to process benefit applications. You notice it is approving at a much higher rate than usual, including people who clearly do not qualify. The approvals go into effect in two hours. Nothing about this is subtle, and yet the common failure here is to keep processing while you decide whether it is really a problem.
Run the sequence. Stop processing applications immediately and approve or deny nothing further. Then document what you observed, using a form of words like this: on [date and time], I noticed the system approving applications that clearly do not meet eligibility criteria; I reviewed [number] applications and found [percentage] were incorrectly approved; this appears to have been occurring for approximately [duration]. Fill in the brackets from what you actually counted, and leave a bracket empty rather than guessing at it.
Then call your supervisor: I have observed a potential serious problem with the AI system approving applications, I have documented it, and we need to notify security and stop processing immediately. Notify your security team and the system owner. Then cooperate while they investigate. This kind of quick response is what prevents hundreds of people from being incorrectly approved and turns a bad afternoon into something smaller than a crisis.
Anti-Patterns to Avoid
Each of these is something a reasonable person does under pressure, and each one makes the incident worse.
- Trying to fix it yourself. You notice the system behaving badly, so you reconfigure it, retry it, or quietly correct its output. You have now destroyed the evidence of what it originally did, you may have made the underlying problem worse, and you have almost certainly stepped outside your policy. The impulse is competence. The effect is obstruction.
- Waiting to see if it resolves itself. You observed something suspicious but you are not sure it counts as an incident, so you wait. Meanwhile the problem keeps running, more people are affected, and recovery gets harder with every hour. Suspicion is a sufficient reason to report. Certainty is not the entry requirement.
- Reporting casually. You mention it to a colleague in passing: the AI system did something weird. Nothing is formally logged, no one owns it, and the response either happens late or does not happen at all. A hallway comment is not a report. Say the words "I am reporting a potential AI system incident" to someone whose job includes receiving them.
- Overstating or downplaying it. You panic and describe something minor as catastrophic, or you soften something serious so it sounds manageable. Both are costly: one burns credibility you will need next time, the other delays the response. Describe what you observed factually and let the responders size it.
- Treating your own evidence as the conclusion. Your screenshot shows one record, so you report that only one record was affected. What it actually shows is one record on your screen. Scope is established from logs and vendor records, not from what you happened to capture.
- Deleting to tidy up. Closing the window, clearing the history, or deleting the account feels like containment. It removes none of the data from anyone else's systems and removes all of your ability to show what happened.
Practice Prompts
Do these while nothing is wrong. Every one of them is useless to discover mid-incident.
- Find your responder. Identify who leads AI incident response in your agency and confirm you can actually reach them today, by a route that does not depend on the system that failed.
- Trace the reporting path. Write down how you would report an incident, what the escalation path is above your supervisor, and what you should preserve as evidence. If you cannot answer these, that gap is the finding.
- Fill in the card. Complete the three contact lines on the quick-response card and put it somewhere you will see it from your desk.
- Rehearse the sentence. Say out loud the words you would use to open a report, including the phrase "this involved someone's personal information," so it is not the first time you have said it.
- Practise the documentation template. Take a recent AI output you were unsure about and write the five documentation answers for it: what, when, who is affected, how you know, and what the impact is.
Reflection
Answer these honestly rather than aspirationally. The gap between what you would like to do and what you would actually do late on a Friday is the thing worth finding.
- If I suspected an AI incident right now, what would my very first action be, and would it preserve evidence or destroy it?
- Have I ever noticed an AI output that was wrong and simply not used it, without telling anyone? What made that feel like enough?
- If the AI system made a decision that harmed someone in my caseload, would I recognise it as an incident, or as a bad day?
- What would make me hesitate to report, and is that hesitation about the agency or about me?
- Do I know, without looking anything up, which office to call when personal information is involved?
Glossary
- Incident. An event where an AI system has failed, been compromised, or caused harm. Suspicion of one counts, and confirmation is not a prerequisite for reporting.
- Incident response. The process of addressing a security incident or system failure, from first notice through containment, investigation, and follow-up.
- Evidence preservation. Keeping records, logs, screenshots, and documentation from the incident intact and unmodified so the sequence of events can be reconstructed later.
- Root cause. The underlying reason an incident occurred, which is frequently not the thing that first appeared to be the problem.
- Isolation. Limiting a system's access or shutting it down to stop ongoing harm, taken only by someone with the authority to do it and never in a way that destroys evidence.
- Escalation path. The defined order in which an incident moves upward and outward, from supervisor to security to privacy and legal, so that no report depends on one person being available.
Related Lessons
This lesson is the first five minutes. These lessons are what surrounds it.
- How AI Changes the Threat Landscape explains why these incidents are more frequent and different in kind than the ones your agency trained for.
- Recognizing AI-Generated Threats covers the impersonation and phishing cases that often trigger the compromise category above.
- Data Leakage: When Sensitive Info Enters AI is the detailed treatment of the exposure type, including the pathways you cannot see.
- PII and AI: The Bright Red Lines sets out the categories of data that make an incident a reportable breach rather than a mistake.
- Prompt Injection and Manipulation explains one of the ways an AI system starts behaving in ways nobody at your agency asked for.
- Reporting AI Concerns in Your Agency goes further into the channels, protections, and what happens after your report lands.
Closing
Incident response is serious business, and the part of it that belongs to you is not technical. If you suspect something is wrong with an AI system, report it. Do not minimise it, do not sit on it while you work out whether it counts, and do not cover it up. Quick reporting limits the damage and gives your agency the option of handling the problem while it is still small.
Dwayne did not do anything clever. He stopped, took one screenshot, and picked up the phone inside ten minutes. That was enough to turn an exposure into a contained and documented event with a named affected person and a fix that protected everyone who came after him. The sequence is stop, document, report, cooperate, follow up. Write the three phone numbers on the card today and you have done the only preparation this lesson asks of you.
Key Takeaways
- Five steps, in order: stop, document, report, cooperate, follow up. You do not need to be a technologist to handle an AI incident correctly, only to follow the sequence.
- Stop means stop now. The instant something feels wrong, stop using the tool for that task, do not send or act on its output, and do not try to fix the system yourself.
- Preserve, do not delete. A screenshot of what you entered and what the AI returned is the single most useful thing you can do, and deleting the chat removes your proof without removing anyone else's copy.
- Document five things. What happened, when, who is affected, how you know, and what the impact is, written down rather than trusted to memory.
- Your evidence starts the inquiry, it does not size it. A screenshot shows what was on your screen, not what a vendor retained, logged, cached, or shared.
- Report fast, not perfectly. A phone call to your supervisor within the hour beats a flawless email tomorrow, and suspicion is enough of a reason.
- Say the words "personal information" out loud. That is what pulls the privacy office in rather than leaving them to find out weeks later.
- Reporting is the containable path, not an indemnity. Agencies are expected to encourage reporting rather than punish it, and concealment reliably turns one problem into two.
- Match the type to the contact. Exposure goes to security and privacy, a harmful decision goes to the supervisor, system owner and legal, compromise goes straight to security, and strange behavior goes to the system owner.
- Fill in the card today. Supervisor, IT and security, privacy officer, written down now, so the next person in trouble does not have to search.
Frequently Asked Questions
I am not sure it was really an incident. Should I still report it? Yes. An incident includes the case where you suspect something is wrong and cannot articulate why, and the entry requirement for reporting is suspicion rather than certainty. The cost of a report that turns out to be nothing is a short conversation. The cost of waiting until you are sure is that the system keeps running, more people are affected, and recovery gets harder. Describe what you observed factually and let the responders decide what it is.
I already deleted the chat before I read this. What now? Report it anyway, and say clearly that the chat was deleted and roughly when. Deleting your copy does not delete the vendor's, and system logs, account records and vendor-side retention may still let the security team reconstruct what happened. What makes this much worse is not the deletion but the silence that usually follows it, because the team then discovers both the incident and the missing evidence at the same time and from someone other than you.
My supervisor told me not to escalate it. Do I stop there? The escalation path exists precisely so that no single person can end a report. If personal or sensitive information was involved, the privacy office has an independent interest in knowing, and the security team triages technical incidents regardless of where the report originates. Ask your supervisor to route it, and if that does not happen, contact the security team or privacy office directly and say that you have already raised it with your supervisor. Document that you did so.
How much detail should I put in the first report? Enough to let someone act, which is less than you think. What you observed, when, which system, and whether personal information was involved is a complete first report. Do not delay the call to assemble a full account, and do not fill gaps with inference to make the account sound tidy. Say what you saw, say what you do not know, and hand over your notes and screenshots when asked.
Will I get in trouble for reporting my own mistake? Agencies in line with federal AI guidance are expected to encourage reporting rather than punish it, because staff hiding mistakes is the far more dangerous pattern. That is not a promise of immunity, and it does not erase the consequences of knowingly breaking a rule. What it does reliably change is the shape of the problem: a reported incident is an incident, while a concealed one becomes a separate and more serious matter about the concealment. Reporting is also the only version of this where the affected person gets told.
Skill.re