Data Leakage: When Sensitive Info Enters AI
Janet Okoye, a caseworker at a county social services office, was drowning in a 14-page incident report due by five o'clock. So she did the reasonable thing: she pasted the whole report into a free AI chatbot on her personal phone and asked it to summarize. The summary was excellent. It also contained a child's full name, address, and the details of an abuse allegation, and so now did the chatbot's servers, possibly for a long time, possibly used to train the next version of the tool, possibly readable by that vendor's staff. Janet was not careless or malicious. She was busy, and nobody had ever shown her exactly where the data goes when you hit send.
That is what this lesson does. You do not need to be technical. You need an accurate picture of the pathways by which sensitive information escapes into AI tools, and a set of habits that close most of them. Get these right and you can use AI to do real work without becoming the next breach in the news.
The send button is a property line
Here is the one idea that makes everything else click. When you type into most AI tools, your words leave your agency's control the instant you hit send. They travel across the internet to a company's computers. You are no longer the only party holding that information, and you no longer decide what happens to it next.
Inside your agency, data sits behind locked doors, meaning access controls, logs, and rules about who may see what. The send button is the property line. On your side, your rules protect the data. On the other side, the vendor's rules apply, and you have almost certainly never read them. Janet crossed that line with a child's name and never registered that she had left the building. Treat the AI text box as an open window facing the street: assume anything you put there could be read by a stranger.
Approved is not the same as safe
Many people carry an assumption that is comforting and incomplete: if I use an approved AI tool, my data is safe. The tool may well be approved. The way you use it can still expose data. Approval is a judgment about the tool, its contract, its security assessment and the categories of information permitted to enter it. It is not a judgment about the specific thing you are about to paste, and it cannot be, because your agency did not know what you were going to paste.
Data leakage through AI tools is one of the most common ways sensitive information escapes government control, and leakage is breach. When sensitive data leaks, citizens are harmed, the government faces legal liability, and public trust erodes in ways that are slow to repair. The reason to understand the pathways in detail is not curiosity. It is that you cannot avoid a route you cannot see, and several of these routes operate entirely inside tools your agency has approved.
Five pathways, and what each one looks like
Leakage is not one event. It is five distinct routes, and the habits that close one do nothing about another. Work through them in order and notice how few involve an attacker.
Pathway one: cloud processing and storage
Most AI tools do not run on your phone or laptop. They run on the vendor's servers in some data center, and whatever you send travels there to be processed. That alone means sensitive data has left your agency. The data might be stored on servers in other countries. The vendor might retain it for training or research, might share it with third parties, and might be hacked. For a free consumer tool you have no contract, no guarantee of deletion, and no way to find out who can see it. Janet's report did not stay on her phone for a second.
What this looks like in practice: an agency uploaded a spreadsheet containing employee names and salaries to a cloud-based AI analysis tool. The tool's vendor retained a copy. Later the vendor was hacked, and the data was exposed. Nobody at the agency did anything unusual at any point in that sequence.
Pathway two: model training and extraction
Many free and consumer AI tools reserve the right to use what you type to train their future models. Your input can be absorbed into the tool itself, and once it is training material you cannot pull it back. This is the leak people least expect, because nothing was stolen. You gave it away in the terms of service.
There is a sharper version of this risk for agencies that train their own models. A trained machine learning model is in effect a compressed version of its training data, and under the right conditions the original data can be extracted or inferred from it. Membership inference determines whether a specific individual's data was in the training set. Model inversion reconstructs original data from the model. This is not theoretical, and it has been demonstrated in research. The practical consequence is that sharing a model trained on confidential government data with external partners can be closer to sharing the underlying records than it feels.
Pathway three: metadata
When you upload a document, you send more than the words on the page. Files carry hidden details called metadata: who created the file, when, the edit history, sometimes location data, and tracked changes you believed were removed. Author names, organization details, timelines and system information all travel with the file. A colleague who uploads the original document instead of pasting the text sends the file's entire hidden history along with it.
What this looks like in practice: a government employee uploaded a policy document to an AI analysis tool. The document's metadata revealed which agency it came from, who had been editing it, and internal names for classified projects. The text on the page was unremarkable. The properties attached to it were not.
Pathway four: cache and logging
AI systems cache data for performance and log queries for debugging and improvement. That cached and logged material has to live somewhere. It might sit on the vendor's servers, might be accessible to other users, might be used to train other models, and might be retained indefinitely. You will generally have no visibility into which of those applies, and deleting your conversation does not touch any of it.
What this looks like in practice: a government employee submitted a query to an AI system that included specific details about a sensitive investigation. The query was cached. Months later, when the cache was searched, the investigation details were visible.
Pathway five: the output you share
The last pathway runs in the opposite direction. You ask an AI system a question, it produces output, and you screenshot or copy that output and pass it on. The output may contain material you did not put there: information absorbed from its training data, inferences the system drew and stated as fact, or content from previously asked questions where the system displays history. The threat is simply sharing output without reading what is actually in it, which is easy to do when the output looks like a clean summary of your own request.
What happens after you hit send
It helps to see the sequence rather than the endpoints, because each step is its own opportunity for loss. Your data is transmitted to the service, ideally encrypted in transit. It is processed on the service's servers. It is stored, usually at least temporarily. It may be logged, meaning a record exists that this request was made with this content. It may be used for training. It may be retained beyond what is necessary for the task. The access controls governing all of that may be inadequate. And the vendor may be hacked.
Eight steps, each of which is a risk in its own right. Encryption in transit protects step one and nothing after it, which is worth knowing because "it is encrypted" is frequently offered as if it settled the question. It settles one step out of eight. The remaining seven are governed by a contract, and whether that contract exists and what it says is exactly the difference an approved tool makes.
Know what you are protecting
You cannot protect what you do not recognise. Four categories should stop you before you paste, and the boundaries between them matter less than the reflex.
- Personal information. Names, addresses, Social Security numbers, dates of birth, and anything else that identifies a specific person alone or in combination.
- Health information. Medical conditions, treatments and diagnoses tied to a person.
- Controlled or law-enforcement-sensitive information. Investigation details, security information, and anything marked for limited distribution.
- Classified information. The hardest line of all. Classified material must never touch a public AI tool, without exception and regardless of urgency.
The simple test underneath all four: if this got out, could it harm a person, an investigation, or national security? If the answer is yes, it does not go into a public AI tool. Notice that the test asks about consequences rather than labels, which is deliberate. Plenty of material that carries no marking at all would fail it, and the marking is not what protects the person.
Seven ways to close the gap
Janet's office adopted these after her near miss. None of them requires technical skill, and each one closes a different pathway.
- Prefer systems your agency controls. An AI system hosted on government infrastructure removes the vendor from the picture entirely, which takes vendor misuse, vendor sharing and vendor breach off the table. It does not make the system unhackable and it does not stop insider misuse. It moves the responsibility to your agency, which is where it can be governed.
- Minimize before you upload. Do not send full datasets. Send only what the task requires. The wrong version is uploading a spreadsheet of full employee records with names, salaries and addresses. The right version is uploading the aggregate you actually need, such as "15 employees in Department X, average salary $65K."
- De-identify before you upload. Remove direct identifiers first. Replace names with ID numbers, replace addresses with regions, and remove or generalise dates. Then read the result again, because de-identification can fail and the next section explains how.
- Strip metadata, and check the body too. Most office software has a tool for removing document properties. Use it. Then remember that tracked changes, comments and embedded content live in the document body rather than in its properties, so the stripping tool will not touch them. Better still, paste the clean text you need instead of uploading the file.
- Review AI output before you share it. Read what the system actually produced rather than what you asked for. Check for information you did not supply, inferences stated as fact, and anything carried over from earlier in the session.
- Read your tool's policies. For any system you use, find out whether the vendor retains data, for how long, whether they use it for training, who has access, and what happens if they are breached. If you do not like the answers, or you cannot find them, do not use that system with sensitive data.
- When in doubt, do not. The deadline pressure that pushed Janet to paste is precisely the moment to pause. A late report is recoverable. A child's identity in a vendor's training corpus is not.
What de-identification actually buys you
This is where well-intentioned staff most often go wrong, and it is worth being exact. Removing names is a control. It is not a proof of anonymity, and the two are not interchangeable. Quasi-identifiers such as age combined with zip code and a condition can re-identify individuals through linkage to other datasets. Narrative detail can do it without any structured field at all: a prompt about a household with two minors and a specific set of circumstances in a specific county may describe exactly one family, and every identifier in it was already bracketed out.
The same caution applies to aggregation. Small groups can be reversed, and a statistic about a category containing a handful of people is a statement about those people. This is why "only aggregated data, so no risk" is a mistake rather than a shortcut. De-identification and aggregation both reduce risk substantially and neither eliminates it, which means they change the category of data you are handling rather than removing the obligation to handle it properly.
Differential privacy is the most rigorous of these techniques and deserves the same precision. It adds controlled noise so that the presence or absence of any single record has a bounded effect on the result, while aggregate patterns are preserved. That is a measurable limit on how much any individual can be inferred, set by a privacy budget somebody chooses, and not a guarantee that no individual can ever be identified. The guarantee it offers is mathematical and conditional, it degrades as more queries are run against the same data, and it depends on the implementation being correct. It is a strong tool that some government agencies use, and it is advanced enough that the honest answer for most staff is to ask the person who configured it what the budget was.
The practical rule that follows: state what you did rather than what you achieved. "I removed the direct identifiers" is a true and useful sentence. "This is anonymized, so nobody can be identified" is a claim you are almost never in a position to make, and it is the sentence that ends the conversation right before somebody gets identified.
A ten-second decision card
Tape this next to your screen and run it before every paste or upload. It is deliberately short, because a check that takes longer than the shortcut will lose to the shortcut.
| Ask yourself | If the answer is | Then |
|---|---|---|
| Is this an agency-approved AI tool, and does my approval cover this kind of data? | No, or not sure | Stop. Ask first. |
| Does the text contain personal, health, controlled or classified information? | Yes | Remove it, then reread what remains. |
| Would the remaining detail let a colleague work out who this is? | Yes | Cut the detail or do not send the passage. |
| Am I uploading a file rather than pasting text? | Yes | Paste only the text you need, and strip the file if you must send it. |
| Could this harm a person, a case or national security if it leaked? | Yes | Do not use a public AI tool. |
A worked scenario: 500 comment letters
You are a policy analyst who wants to use an AI system to analyze citizen feedback on a proposed policy, and you have 500 comment letters. The risky approach is to paste all 500 into a summarization tool and ask it to identify common themes. Those letters contain names, addresses and contact information. The vendor may retain them, metadata in the original documents may reveal more, the names and addresses now sit on the tool's servers, and the material may be used for training. Four pathways opened by one paste.
The safer approach inverts the order of work. Read the letters yourself and extract the themes without the personal information: 50 comments expressed concern about cost, 120 about the implementation timeline, 80 in support, and so on through the rest of the set. Then upload only that aggregate summary and ask the AI to help organise the themes into categories. The system now works with counts rather than citizens, and no personal information has been uploaded at all.
Note what this does and does not achieve. It closes the pathway completely for the citizens' identities, because their letters never left your desk. It does not make the exercise cost-free: the aggregate still goes to a vendor, is still logged and cached, and the categories you are exploring still tell that vendor something about your agency's work. That residual exposure is usually acceptable and it is not zero, and knowing the difference is what lets you make the call rather than guess at it.
Janet still uses AI most weeks to draft and summarize. What changed is that she works from a de-identified version inside her agency's approved tool, and she rereads the de-identified version before sending it to check that the remaining detail does not point at one family. The work still gets faster. The child's name stays in the building.
Anti-Patterns to Avoid
Each of these is a reasonable-sounding sentence that has preceded a real exposure.
- "It is an approved tool, so it is safe." Approval covers the tool under stated conditions, not your particular use of it. You can leak thoroughly inside an approved system by pasting a category of data the approval never contemplated, and the software will not stop you or warn you.
- "We removed the names, so it is anonymized." Removing direct identifiers is one control. Age plus zip code plus a condition can re-identify through linkage, and a distinctive narrative can identify someone with no structured fields at all. Say what you removed, not what you believe you achieved.
- "Only aggregated data, so there is no risk." Aggregates over small groups are reversible, and a statistic over a handful of people is a statement about those people. Aggregation lowers the risk category. It does not exit the question.
- "Differential privacy means individuals cannot be identified." It bounds the influence any single record has on a result, under a privacy budget somebody selected, degrading as more queries are run. That is a measured limit, not an impossibility proof, and treating it as one is how a strong technique gets used carelessly.
- Sharing output without reading it. The system produces a summary, it looks like your request, you forward it. Outputs can carry training-derived content, confident inferences that are not facts, and material from earlier in the session that has nothing to do with the recipient.
- "It is a major vendor, they would not misuse data." Every vendor is a target for attackers, and a vendor's published policies may permit uses you never considered. Trustworthiness is not the same as a contractual restriction, and only one of the two is enforceable.
- Testing a system with real personal information. Trying a tool out with genuine records to see how it copes puts those records into the system's logs, where they may be retained and exposed whatever you concluded from the test. Test with fabricated data.
- Deleting the conversation to clean up. Deleting your side of a chat removes your record. It does not remove the vendor's copy, the cache, the logs, or anything already absorbed into training.
Practice Prompts
Work these against real material, because the hypothetical version is always cleaner than your actual caseload.
- Trace one upload. Pick something you have put into an AI system. Write down what the data was, where it is now, what could happen if it leaked, and which of the five pathways it travelled.
- Answer the five vendor questions. For the AI tool you use most, find out whether the vendor retains data, for how long, whether it trains on it, who has access, and what happens in a breach. If you cannot find the answers, that is an answer.
- De-identify and then reread. Take a real passage you would want AI help with, remove the direct identifiers, then read what remains and ask whether a colleague could still name the person. Cut until they could not.
- Strip a document. Take a file you would have uploaded, run your office software's metadata removal on it, then check the body for tracked changes and comments the tool did not touch.
- Audit an output. Take an AI output you shared recently and read it line by line for anything you did not supply, any inference stated as fact, and any content carried over from earlier in that session.
Reflection
Answer these honestly. Any question you cannot answer is the assignment.
- What have I actually put into an AI system this year, and could I list it if someone asked me tomorrow?
- Do I know where that data is now, or have I been assuming it went nowhere because nothing happened?
- What does my agency do to prevent leakage through AI tools, and could I describe it to a colleague?
- Have I ever called something anonymized when what I really did was remove the names?
- If I discovered today that data had leaked through a tool I used, would I report it, and how quickly?
Glossary
- Data leakage. When sensitive information escapes from where it was supposed to be protected, whether or not anyone intended it or attacked anything.
- Metadata. Data about data, including author, timestamps, edit history and system information, which travels with a file whether or not you can see it.
- Membership inference. Determining whether a specific individual's data was in a model's training set.
- Model inversion. Reconstructing original training data from a trained model.
- De-identification. Removing information that directly identifies individuals. It reduces risk rather than eliminating it, because quasi-identifiers and narrative detail can support re-identification.
- Quasi-identifier. A field that identifies nobody on its own but identifies individuals in combination with others, such as age plus zip code plus condition.
- Differential privacy. Adding controlled noise so that any single record has a bounded effect on a result while aggregate patterns survive. The bound is set by a privacy budget and weakens as more queries are run.
- Cache. Temporarily stored copies of data or queries kept for performance, which may persist and be searchable long after you have finished with them.
Related Lessons
Leakage sits inside a wider set of obligations, and these lessons supply what this one assumes.
- PII and AI: The Bright Red Lines defines the categories of citizen data that must never enter an unapproved tool, and the approvals required for the ones that may.
- Data Sensitivity and Classification teaches the labelling scheme that decides which of the four stop categories a given file falls into.
- Approved vs. Shadow AI covers the tool-choice half of this problem, including what an approval actually contracts for.
- Privacy Engineering for AI goes deeper on de-identification, synthetic data and systems that leak less by design.
- Privacy Impact Assessments for AI Systems is the formal instrument for documenting these risks before a system goes live rather than after.
- AI Incident Response: What to Do is the sequence to run in the first minutes after something has already gone out.
Closing
Data leakage is preventable, and preventing it is unglamorous. It requires awareness of the five pathways, discipline about which tool you open, the habit of asking questions about the systems you use, minimising data before you upload it, and reading output before you pass it on. Do those things and you cut the risk dramatically. Skip any one of them and you have left a route open that the others do not cover.
What makes this worth the effort is not compliance. It is that every one of these pathways ends at a specific person who had no choice about giving their information to the government. Janet's report was about a child. The summary that arrived in seconds was genuinely useful, and it was not worth what it nearly cost. Keep the productivity, close the pathway, and when you are not sure which side of the property line you are standing on, stop and ask before you send.
Key Takeaways
- The send button is a property line. Whatever you type into most AI tools leaves your agency's control immediately and lands on a company's servers under their rules.
- Approved is not the same as safe. Approval covers the tool and the permitted data categories, not the particular thing you are about to paste into it.
- There are five pathways, not one. Cloud storage, model training and extraction, metadata, cache and logging, and the output you forward without reading.
- Models can leak their training data. Membership inference and model inversion are demonstrated in research, so sharing a model trained on sensitive data can approach sharing the records.
- Metadata travels with the file. Author, edit history, timelines and internal project names have all been exposed by uploading a document whose visible text was unremarkable.
- Caches and logs outlive your conversation. Deleting the chat removes your copy and nothing else, and cached queries have surfaced months later.
- Minimize, then de-identify, then reread. Send the least data that does the job, remove direct identifiers, and check that the remaining narrative does not point at one person.
- De-identification reduces risk, it does not confer anonymity. Quasi-identifiers and distinctive detail support re-identification, and differential privacy bounds inference under a budget rather than making it impossible.
- Read the output before you share it. Outputs can carry training-derived content, inferences stated as fact, and material from earlier in the session.
- Know your tool's policies. Retention, duration, training use, access and breach handling. If you dislike the answers or cannot find them, keep sensitive data out.
Frequently Asked Questions
The tool says it does not train on my data. Does that close the risk? It closes one pathway out of five. The data has still crossed to the vendor's servers, where it is processed, stored at least temporarily, very likely logged, possibly cached, and subject to whatever retention and access controls the vendor operates. It is also reachable by legal discovery and exposed if the vendor is breached. A training toggle is a useful setting and it is the vendor's assurance to a user, not your agency's authorization for a category of data.
I pasted text rather than uploading the file. Am I clear of the metadata problem? Of that pathway, largely yes, and that is exactly why pasting clean text is the better habit. You have not avoided the others. The text still reaches the vendor's servers, is still logged and cached, and may still be retained or used for training depending on the terms. Pasting solves the hidden-history problem and leaves the property-line problem entirely intact.
How much detail do I have to remove before a passage is safe to send? Enough that a colleague reading it could not work out who it is about, which is a higher bar than removing the identifiers. Bracket the names, numbers and addresses first, then read what remains and ask that question honestly. A household with two minors and one distinctive circumstance in one county is frequently a unique description. If the passage still points at a person after you have stripped it, cut the detail or do not send the passage.
What if the leak has already happened? Report it, even late. Capture what was sent and when, do not delete the conversation or the account, and tell your supervisor and your privacy or security officer. Deleting removes your evidence and none of the vendor's copies. Agencies are judged on how they respond, and the option to notify the affected person only exists while somebody knows there is one. A late report is always better than a discovered one.
Is an on-premise or government-hosted AI system the answer to all of this? It is the answer to some of it. Hosting on infrastructure your agency controls removes vendor misuse, vendor sharing and vendor breach, which is a large share of the problem. It does not make the system unhackable, does not prevent insider misuse, and does nothing about the training-extraction pathway if you train models on sensitive data and then share them. What it changes is who is accountable, which is the point: it puts the risk somewhere your agency can actually govern it.
Skill.re