←
AI for Government
Aware · M9 · lesson 9 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
PII and AI: The Bright Red Lines
📖
now learning

PII and AI: The Bright Red Lines

10 min

Marcus Bell, an eligibility caseworker at a county human-services office, had forty-one applications to clear before lunch. To speed up a tricky denial letter, he pasted the whole case file into a public chatbot and asked it to "rewrite this so it's clearer and kinder." The file held the applicant's full name, Social Security number, address, two children's names, and a note about a domestic-violence shelter. The letter came back beautifully worded. By then the damage was already done: that data had left the building.

Nothing exploded. No alarm sounded. That is exactly what makes this the most common and most dangerous AI mistake in government. Marcus did not hack anything or break a rule he had read. He just did not know where the bright red line was. This lesson draws that line, explains why it sits where it sits, and gives you concrete strategies for doing AI work safely when you legitimately need to work with data that relates to real people.

Be clear from the start about the register here. Some things you absolutely cannot do. You cannot upload a citizen's Social Security number into an AI system. You cannot use personal medical information to train a model. You cannot feed an AI system the names and addresses of vulnerable populations. Not because it is inconvenient, not because it is a best practice worth reconsidering, and not because in your specific case the risk might be low. These are hard lines because the risks are existential: to individuals, to your agency, and to public trust.

What actually happened when Marcus hit enter

"Personally identifiable information," usually shortened to PII, is any information that can be used to identify a specific person, alone or combined with other data. A Social Security number is obvious PII. So is a full name plus a home address. So is the fact that someone uses a particular shelter.

When Marcus pasted the case file into a public AI tool, three things happened at once that he could not see:

  • The text was sent over the internet to a company's servers, often in another state or country.
  • Depending on the tool's settings, that text could be stored and used to train future versions of the model. It may not be deletable on request.
  • The county lost control of the data. It cannot say who at the vendor can read it, how long it is kept, or whether it will surface in someone else's answer later.

For a private person sharing their own recipe, none of this matters. For a government employee sharing a citizen's data, it can violate the Privacy Act of 1974, breach the agency's terms of service with the vendor, and trigger a breach notification that reaches the applicant, agency leadership and sometimes the press. One paste can become a months-long incident. The AI did not steal anything. A trusted employee handed it over, politely, in plain sight.

Why the stakes are this high

PII is the most personal information an individual shares with government. When someone hands over a Social Security number, a medical history or a description of their family composition, they are trusting government to keep it secure and use it only for the purposes they understood. When that trust fails, the consequences land on the person, not on the system: identity theft, fraud victimization, stalking or harassment, wrongful investigation, discrimination, and the plain loss of privacy and dignity.

The agency pays too, in legal liability, loss of public trust, operational disruption and regulatory scrutiny. Those costs are real but they are recoverable. The harm to an individual whose shelter address became public is frequently not.

AI amplifies every one of those risks in four specific ways. These systems process data at scale, thousands of records at a time rather than one file on one desk. They combine data in new ways, producing inferences and predictions that can reveal sensitive facts nobody entered. They are sometimes opaque, which makes them hard to audit and hard to know what is happening inside. And they are sometimes shared or deployed broadly, which multiplies the surface area for exposure. That combination of scale, inference, opacity and reach is why the red lines exist rather than a set of gentle recommendations.

The bright red lines: what never goes in

You do not need to memorize privacy law to stay safe. You need to recognize a short list of data types that must never be typed or pasted into any AI tool that has not been formally approved by your agency for that exact purpose. Treat each of these as a hard stop.

  • Direct identifiers: Social Security numbers, passport or driver's-license numbers, full dates of birth, financial account numbers, biometric data.
  • Names tied to a sensitive status: who is on benefits, who filed a complaint, who is under investigation, who is a victim, who has a medical or immigration condition.
  • Protected health information: diagnoses, treatment notes, prescriptions, anything that would be covered by health-privacy rules.
  • Law-enforcement and security data: case details, informant identities, investigative techniques, anything marked Controlled Unclassified Information.
  • Classified information of any level: never, under any circumstance, in any commercial tool.
  • Other people's PII in bulk: spreadsheets, rosters, mailing lists, exports from your case system.

The simplest test: if the information could embarrass, endanger, or expose a real person were it made public, it does not go into an unapproved AI tool.

Five rules behind the list

The data types above are what you scan for. The five rules below are why the list looks the way it does, and they cover situations a scan alone would miss.

Never upload PII to unauthorized external AI systems. This is the clearest line. You cannot copy a spreadsheet of citizen names, addresses and personal details and upload it to a public, consumer-facing AI service. Not in a test environment. Not to quickly see whether the AI can help. Not because you trust the service and assume it is probably secure. Not if you "anonymize" the names but leave the other details intact. Once data leaves your secure government systems it can be accessed by foreign adversaries, used to train other models, retained longer than you expect, breached, subject to legal discovery if the company is sued, or shared with data brokers. It is no longer yours to protect.

The pattern has played out in reported cases: staff uploading code containing PII that became available to competitors, a hospital employee uploading patient data to a cloud service to test a workflow only for it to be retained and later included in model training, and government employees sharing citizen data with commercial services that then used it for research and passed it to affiliates. The penalties run from disciplinary action and termination for cause through loss of security clearance and legal liability, up to criminal charges in severe cases.

Never use an AI system with PII without explicit legal authority and security review. This applies to internal government systems too, not only to public tools. You cannot decide on your own that your agency's need is urgent enough to skip the security review and the legal check, because "we need this AI output today" does not override security requirements. Before using any AI system with PII, get legal approval confirming your statute authorizes the use, security approval confirming the system is secure enough for that data, and privacy office approval confirming the use respects privacy expectations. Then document the approval, because an approval nobody can find is functionally an approval that never happened.

Never combine data sources to re-identify de-identified data. You have a dataset with names removed and ages generalized, and linking it to other databases using characteristics such as age, zip code and condition would give you more to work with. Do not do it. That is a direct violation of the de-identification promise made to the people in the dataset, and it re-identifies data that was supposed to be anonymous.

Never use one purpose's data for another purpose. Data collected from citizens for benefits determination cannot be repurposed for immigration enforcement, or for fraud detection in a different programme, without legal authority for the new use, notice to the individuals (usually), and new privacy protections specific to the new purpose. Reusing data for new purposes without authorization is a breach of trust and often a violation of law, and it is the failure mode that most reliably destroys public willingness to give government accurate information.

Never train a model on sensitive data and then share the model. A model trained on PII that turns out accurate and useful is exactly the thing another agency or a contractor will ask for, and handing it over without extensive review is riskier than it looks. The model may have memorized PII from its training data, and attacks exist to get it back out: membership inference, which determines whether a specific individual's data was in the training set, and model inversion, which reconstructs original data from the model. Sharing a model trained on PII can be as risky as sharing the original records.

How exposure actually happens

Three cases make the abstract concrete. In the first, a law enforcement agency ran an AI facial recognition system trained on mugshot databases, millions of images with names and identifying information attached. An investigative journalist obtained the training dataset and found facial images of ordinary citizens, people never convicted of anything, collected at driver's license bureaus and used to train a system capable of tracking them. They had never consented and had no idea. The exposure was facial images plus identities, used for purposes the subjects never understood or approved, and it ended in legal suits, regulatory scrutiny, reputational damage and the system's retirement.

In the second, a state agency built a model to predict which welfare applicants might be engaging in fraud, trained on years of case data carrying employment history, family composition, financial details and even mental health information. The model performed well, so the agency planned to share it with other states. Before that happened, a security researcher showed that membership inference attacks were possible, meaning an attacker could determine whether a specific individual's data was in the training set. The identities of people investigated for fraud were extractable from the model itself. Sharing plans were abandoned, the training data was anonymized and the work redone, and the agency took the reputational hit anyway.

In the third, an agency contracted a vendor to build an AI system and supplied de-identified case records with detailed personal information as training data. The vendor delivered the system. Nobody had explicitly documented that the vendor could not reuse the data for other purposes. Years later the vendor used the same dataset to train a different model for a private company, which used it in ways that violated the original individuals' privacy expectations. The result was legal liability, breach notification obligations and a permanent loss of trust in the vendor relationship. Notice what failed there: not the technology, and not anyone's good intentions, but a missing clause.

How to sanitize an input in under a minute

Here is the good news. Marcus could have gotten his clearer, kinder letter with far less risk. The trick is to strip the specifics before you prompt, then add them back yourself afterward. AI is excellent at structure and tone, and it does not need the real names to help with either.

Watch the same task done safely. This is the placeholder method:

  • Before (unsafe): "Rewrite this denial letter to Maria Gonzalez, SSN 123-45-6789, at 412 Oak Street, who was denied SNAP because her income of $2,140 exceeds the limit..."
  • After (safer): "Rewrite this benefits-denial letter to be clearer and more compassionate. Keep it under 200 words. The applicant is [NAME]. The benefit is [PROGRAM]. The reason for denial is that reported monthly income of [AMOUNT] exceeds the program limit of [LIMIT]. Explain the appeal right plainly."

The AI returns a polished template with the brackets intact. Marcus then opens his own secure case system and fills [NAME], [AMOUNT] and the rest by hand, so the citizen's real identifiers never travel. He gets the quality without the paste.

Understand precisely what the placeholder method buys and what it does not. It removes direct identifiers from the text you send. It does not make arbitrary text safe, because narrative detail can identify someone on its own. A rare diagnosis, a named shelter, a programme with a handful of participants in one county, an unusual sequence of dates: any of these can point at one person with the name already stripped. The rule from the red lines applies here too. Anonymizing the names while leaving the other details intact is not sanitization. Strip the detail as well, or do not send the passage.

The five-second pre-flight check

Before you press enter on any prompt, run this quick scan. Print it, tape it to your monitor, and use it until it becomes a reflex.

  1. Approved tool? Am I in a tool my agency has cleared for work data? If not, only generic, non-sensitive content.
  2. Any real names or numbers? Scan for names, SSNs, addresses, account numbers. Replace each with a bracketed placeholder.
  3. Sensitive status? Does the text reveal a health, legal, financial or victim status tied to a person? Strip it or stop.
  4. Bulk records? Am I pasting a spreadsheet or export of many people? Hard stop.
  5. Would I email this to a stranger? If the answer is no, do not put it in the AI.

Five ways to work safely with sensitive data

Realistically, you will sometimes need AI on work that touches personal data. The question is how to do it without crossing a line, and there are five approaches worth knowing, each with its own limit.

  • Use aggregated data instead of individual data. Rather than "here are 10,000 individual benefit applications, analyze them," work from summary statistics: 40% of applicants in region X are single parents, 25% are unemployed, average age is 34. You get insight without individual PII. The limit: aggregation can sometimes be reversed to reveal individuals, especially when the groups are small, so small cells deserve the same suspicion as raw records.
  • De-identify carefully. If you genuinely need individual-level data, remove direct identifiers such as name, Social Security number and exact address; generalize quasi-identifiers, turning exact age into an age range and exact address into a zip code; test whether re-identification is possible through linkage to other datasets; document your method; and have a security expert review it. Each of those steps is load-bearing, and skipping the testing step is how organizations discover their de-identification failed after publication rather than before.
  • Use synthetic data. Train a model on real data, keep the original out of circulation, and use the model to generate artificial records with the same statistical properties but no real individuals behind them. Train on 100,000 employment records containing sensitive information, generate 100,000 synthetic ones, and work with those. The limit: synthetic data can sometimes be reverse-engineered or linked back to real data, so it is much safer rather than absolutely safe.
  • Use only approved government AI systems. Some agencies run secure AI systems built specifically for sensitive work: hosted on secure government servers, layered security, audit logging, security-cleared staff, and compliance with the applicable data handling requirements. Using these is vastly safer than commercial services or unsecured systems, and it is usually faster than arguing for an exception.
  • Minimize retention. If PII does go into an AI system, keep it only as long as you need it, delete it when the project completes, resist keeping it "just in case," and document the deletion and have it verified. The less PII you hold, the less can be exposed when something goes wrong, and something eventually does.

A worked scenario

You work in a social services agency and want to use AI to improve intake, helping caseworkers understand which clients need which services. The data you would like is casework notes, which often carry detailed personal information, along with assessment scores, service history and client demographics. The bright red line is immediate: you cannot upload this to a public AI service. That does not end the project. It shapes it.

Here is what you can do. Get legal approval for using this data for this purpose, asking whether your statute allows it and whether clients need notice first. Get security approval, confirming the data can be used on an approved government AI system. Establish data handling procedures covering who has access, how long data is retained and how it is deleted. Test the system for privacy leakage, checking whether anyone can reverse-engineer outputs to work out what personal information sat in the training data. Limit deployment to controlled environments before going agency-wide. Then monitor and audit: who is using the system, what they are querying, and whether anything suspicious has happened.

When something slips through

People make mistakes under pressure. If you realize you have pasted PII into an unapproved tool, the worst move is to stay quiet and hope. The clock starts the moment it happens, and agencies are judged far more on how fast they respond than on the slip itself.

Do three things immediately:

  • Stop and screenshot. Capture what you sent and when. Do not delete your account to "clean up." Investigators need the record.
  • Report it now, the same hour, to your supervisor and your privacy or security officer. Most agencies have a 1-hour or same-day reporting expectation for suspected breaches.
  • Do not keep using the tool for that work until you are told it is cleared.

Reporting a paste is a five-minute conversation. Hiding one that later surfaces is a career event. Choose the five minutes.

Anti-Patterns to Avoid

Every one of these is a sentence someone said out loud, in good faith, shortly before an incident.

  • "Just one quick test." Someone needs to try an AI approach and is impatient with the approval process, so they run a quick test on real data containing PII. That test is how a vulnerability gets exposed, how PII leaves the building, and how a precedent forms that the next person will cite. The urgency is real; it is still not authority.
  • "Anonymized, so it is safe now." Names come off and the data is treated as safe to use anywhere. But age plus zip code plus medical condition can re-identify individuals, and narrative detail can do it without any structured field at all. Removing the names is one control, not a proof of anonymity, and treating the two as equivalent is the single most common way well-intentioned staff expose people.
  • "The vendor promised confidentiality." An agency shares sensitive data with a contractor on a verbal promise or a thin contract. The contractor misuses it, sells it, or gets breached, and the agency finds it has no clear contractual recourse. Write the restriction down, including what happens to the data when the work ends.
  • "Everyone handles PII, why can't AI?" The argument sounds fair and misses the scale and opacity differences. A caseworker reads one file; a system processes thousands and infers across them, and the exposure is harder to detect precisely because it is automated.
  • Sanitizing the names but not the story. A prompt with every identifier bracketed can still describe exactly one person in your county. If the surrounding narrative would let a colleague name the client, the placeholders have not done their job.

Practice Prompts

Work these against your real caseload, not a hypothetical one.

  • Inventory your PII. Write down what personal data you personally work with, how it is currently stored and protected, and whether anyone has proposed using it in an AI system.
  • Check the approvals. For any AI system your agency uses or is considering that touches PII, find out whether it has legal approval and security approval. If you cannot find out, that is your first finding.
  • Sanitize a real prompt. Take a task you would genuinely want AI help with, rewrite it using bracketed placeholders, then reread the result and ask whether the remaining detail still identifies anyone.
  • Design a de-identification. If you had a legitimate need to analyze data containing PII but no approved system, write out how you would de-identify it, including how you would test the result for re-identification risk.
  • Name your contact. Identify the specific person or office you would ask about whether something is safe to do with PII and AI, and confirm you know how to reach them today.

Reflection

Give these ten honest minutes. If you cannot answer one of them, that gap is the assignment.

  • What PII do I work with, and would I be able to list every place it currently lives?
  • Have I ever been tempted to do something with data "just as a test" without proper approval? What stopped me, or did not?
  • If a prompt of mine leaked tomorrow, whose life would it complicate, and how much?
  • Do the approvals covering the AI tools on my desk actually exist in writing, or do I assume they do?
  • If I made Marcus's mistake late on a Friday, do I know who I would call, and would I call them?

Glossary

  • PII (Personally Identifiable Information). Information that can identify an individual, alone or combined with other data: name, Social Security number, address, phone, email and more.
  • De-identification. Removing direct identifiers from data so that individuals are not directly identifiable. It reduces risk rather than eliminating it, because quasi-identifiers and narrative detail can still support re-identification.
  • Re-identification. Determining the identity of individuals in supposedly de-identified data through linkage to other datasets or pattern matching.
  • Quasi-identifier. A field that is not personally identifying on its own but can identify individuals in combination with others, such as age plus zip code plus medical condition.
  • Membership inference. An attack that determines whether a specific individual's data was used to train a machine learning model.
  • Model inversion. An attack that reconstructs original training data from a trained model.
  • Synthetic data. Artificial records generated to share the statistical properties of real data without representing real individuals, safer than the original but not immune to reverse-engineering.
  • Data minimization. The principle of collecting and using only the minimum data necessary for a stated purpose.

The red lines sit inside a wider set of obligations, and these lessons supply the parts this one assumes.

Closing

The bright red lines are not bureaucratic obstacles. They are the minimum standard for a government people can safely tell the truth to. Every time you respect them you are protecting a specific person who never got to vote on how their file was processed, and every time you cross one you are risking real harm to someone who trusted you with information they had no choice but to give.

None of this requires you to give up the tool. Sanitize the input, use the approved system, get the approvals documented, keep the retention short, and ask when you are unsure. Compliance is not the enemy of speed here; it is the thing that keeps a five-minute task from becoming a months-long incident. Do the right thing, follow the lines, and ask for help when you are uncertain.

Key Takeaways

  • PII leaving the building is the real risk. The danger is not the AI itself but trusted staff handing citizen data to outside servers the agency cannot control or recall.
  • Memorize the hard stops. Direct identifiers, names tied to a sensitive status, health data, law-enforcement and CUI material, classified information and bulk records never go into an unapproved AI tool.
  • Approval is required even internally. Using PII in any AI system needs explicit legal authority, security review and privacy office sign-off, documented, and urgency never overrides that.
  • Purpose limits travel with the data. Data collected for one purpose cannot be reused for another without legal authority, notice and new protections built for the new use.
  • Models can leak their training data. Membership inference and model inversion mean sharing a model trained on PII can be as risky as sharing the records themselves.
  • Sanitize with placeholders, and sanitize the story too. Bracketed placeholders remove identifiers; distinctive narrative detail can still identify one person, so strip that as well or do not send the passage.
  • De-identification reduces risk, it does not guarantee anonymity. Quasi-identifiers such as age, zip code and condition support re-identification through linkage, so test it rather than assuming it.
  • Run the five-second pre-flight check. Approved tool, no real identifiers, no sensitive status, no bulk records, and the email-a-stranger gut check before every prompt.
  • Report fast, do not hide. If PII slips through, screenshot it, report the same hour, stop using the tool, and let the process protect both the citizen and you.

Frequently Asked Questions

I only pasted one paragraph, not a whole file. Is that different? Not in kind. The question is whether the text can identify a person, and one paragraph frequently can: a name, an address, a case number, or a description distinctive enough to point at one household. The volume affects how large the incident is, not whether one occurred. Run the pre-flight check on the paragraph exactly as you would on the file.

The tool says my conversations are not used for training. Is it safe then? That setting narrows one risk out of several. The data has still left your secure systems, where it can be retained longer than you expect, reached by legal discovery if the company is sued, exposed in a breach, or read by people at the vendor. And a setting is a vendor's assurance, not your agency's authorization. Approval by your agency for that exact category of data is what makes a tool usable, not a toggle in its interface.

What if my agency has no approved tool and I have a real deadline? Then the work gets done without AI on that data, or you use AI on a sanitized version with every identifier and every distinctive detail removed. Urgency is not authority: "we need this output today" does not override a security requirement, and the deadline you miss is recoverable in a way that a disclosed shelter address is not. Raise the gap with your AI or privacy point of contact, because a missing approved tool is a problem your agency should be told it has.

Is aggregated or synthetic data always safe to use? Safer, not safe. Aggregation can be reversed when the groups are small, and synthetic data can sometimes be reverse-engineered or linked back to the real records it was derived from. Both are strong risk reductions and neither is a licence to skip the approval question. Treat them as ways to lower the category of data you are handling, then handle it at that lower category properly.

Someone on my team already pasted a case file last month. What now? Report it, even late. Capture what was sent and when, do not delete the account or the history, and tell your supervisor and your privacy or security officer. Reporting expectations at most agencies are measured in hours, so a month-old paste is already overdue, which is an argument for reporting it now rather than for continuing not to. Agencies are judged on how they respond, and a late report still beats a discovered one.