←
AI for Government
Aware · M8 · lesson 8 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Data Sensitivity and Classification

10 min

Marcus Hill, a budget analyst at a federal agency, had a clever idea. He wanted to use an AI tool to spot patterns in a spreadsheet of contractor payments. Harmless numbers, he figured. Before pasting, he asked a colleague, who asked one question that stopped him cold: "Does that spreadsheet have vendor tax IDs in it?" It did. Tax identification numbers are sensitive personal data. What Marcus thought was "just budget numbers" was actually a mix of public spending figures and protected information, and he had nearly fed it all into a tool he had not checked. He was not reckless. He just had not learned to look at his own data and see what category it falls into.

That skill, looking at any piece of data and knowing how sensitive it is, is the foundation of using AI safely in government. You do not need a security clearance or a law degree. You need to recognize a handful of data categories and understand how AI changes the rules for handling each. This lesson gives you both, with a checklist you can run in seconds. It reads like a dry, bureaucratic topic and it is not: data sensitivity is the first line of defense against harm, and no amount of AI policy will save you from a breach if you do not know what data you are holding.

Why government sorts data into buckets

Every piece of government data carries a level of harm-if-exposed. A published press release: zero. A list of citizens' Social Security numbers: severe. Classification is simply the system for sorting data by that harm so everyone knows how carefully to handle it. The label travels with the data and tells you who may see it, where it may go, and, now, whether it can touch an AI tool. Marcus's mistake was assuming a spreadsheet had one uniform sensitivity. In reality a single file can hold several categories at once, and the most sensitive one sets the rules for the whole file.

What is actually at stake

Classification is foundational to government AI governance for four separate reasons, and it helps to keep all four in view, because people who dismiss one usually have not thought about the others.

  • Legal requirements. Agencies are subject to laws and regulations about data handling, and different types of data require different levels of protection. Using AI with data you are not supposed to use that way can violate law.
  • Operational risk. A breach involving sensitive personal information can cripple an agency's operations, damage public trust and produce legal liability.
  • Citizen privacy. Government collects information from citizens under an implicit promise that it will be handled responsibly and securely. Using that data for purposes citizens do not expect, or failing to protect it, violates that trust.
  • AI-specific risks. AI systems change how data is used and combined. A dataset that was safe to share in one context can become dangerous in a system where it is joined to other datasets and processed in new ways.

The categories you must recognize

Learn these five labels well enough to spot them on sight. The acronyms are everywhere in government; here they are in plain English.

  • Public / Unclassified. Information meant for, or safe for, public release: published reports, press releases, open datasets. Lowest sensitivity, but check that "unclassified" does not actually mean "controlled."
  • PII, Personally Identifiable Information. Anything that identifies a specific person: name, Social Security number, address, date of birth, tax ID. Marcus's contractor tax IDs were PII hiding in a budget file.
  • PHI, Protected Health Information. Health and medical details tied to a person, such as diagnoses, treatments and records. Protected by health-privacy law on top of normal privacy rules.
  • CUI, Controlled Unclassified Information. The big, easy-to-miss category. Not classified, but still requiring protection: law-enforcement-sensitive material, certain procurement data, infrastructure details. "Unclassified" does not mean "free to share." CUI is the trap most people fall into, and the most dangerous label is always the one that sounds safe.
  • Classified. Information whose exposure would damage national security, in tiers (Confidential, Secret, Top Secret). The brightest line of all. Classified data never touches a public AI tool, never, under any circumstance.

How each category is actually handled

Recognition is step one. Step two is knowing what each label demands of you, and specifically what it demands once an AI system enters the picture. Work through them in order of rising risk.

Unclassified and public data

This is information that poses no known risk if released: weather data, published economic statistics, agency press releases, general programme information such as eligibility requirements and how to apply, publicly available organizational charts. The handling requirement is low risk, with no special access controls, and it can typically be shared with external partners and used in public-facing systems. The AI wrinkle is combination. Neighborhood median income is public. Home addresses may be publicly available. A system that joins them to predict property values can raise privacy concerns that neither dataset raised alone, so ask not only whether each input is public but what emerges when they meet.

Controlled Unclassified Information

CUI is not classified but needs protection for reasons other than national security. The United States maintains official CUI markings and handling procedures, and other countries run equivalent categories under names such as Official Sensitive or Internal Only. Typical examples: certain law enforcement records, pre-decisional deliberations, proprietary business information, personal information that does not rise to PII status, and internal communications that could cause operational disruption if released.

CUI often cannot be used in an AI system without explicit safeguards: anonymize or de-identify the data, limit access to the system itself, document who has access and why, and treat the system's outputs as CUI too. Handling sits at moderate to high risk, with access limited to authorized personnel, storage in secure systems, encryption during transmission, no external sharing without authorization, and clear records of who accessed what and when.

Personally Identifiable Information

PII is information about an individual that can be used to identify that person, and it is one of the most commonly regulated categories in government AI work. It runs wider than most people assume: name, Social Security number, date of birth, driver's license number, passport number, email address, home address, phone number, biometric data such as fingerprints and iris scans and facial recognition data, financial account numbers, medical information, criminal history and employment records. PII is often the most valuable data for an AI system precisely because it is specific, personal and predictive, and that is exactly why it is the most tightly controlled.

Using PII in an AI system means having explicit legal authority to use it for that purpose, notifying individuals that their data is being used this way (usually), implementing strong security measures, holding a plan for data retention and deletion, testing whether the system's outputs can reveal the original PII, and considering whether the person has to consent to the use. Handling is high risk. PII should never be used in unapproved AI systems, should be encrypted at rest and in transit, and access should be strictly limited to essential personnel. Never, ever upload PII to a public, consumer-facing cloud AI service unless you have explicit authorization and a security review.

Protected Health Information

PHI is a subset of PII: medical and health information about individuals, subject in many countries to specific privacy laws, HIPAA in the United States being the obvious example. It covers medical diagnoses, treatment history, prescription information, mental health records, genetic information, health insurance information and immunization records. If you work in health, veterans benefits or a similar agency, you need to understand PHI handling thoroughly rather than approximately, because it is among the most sensitive data government holds.

Using PHI in an AI system requires explicit legal authority under health privacy laws, security requirements often higher than for other PII, individual consent in many contexts, audit trails documenting every access, and strong de-identification if you are trying to reduce risk, done carefully enough that the data cannot be re-identified. Handling is the highest civilian risk tier. PHI should be used in AI systems only when absolutely necessary and only with explicit legal authority and individual consent. Encryption is mandatory. Audit logging is mandatory. Access is restricted to the minimum necessary personnel, and PHI is never shared with external partners without ironclad legal agreements.

Classified information

Classified material relates to national security and is marked at levels such as Top Secret, Secret and Confidential in the United States, with equivalents elsewhere. It covers intelligence information, military capabilities and tactics, diplomatic cables, nuclear weapons information and some cybersecurity information. Classified information should generally not be used in AI systems at all, and certainly not in commercial or general cloud systems. If you work in an agency that handles classified material and you are thinking about AI, consult your security and legal teams first.

The risk is extreme: classified information inside an AI system could compromise national security. It can only be processed on classified systems with multiple layers of security, and it should not be submitted to any AI system not specifically approved for classified work and physically secured at the appropriate level.

How AI changes the handling rules

You already follow handling rules for sensitive data. You do not email Social Security numbers to strangers. AI introduces three new wrinkles that old habits do not cover, and each one breaks a different assumption you have been relying on without noticing.

The data leaves your control. Typing into most AI tools sends your text to a company's servers. For PII, PHI, CUI or anything classified, that is the same as handing it to an outside party. The classification rules that say "keep this inside the agency" are violated the moment you press send on an unapproved tool, and no intent on your part changes that.

The data may become permanent. Many consumer AI tools learn from what users type. Sensitive data fed in could be absorbed into the tool and resurface later. You cannot reclassify it, you cannot recall it, and you generally cannot verify its deletion once it is gone.

Mixing creates new sensitivity. AI is good at combining scattered facts. Two pieces of low-sensitivity data, joined, can identify a person, turning public into PII. A model can connect dots you never intended to connect, which means the sensitivity of an output is not capped by the sensitivity of any single input.

A four-step classification check before any AI use

Run this on any data before it goes near an AI tool. Marcus uses it now on every file, and it takes under a minute once it is habit.

StepAskAction
1. Look for markingsIs the file or system labeled (CUI, PII, classified)?If marked sensitive, stop and follow that label's rules.
2. Scan the contentsDoes it contain names, IDs, health, financial or controlled details, even buried in a column?If yes, treat the whole file at that sensitivity.
3. Find the highest categoryWhat is the single most sensitive item in here?That item sets the rule for the entire file.
4. Match to an approved toolIs there an agency-approved AI tool cleared for this category?Public data: an approved tool is generally fine, but consider what combination could reveal. Sensitive: only a cleared tool, or de-identify first. Classified: never a public tool.

De-identification and its limits

One strategy for handling sensitive data in AI systems is to remove the information that makes it personally identifiable. This is called de-identification, or anonymization when it goes further. The idea is that once you strip names, addresses and other direct identifiers, what remains is data about characteristics, behaviors or outcomes rather than about people. That idea is sound and it is also where most privacy failures begin, because de-identification is not a switch. It is a process, and it can fail.

True anonymization requires three things: that direct identifiers such as name and Social Security number are removed, that the data cannot be re-identified by combining it with other data, and that no residual risk of re-identification remains through pattern matching or linkage. The third requirement is the hard one. In an era of large datasets and capable models, re-identification is increasingly possible. If your de-identified dataset includes age, gender, zip code and medical condition, researchers can often re-identify individuals by linking it to other public datasets carrying the same combination of characteristics.

So treat de-identification with caution in government. Do not assume de-identified data is safe to use in any AI system. Have a data scientist or security expert review your approach. Test whether re-identification is possible by actually trying to link your data to other known datasets. Document the process and the assumptions it rests on. And plan for the possibility that it fails, which means knowing your fallback before you need it rather than during an incident.

Return to Marcus with that framing. He still wanted his pattern analysis, and classification does not mean you cannot use AI, only that you handle the data correctly first. He removed the tax IDs and vendor names, replacing them with neutral labels (Vendor A, Vendor B), and ran his analysis on what was left. The patterns he needed lived in the amounts, not the identities.

What that buys is real and it is also bounded: he kept direct identifiers out of the tool. It is not proof that nobody could ever work out which vendor is which, because a small contract set with distinctive amounts can still point back at a single firm. Removing the names is a control, not a guarantee, and the difference between those two words is where agencies get into trouble.

That habit still aligns with the responsible-AI expectations agencies operate under, which start from a simple principle: protect people's information by default and relax that protection only when you have confirmed it is safe to do so. Classification is how you make that judgment quickly and consistently instead of guessing under deadline pressure.

A worked scenario

Imagine your agency wants to use AI to improve service delivery by predicting which eligible citizens are at risk of not applying for benefits they qualify for, so outreach can reach them. This is a good use of AI in principle; it helps people get services they are entitled to. Now ask what data it needs: demographic information such as age, location and family size, employment information, previous benefit application history, perhaps economic indicators for the area. Most of that is sensitive, either PII or CUI.

Work it through in order. What data do you need? Employment history, family size, address, all PII. Do you have legal authority to use it for this purpose? Check your agency's organic statute and relevant regulations. Using benefit application data to predict unmet need is usually within the mission; using it to train an AI is a separate question you have to check, and you may need to notify people first. How will you handle it? Store it encrypted, limit access to the specific team building the model, keep it out of public cloud, audit who accesses it, hold a retention policy that deletes it when the project ends, and test the model to confirm it does not reveal PII.

Then handle the outputs with the same care as the inputs, which is the step teams skip. The model produces a list of people to contact, and that list itself reveals something about them, namely a suspected need for benefits. Treat it as you would any sensitive benefit information. And when you want to test, use test data rather than real records with real PII, or work with real data inside a controlled, secure environment, never in a public AI system. That is what operationalizing data sensitivity actually looks like.

Anti-Patterns to Avoid

These are the specific ways organizations mishandle classification once AI enters the workflow.

  • Not classifying data at all. A team starts building and collects whatever it needs without explicitly thinking about sensitivity levels, documenting neither what data is in use nor how sensitive it is. They end up using data they are not legally authorized to use, or mishandling data that is sensitive, and when it is discovered they face liability and have to shut the system down.
  • Treating "we removed the names" as anonymization. An organization strips names from a dataset and assumes it is now safe to use anywhere, not realizing that age, gender, zip code and medical condition together can re-identify individuals in many cases. The data is less anonymized than they believe, individuals can be re-identified, and the privacy promise made to those people is broken without anyone noticing they broke it.
  • Uploading PII to unauthorized systems. Someone needs to test an AI approach, copies a dataset containing PII and uploads it to a consumer cloud AI service to experiment. That violates security and privacy policy outright. The data is now outside government control, where it can be accessed by others, used to train other models, or leaked, and the agency faces liability, reputational damage and the obligation to notify affected individuals.
  • Not reclassifying when AI combines datasets. Two datasets are each classified low-sensitivity, then combined in an AI system that produces insights more sensitive than either alone. The team keeps treating the output as low-sensitivity and shares it with people who should not see it.
  • Reading "unclassified" as "public." CUI is unclassified and still protected. The word on the marking describes national security status, not shareability, and the gap between those two meanings is where most well-intentioned disclosures happen.

Practice Prompts

Do these against real data you touch, not a hypothetical.

  • Classify three datasets. Write down three datasets your agency uses, the sensitivity level of each (unclassified or public, CUI, PII, PHI, classified), your reasoning, and what you would have to do differently to use each in an AI system. Flag any you are unsure about for follow-up with your security or privacy office.
  • Find the hidden column. Pick one dataset and identify whether any part of it is more sensitive than the rest, the way tax IDs sat inside a budget spreadsheet. Then confirm what that most sensitive item does to the handling rules for the whole file.
  • Write the safety conditions. If your agency wanted to use that data in an AI system, what would have to be true for it to be safe? Name the legal authority required and the specific security measures.
  • Test a de-identification. Take a de-identified extract and try to link it to a public dataset using quasi-identifiers. Document what you found, including finding nothing, and what assumptions your method depends on.
  • Answer the question you will be asked. If a colleague asked "is this data safe to use in an AI system," write the answer you would give today, then note which parts you could not support with a policy or a marking.

Reflection

These are worth sitting with, because the wrong answer to any of them is usually discovered in an incident report.

  • What data do I handle in my current role, and are the classification and handling procedures for it actually clear to me?
  • Have I ever been surprised by how sensitive data became when combined with other information? What did that teach me about combination?
  • Where in my work would deadline pressure most tempt me to skip the classification check, and what would I do instead?
  • If a de-identification I relied on turned out to be reversible, who would need to know, and how quickly?
  • Who in my agency do I ask when I genuinely cannot tell, and do I know how to reach them today rather than next week?

Glossary

  • Data classification. The process of assigning sensitivity levels to data based on the harm that would result from unauthorized disclosure.
  • CUI (Controlled Unclassified Information). Unclassified information that requires protection for reasons other than national security, such as law enforcement sensitivity or pre-decisional status.
  • PII (Personally Identifiable Information). Information that can be used to identify an individual, including name, Social Security number and address.
  • PHI (Protected Health Information). Medical or health-related information about an individual, subject to special privacy protections.
  • De-identification. The process of removing direct identifiers from data so that it is no longer personally identifiable.
  • Anonymization. The process of removing all information that could identify an individual, either directly or through combination with other data.
  • Re-identification. Determining the identity of individuals in supposedly de-identified data through linkage to other datasets or pattern matching.

Classification is the entry point to the responsible-use chapter, and each of these takes one part of it further.

Closing

Data sensitivity looks like a compliance checkbox, something to worry about once and then forget. It is closer to the foundation. If you cannot keep data secure and handle it appropriately, no amount of algorithmic sophistication matters, because the system is built on sand. Classification is the fast, repeatable judgment that keeps that foundation solid under deadline pressure.

So keep it front and center. Classify your data. Handle it according to its highest category. Do not cut corners to move faster, because a breach or a privacy violation costs far more time and credibility than the weeks spent doing governance properly. And when you are genuinely in doubt about whether data is safe to use in an AI system, ask your security team, privacy office or legal counsel. Asking before you cause a problem is always cheaper than finding out afterward.

Key Takeaways

  • Classification sorts data by harm-if-exposed. The label travels with the data and tells you who may see it, where it may go, and whether it can touch an AI tool.
  • Learn five categories on sight. Public, PII, PHI, CUI and Classified, each with stricter handling than the last, and each with its own AI-specific conditions.
  • CUI is the trap. "Unclassified" does not mean "public"; Controlled Unclassified Information still requires protection and trips up most people.
  • A file's most sensitive item sets the rule. Sensitive data hides in single columns, like tax IDs in a budget spreadsheet, and governs the whole file.
  • AI changes handling three ways. Data leaves your control, may become permanent through training, and can gain new sensitivity when a system combines facts.
  • PII and PHI never go to unapproved cloud AI systems. This is not a gray area, and PHI carries the additional legal-authority, consent and audit requirements of health privacy law.
  • Classified data never touches a public AI tool. No exceptions, no deadlines, no clever workarounds.
  • De-identification is a process that can fail. Removing names is a control, not a guarantee; quasi-identifiers such as age, gender, zip code and condition can re-identify people through linkage.
  • Document your classifications and your reasoning. If something goes wrong later, that documentation is your protection.

Frequently Asked Questions

My file is marked unclassified. Can I paste it into an AI tool? Not on that basis alone. "Unclassified" describes national security status, not shareability, and Controlled Unclassified Information is unclassified by definition while still requiring protection. Run the four-step check instead: look for markings, scan the contents for names, identifiers, health, financial or controlled details, find the single most sensitive item, and match that to a tool your agency has cleared for that category.

If I remove the names, is the data safe? Removing direct identifiers is a real and useful control, and it is not the same as anonymization. Anonymization also requires that the data cannot be re-identified by combining it with other data and that no residual linkage risk remains, which is the hard part. Age, gender, zip code and medical condition together are often enough to pick individuals out by linking to public datasets. Have someone review the method, test the linkage yourself, and document what your approach assumes.

Two datasets are both low-sensitivity. Is combining them low-sensitivity? No, and this is one of the four common anti-patterns. Combination can produce insights more sensitive than either input, and the output then needs its own classification rather than inheriting the lower one. Neighborhood income and home addresses are both ordinary on their own; joined and modeled, they can raise privacy concerns neither raised alone.

What is different about PHI compared with other PII? PHI is a subset of PII with an extra legal layer on top, health privacy law such as HIPAA in the United States. In practice that means explicit legal authority for the use, security requirements often stricter than for other PII, individual consent in many contexts, mandatory encryption, mandatory audit logging of every access, minimum-necessary access, and no sharing with external partners without ironclad legal agreements.

I am not sure which category applies. What do I do? Stop before the data goes anywhere and ask your security team, privacy office or legal counsel. Uncertainty is a reason to pause, not a reason to guess, and treating the file at the highest plausible category while you wait costs you a delay rather than an incident. Flag it for follow-up and write down what made it ambiguous, because that ambiguity is usually a gap in your agency's markings that someone should fix.