AI in Government Today
Diane Whitfield has worked the front counter at a county benefits office for nineteen years. She is not a programmer and never wanted to be. But last month her supervisor mentioned that the agency was "looking at AI," and a citizen at her window asked whether a robot was going to decide his food-assistance case. Diane did not know the answer. She does now, and she sleeps fine, because she learned the one thing that matters for someone in her seat: AI in government is already here, it is mostly doing boring helpful things, and a human still owns every decision that affects a person's life.
This lesson is for everyone in public service, at any level. You do not need to build anything. You need to recognize what AI actually is when you encounter it at work, where it already lives across government, and what your one job is when it shows up at your desk. Think of AI here as a very fast, very literal new clerk: tireless at sorting and drafting, terrible at judgment, and never, ever the person who signs the decision.
What "AI" actually means at work
Forget the movies. The AI arriving in government is mostly software that does one of three things. It finds patterns in large piles of data, like flagging which tax returns look unusual. It understands and produces language, like answering a routine question or drafting a first version of a letter. And it predicts, like estimating which bridge needs inspection first.
That is it. It is a tool that is fast at narrow tasks and confidently wrong on a regular basis. It does not understand the way Diane understands a frightened applicant. It pattern-matches. Holding onto that one distinction, pattern-matching versus understanding, is most of what any public servant needs to carry into a meeting where somebody is selling a system.
Why the landscape is worth studying
Government AI is not hypothetical anymore. It is deployed, it is affecting citizens, and understanding what is actually happening is essential to your credibility and your agency's success. The questions worth answering are practical ones. What are peers doing? What is proven? What is risky? Where are the lessons already learned, at somebody else's expense, so you do not have to pay for them again?
Understanding the current state also grounds abstract concepts in reality. When your agency proposes an AI initiative, you will recognize patterns from deployments that worked elsewhere and avoid mistakes others have already made. Some of what follows is brilliant. Some of it is a cautionary tale that cost real people their benefits, their jobs, or their liberty. All of it is instructive, and the failures teach harder lessons than the successes do.
Where AI already lives across the federal government
This is not coming. It is running today. A tour of real federal use makes it concrete, and the pattern repeats: the system speeds up sorting and flagging, and a human keeps the determination.
Veterans Affairs. The VA processes benefit claims, manages medical records, and coordinates care for more than 9 million veterans. AI classifies incoming documents such as discharge papers, medical records, and benefit requests so they route to the right desk. It extracts key data such as service dates, disability ratings, and medical conditions from unstructured documents for entry. It also helps optimize appointment scheduling to reduce wait times. These are low-risk assistance tasks: the AI speeds the work without making the final call, and veterans keep human review and appeal. The VA had to address privacy concerns around sensitive veteran health data, so it works with encrypted systems and only VA-approved deployments.
Internal Revenue Service. The IRS receives hundreds of millions of tax documents annually. Risk scoring flags returns for human review based on patterns such as inconsistencies and unusual deductions. Optical character recognition and classification process documents at scale. Matching identifies mismatches between the income employers report and what taxpayers file. The upside is that it targets scarce examiner time where it is most needed, and humans make every final decision. The live challenge is fairness: ensuring the system does not disproportionately flag particular taxpayer demographics. The IRS has had to audit its own system for exactly that.
Postal Service. Post offices process more than 470 million mail pieces daily. Optical character recognition reads handwritten and printed addresses, sorting routes mail to regions, and package identification separates letters from parcels at scale. This is the ideal AI task: clear, repetitive, enormous volume, low stakes on any single item. The honest limitation is that handwriting varies widely and the systems perform worse on certain handwriting styles, which the Postal Service continues to refine.
Environmental Protection Agency. Environmental agencies monitor air quality, water pollution, and land use. Satellite image analysis detects illegal dumping, coastal erosion, and deforestation. Sensor data analysis finds patterns in pollution levels. Report automation extracts key information from environmental filings. This enables monitoring at a scale humans could not achieve and gives scientists more time for analysis. The caveat is interpretive: satellite imagery varies seasonally and by weather, so the systems need context to read correctly.
Defense and security agencies. Military and intelligence agencies use AI for threat detection, pattern analysis, and logistics. Details are limited for security reasons, but public information describes pattern recognition in communications and financial data, logistics optimization for supply chains, and image analysis for surveillance. Governance here is heavy by design: restrictions on use, limits to specific scenarios, extensive human oversight, and regular audits for bias. Humans remain in command of any consequential action.
Social Security Administration. The SSA receives millions of retirement, disability, and survivor benefit applications. Eligibility screening flags applications that clearly meet or clearly fail criteria so they can move faster. Document processing extracts information from applications. Fraud detection identifies suspicious patterns. Clear cases move quickly and humans review the borderline ones. The challenge the SSA had to work through was making sure the system did not disadvantage people with disabilities or particular demographic groups.
State and local: where the hardest lessons live
State and local government is moving too, and often faster than federal. Cities use AI chatbots to answer routine questions about permits at two in the morning so a resident does not wait until Monday. Transit agencies predict where buses will bunch up. Diane's own county is "looking at AI" for exactly this kind of routine-question relief, not to decide cases. But the local record also contains the campaign's clearest warnings, and they are worth knowing by name.
Michigan unemployment fraud. Michigan deployed an AI system to detect unemployment insurance fraud. It flagged 40,000 people as fraud perpetrators. Most were innocent. The system had mathematical flaws. Thousands lost benefits they were entitled to, lawsuits followed, and settlements cost the state millions. The lesson is blunt: do not deploy AI to high-stakes decisions without extensive testing, and do not ignore the early signs that something is wrong.
California benefits automation. California built an automated benefits system that required citizens to report changes online. The system was poorly designed. Thousands of eligible people lost benefits through technical failures and confusing instructions, and litigation and negative publicity followed. The lesson is that automation must account for human diversity. Not everyone is comfortable with online systems, and an easy appeals process is not a nicety, it is essential.
Boston police hiring. Boston considered using an AI system to predict which police candidates would be disciplined during their careers. The data showed bias: candidates from certain demographic backgrounds had higher discipline rates, likely because of biased policing practices rather than because they were worse police officers. The city rejected the system. The lesson is to audit whether your training data contains the very biases you would be automating, before you deploy, not after.
New York predictive policing. New York explored using AI to predict where crime would occur and deploy officers there. The practice amplified over-policing in certain neighborhoods and created a feedback loop: more police presence produced more arrests, which fed back into training data, which produced more predictions and more policing. The practice was discontinued. Prediction systems can amplify existing inequities if nobody is watching for it.
Chicago risk assessment. Chicago used a risk assessment tool to identify individuals likely to be involved in violence. The tool was biased against African Americans. People on the list faced increased police attention and harassment. The city eventually recognized the bias. High-stakes decisions need fairness audits and human oversight, and recognizing the problem years later is not the same as preventing it.
Two lessons from outside the United States
The UK Post Office deployed a system to detect branch accounting errors. The system was flawed, but post office workers were blamed for the discrepancies it produced. Hundreds were prosecuted and some were imprisoned. The errors were eventually attributed to the system, and the Post Office had to overturn convictions and pay settlements. The lesson belongs on a wall in every agency: when a system flags a problem, do not assume the human is the cause. Investigate the system itself.
The European Union is rolling out the AI Act, which classifies AI systems by risk level and requires different governance for each. The structure it sets out is worth knowing because regulatory frameworks are coming, and governance that works tends to become standard practice whether or not a statute forces it.
| Risk level as the AI Act describes it | What it covers | What is required |
|---|---|---|
| High risk | Systems affecting fundamental rights, including criminal justice, benefits, and hiring | Extensive documentation, testing, and human oversight |
| Medium risk | Systems such as chatbots | Transparency disclosures |
| Low risk | Most other systems | The baseline |
The line that protects you and the public
Here is the rule that makes Diane sleep fine, and it is the most important sentence in this lesson. For any decision that affects a person's rights, benefits, or safety, a human must remain accountable. The model can suggest. It cannot decide alone.
This is not just good ethics; it is becoming policy. Federal guidance issued through the Office of Management and Budget draws a bright line around what it calls "rights-impacting" AI, meaning anything affecting someone's access to benefits, jobs, housing, or due process. For those uses, agencies must keep a human in the loop, let people appeal, and offer a non-AI path. So when the citizen at Diane's window asks whether a robot is deciding his food-assistance case, the honest, accurate answer is: no, a person decides, and you can appeal to a person.
Notice that this line is also the thing every failure above crossed. Michigan removed the human from a determination. The Post Office treated a system's output as evidence against a person. Chicago let a score follow individuals into real police contact. In each case the technology was a contributing cause and the missing human accountability was the decisive one.
What separates a help from a harm
Read enough deployments and the sorting rule becomes obvious. Successful ones share a short list of characteristics: human oversight, a working appeals process, ongoing fairness monitoring, and clear governance about who is responsible for what. Failures share their own list just as consistently: insufficient testing before launch, ignoring bias when it first appears, removing human judgment from a decision that needed it, and no oversight after go-live. Those two lists are the practical takeaway from the whole landscape.
A simple two-column map keeps the same idea at desk level.
| AI is genuinely good at | AI is unreliable or dangerous at |
|---|---|
| Sorting huge volumes fast (mail, forms, alerts) | Understanding context, tone, or a person's situation |
| Drafting a first version of routine text | Being correct without a human checking the facts |
| Spotting unusual patterns for a human to review | Making fair decisions when its training data was biased |
| Answering common, well-defined questions | Handling the unusual case that does not fit the pattern |
| Translating and transcribing | Owning any decision about a person's rights or benefits |
Notice that AI's weaknesses are precisely the things public servants are trained for: judgment, fairness, and handling the case that does not fit the form. That is not a coincidence. It is the reassurance. AI does not replace the public servant. It clears the clerical pile so the public servant can do the human part, which is the part that was always the job.
Your job when AI shows up at your desk
You will not be asked to build a model. You will be asked to use a tool, or to answer a citizen about one. Five simple habits cover almost every situation.
- Treat every AI output as a draft. Whether it is a summary, a suggested answer, or a flagged case, check it before you act. It is a fast first guess, not a final word.
- Never paste private citizen data into a public AI tool. Names, Social Security numbers, and case details go only into systems your agency has approved.
- Know who owns the decision. If a tool is influencing a person's benefits or rights, confirm a human signs off and can be named.
- Tell the truth to citizens. If asked, you can say a tool helped sort or draft, but a person decides and they can appeal to a person.
- Speak up when something looks wrong. If an AI tool flags a case that obviously does not fit, your judgment is the safeguard. Use it and report it.
That is the whole job for most public servants. Recognize the tool, treat it as a draft, protect citizen data, and remember that the human, you, still owns the outcome. Every agency on the harm list above had someone in Diane's seat who noticed something was off. What separated the recoverable failures from the catastrophic ones was whether anyone listened.
Anti-patterns
Three failure shapes account for most of the harm in the record above. They are worth naming because each one looked reasonable to the people who chose it.
- Rolling out without testing on real-world data. Systems work beautifully in testing and fail in production because real-world data is messier than test data. Always test on data that matches production conditions before anyone's case depends on the result.
- Ignoring early signs of bias. Newly deployed systems often show unequal performance across groups in their first weeks. Monitor from day one. If bias appears, fix it before it harms thousands, which is the number Michigan reached.
- Removing human review from high-stakes decisions. Automating for efficiency and quietly deleting the human judgment that should have stayed. AI should assist humans in high-stakes decisions, not replace them.
- Treating a system's output as evidence about a person. The Post Office prosecuted people on the strength of a flawed system's arithmetic. A flag is a reason to look, never a finding of fact about someone's conduct.
Practice prompts
- Landscape mapping. Which federal agencies does your agency interact with? Research their AI deployments. How might those systems affect your work or your applicants?
- Local deployment. Does your state or city use AI in any capacity? Research it. What are citizens saying about it?
- Learn from failures. Pick one government AI failure from this lesson, such as the Michigan unemployment system or the Post Office accounting system. What went wrong, and what specific control would have caught it?
- Peer learning. Find an agency similar to yours that is using AI. How did they approach governance, and what can you borrow?
- Readiness assessment. If your agency deployed AI tomorrow, what would you most want to know from other agencies' experiences first?
Reflection
Find one AI deployment in government, at any level, that touches you or your agency. Then answer four questions about it in writing: What does it actually do? Who reviewed it before it went live? Has anyone complained, and what happened to those complaints? What would you do differently if it were yours to run?
Then ask the question Diane had to answer at her window. If a citizen asked you today whether a machine was deciding their case, could you give a true answer without checking with anyone? If not, that gap is the first thing to close, and closing it usually takes one conversation with the person who owns the system.
Glossary
- Production deployment. The point at which an AI system moves from testing into real-world use affecting real people or processes.
- Feedback loop. When a system's predictions influence the data used to train future versions, potentially amplifying biases, as in predictive policing.
- Fairness audit. Testing an AI system to see whether it performs differently for different demographic groups.
- High-stakes decisions. Decisions that significantly affect people's lives, rights, or welfare, such as benefits, criminal justice, and hiring for sensitive positions.
- Real-world data. The messy, diverse data a system encounters after deployment, often meaningfully different from controlled test data.
- Equity. Fair treatment and opportunity. AI systems should not systematically disadvantage groups.
- Governance. The rules, processes, and oversight mechanisms guiding how an AI system is developed, deployed, and monitored.
Related lessons
- What AI Is and Is Not draws the boundary this lesson assumes between what these systems are and what the marketing claims.
- How AI Actually Works explains the mechanism underneath every deployment described here, including why biased data produces biased output.
- Types of AI Systems gives you the vocabulary to say which kind of system an agency is actually proposing.
- What AI Does Well and Where It Fails turns the two-column map above into a test you can run before starting a project.
- When Government AI Goes Wrong takes the cautionary tales further and covers what to do when you are the one who finds the problem.
Closing
You have now seen AI across government: where it works, where it has failed, and the patterns that make the difference. The landscape is your teacher. Learn from other agencies' successes and from their mistakes, which are usually documented in more useful detail because somebody had to explain them.
Your agency is part of this story. When you evaluate an AI tool, you are not only solving your own problem; you are contributing to whether government AI earns public trust or spends it. The deployments in this lesson are affecting people today. Some are working well. Others have harmed people who did nothing wrong. Your responsibility is to apply the lessons to your own context and help your agency do better than the worst example on the list.
Key takeaways
- AI is already here, doing boring useful work. Across the VA, IRS, USPS, DoD, EPA, and SSA it sorts, drafts, extracts, and flags, mostly assisting humans rather than making decisions.
- It pattern-matches; it does not understand. AI is fast at narrow tasks and confidently wrong on a regular basis, which is why it needs a human check.
- A human owns every decision that affects a person. For benefits, rights, and safety, the model may suggest but a named human decides and can be appealed to.
- High-stakes deployments have caused real harm. Michigan's unemployment system and Chicago's risk tool damaged people's lives when testing and fairness monitoring were inadequate.
- Successes and failures each have a signature. Successes share human oversight, appeals, fairness monitoring, and clear governance; failures share thin testing, ignored bias, deleted human judgment, and no post-launch oversight.
- Regulation is arriving. The EU AI Act sorts systems by risk level and attaches heavier duties to systems that touch fundamental rights, and governance that works becomes standard practice regardless.
- Your judgment is the safeguard. When a tool flags something that obviously does not fit, speaking up is part of the job, not an interruption to it.
Frequently Asked Questions
Is AI deciding citizens' benefit cases today?
Not where the rules are followed. Across the federal deployments in this lesson, AI classifies documents, extracts data, screens for clear cases, and flags anomalies, while a human examiner makes the determination. Federal guidance treats systems affecting benefits, jobs, housing, or due process as rights-impacting, which requires a human in the loop, an appeals path, and a non-AI alternative. The failures in this lesson are largely cases where that separation broke down.
If the technology is the same, why did some agencies succeed and others cause harm?
Because the difference was rarely the technology. Successful deployments tested on data matching real conditions, monitored fairness after launch, kept a human accountable for consequential calls, and gave people a way to appeal. Harmful ones skipped testing, ignored early bias signals, or automated a judgment call. The same classifier can be a helpful triage tool or a harm engine depending entirely on what surrounds it.
What should I do if a tool flags a case that looks obviously wrong to me?
Act on your judgment and report it through whatever channel your agency has. A flag is a prompt to look, not a finding about a person. The Post Office case is the cautionary extreme: staff were blamed for discrepancies a flawed system produced, and hundreds were prosecuted before anyone investigated the system itself. When output and reality disagree, the system is a suspect too.
Does the EU AI Act apply to my agency?
This lesson does not make that determination, and you should not assume either way from a training course. What is worth taking from it is the structure: it sorts systems by risk level, and systems touching fundamental rights such as criminal justice, benefits, and hiring carry the heaviest documentation, testing, and oversight duties. That risk-tiered logic is becoming general practice, so building to it is rarely wasted effort.
I do not use AI at all. Why does this matter to me?
Because you will be asked about it by a citizen, a colleague, or a supervisor, and because a system you never touch may still route, score, or flag the work that reaches your desk. Knowing what the tool did and did not decide is what lets you answer honestly, and knowing that your own judgment is a designed safeguard is what lets you use it without hesitation.
Skill.re