What AI Does Well and Where It Fails
A grants analyst named Priya Nair was three weeks behind on summarizing 600 public comments for a new environmental rule. A colleague showed her a generative AI tool. She fed it the comments, and in twenty minutes she had a clean, organized summary by theme. It was genuinely good work, the kind that would have taken her two days. Thrilled, she used the same tool that afternoon to answer a citizen's question about a specific deadline in the regulation. The AI gave her a confident, precise date. She put it in an email to the public. The date was wrong. The AI had invented it. The same tool that saved her two days in the morning created a correction notice and an awkward apology by 4 p.m.
Priya did not use a bad tool. She used a good tool for the wrong job. The single most useful skill in working with AI is knowing the line between what it does brilliantly and what it does dangerously. This lesson draws that line, works through the tasks that sit awkwardly on it, and gives you a test you can apply before you trust any AI output.
Why this is the skill that matters
Most AI failures in government are not failures of the technology. They are failures of fit: someone pointed a capable tool at a task it was never good at. Get the fit right and AI saves real time on real work. Get it wrong and it produces confident, polished, plausible mistakes, which are far more dangerous than obvious ones because they slip past a busy reviewer.
This is where well-intentioned initiatives die. An agency adopts AI for the wrong problem, the system performs technically, and it does not solve the actual challenge, because the task was never suited to it. Or worse, it automates something that should never have been automated. Government has seen this shape before with blockchain for government and connected sensors for everything. AI is the same story with better press.
Underneath it, government carries a duty to citizens to spend resources wisely, and that duty starts with a question asked before the procurement rather than after: is AI actually the right tool here? The reason a generative model both saves and betrays comes down to what it does. It predicts likely words. It is a brilliant pattern-completer, not a fact-checker and not a reasoner. AI is fluent before it is correct, and it will give you a confident answer whether or not it has one.
What AI does well
AI shines when the task is about transforming or organizing information you already have, where fluency matters and you can check the result. The capabilities below cover most of the real wins in government work.
Pattern recognition at scale. AI's core advantage is finding patterns in datasets too large to analyze by hand: thousands of documents, millions of transactions, terabytes of sensor data. The source figures for this comparison are worth reading carefully. A human auditor could review 100 financial documents per day, while an AI system reviews 100,000. A human might catch irregularities in 5 of them; an AI system might catch 50. That comparison is about absolute volume surfaced, not about per-document accuracy, and the two rates are worth working out for yourself before quoting the numbers to anyone.
In practice, an agency receiving millions of small claims cannot manually review each one for fraud, so a system flags suspicious patterns such as transactions at odd hours, unusual amounts, or new vendors from high-risk regions for human investigation. The system does not decide. It prioritizes human attention.
Classification and categorization. Given a document, image, or transaction, AI assigns it to a category remarkably well. Not perfectly, but consistently and at scale. An agency receiving thousands of citizen emails daily can have them sorted into benefits questions, complaints about service, fraud reports, and general inquiries, with humans then handling each category appropriately. Sorting 311 requests or tagging records works the same way: fast, consistent, and easy to spot-check.
Summarization and extraction. AI can read documents and pull out names, dates, amounts, decisions, and main points. Consider an agency with 20 years of case files kept as unstructured written notes, where lawyers need to know what happened in each case for legal defense. A system reads the notes and extracts a structured summary: accused stole $50,000 from the agency, evidence includes bank records and witness statements, case was settled for $30,000. A human lawyer reviews the extraction. This was Priya's morning win. Context and nuance still get lost, which is exactly why the review step is part of the task rather than an optional extra.
Translation. AI translates between languages reasonably well, and noticeably better for high-resource pairs such as English and Spanish than for low-resource pairs. A federal agency serving immigrants who speak many languages can translate documents and interactions and deliver service that otherwise would not happen at all. It is not perfect and context gets lost, so it works as assistance rather than as a replacement for human translation on sensitive documents, and anything official should be checked by a bilingual human.
Generation. Producing first drafts of routine text, such as a memo, a set of frequently asked questions, or a public notice, for a human to refine into something the agency will stand behind.
Prediction, with caveats. AI predicts outcomes from historical patterns, not because it understands cause and effect but because it found correlations. A school district wanting to identify students at risk of dropping out can train a system on historical student data covering attendance, grades, and demographic factors, and flag likely cases for counselor intervention. The caveat is load-bearing. The predictions are only as good as the historical patterns. If past data shows low-income students dropped out more often, the system will flag low-income students more often. That correlation might be real, because poverty affects education, or it might reflect past neglect of low-income students by the school. The system cannot tell you which, and it will not warn you that the question exists.
Notice the common thread in the strong cases: the AI works from material you supply, and a human can verify the output against that material. That is the safe zone, and it is safe only when somebody actually does the verifying. Errors being catchable is not the same as errors being caught.
Where AI fails
AI struggles when a task requires it to know things, reason carefully, or understand a situation it was not given. These failure modes cover most of the trouble.
Reasoning and causal logic. AI does not reason and does not understand cause and effect. It finds correlations. That matters whenever the task involves thinking like "if we do X, then Y follows" or "because of Z, we should do W." Ask a system trained on homelessness data what to do about homelessness and it may report that people in cities with mild climates are less likely to be homeless. Should the policy be to move homeless people to California? Obviously not. The system found a correlation between mild climate and lower homelessness without reasoning about causation, and it ignored the causal pathways the data never contained: housing prices, job availability, social networks. Policy requires reasoning about causation. AI is bad at this.
Factuality without external sources. Language models cannot reliably distinguish true from false. They are trained to generate plausible-sounding text, and plausibility and truthfulness are not the same property. Asked about a Supreme Court ruling, a model will confidently explain a ruling that does not exist, or misrepresent a real one, and it sounds authoritative either way. This is called hallucination, and it is Priya's afternoon. For government it is particularly serious, because official guidance must be factually correct: an agency cannot publish a model-generated explanation of legal obligations without verifying it against the actual law. Never use these systems as authoritative sources, and never trust an AI for a fact you have not verified.
Contextual judgment. AI does not understand context and cannot adapt to "this is a special situation." A welfare agency runs a rules-based threshold: if income is below X, approve benefits. An applicant's income is above X, but they were suddenly laid off and their savings are draining. A human can say the rules point one way and this case warrants a second look. A system cannot, unless someone explicitly programmed that path. Government decisions constantly turn on questions of this shape: does this person deserve a waiver, is this an exception, what is right given these unusual circumstances? Those require human judgment.
Common sense. AI lacks human intuition. A person knows you cannot put a living dog in the freezer to preserve it, without ever having been told; a system trained on data about food preservation might suggest it. The government version is quieter. A scheduling optimizer assigns an elderly caseworker eight office visits a day across the city. The system met its efficiency target. It did not account for human limits, and by the third day the caseworker is exhausted and service quality falls. The output was technically formed and made no real-world sense.
Natural language nuance. AI struggles with sarcasm, irony, metaphor, and cultural context, and sarcasm especially trips up text systems. A citizen writes "Great job on the service here, only three weeks to get an approval." A sentiment analyzer may score that as positive and mark the issue resolved. A human reads it correctly as a complaint. If sentiment scoring feeds any queue that decides whose problem gets attention, this failure mode is not cosmetic.
The grey zone: tasks that look like AI problems and are not
Between the clear wins and the clear failures sits a set of tasks that sound algorithmic and are not. These are where agencies lose the most money and do the most damage, because the pitch is genuinely persuasive.
High-stakes decisions without clear rules. Consider a system that would decide whether to approve disability benefits. Benefits determination means assessing someone's capacity to work given their condition, age, job market, and location. That is not a classification problem; it is a judgment call. Detailed frameworks exist, but they are complex and contextual. A system can assist by extracting information and identifying cases at the clear ends of the distribution. It should not replace human judgment on the determination.
Tasks with changing definitions. If what counts as success shifts over time, AI struggles, because it is trained on historical patterns. Fraud is the standing example. Fraud evolves, criminals adapt, and last year's patterns are not this year's. A system trained on 2023 fraud data will miss 2025 fraud techniques, and it needs continuous retraining just to stay level.
Tasks with insufficient or biased historical data. If you do not have good training data, do not use machine learning. Predicting which state employees will become leaders fails on both counts: you do not have decades of clean data on it, what makes someone a leader is itself changing, and past promotions reflected the biases of past decision-makers. A system will perpetuate those biases rather than predict leadership potential, and it will do so with a confidence the underlying data does not support.
Tasks where failure is catastrophic. If the consequences of being wrong are severe, human judgment makes the decision. A system recommending whether to approve new medications is the clean example: this is life and death, AI can assist by analyzing clinical trial data, and it should not make the call.
A usable artifact: the fit-and-verify test
Before you use AI for any task, run it through two questions. They will keep you in the safe zone and out of Priya's afternoon.
Question 1. Is this a transform task or a knowledge task?
- A transform task reshapes information you provide: summarize this, sort these, draft from this, translate this. AI is strong here.
- A knowledge task asks AI to supply facts, do careful reasoning, or know context it was not given. AI is weak here, and you should be on guard.
Question 2. What happens if it is confidently wrong?
- Low stakes and easy to check: use AI as a normal part of the work, and still glance at the result before it goes anywhere.
- High stakes or hard to check: use AI only as a starting point, and verify every consequential element against an authoritative source before it leaves your hands.
Two constraints sit outside this grid and never move, whichever quadrant you land in. Use only the tools your agency has approved, and never put sensitive citizen data into a tool that has not been cleared for it. A green-quadrant task does not turn an unapproved tool into an approved one, and "low stakes" describes the consequence of a wrong answer, not the consequence of a data disclosure.
| Low stakes / easy to check | High stakes / hard to check | |
|---|---|---|
| Transform task | Green: use AI, review before it goes out (summarize a meeting) | Yellow: use AI, verify carefully (summarize a legal brief) |
| Knowledge task | Yellow: use with caution (brainstorm ideas) | Red: do not rely on AI (state a regulatory deadline) |
Priya's two tasks now sort themselves instantly. Summarizing 600 comments is a transform task, checkable against the comments themselves: green, and a good use of AI, provided somebody spot-checks the themes against the source. Answering a citizen with a specific regulatory deadline is a knowledge task, high stakes, and the answer was never in anything she fed the tool: red, never to be used without checking the regulation itself. Same tool, opposite quadrants. Had she run the test, the morning win would have stood and the afternoon apology would never have happened.
Four pairs that make the line concrete
Good: mail sorting. A postal service needs to sort millions of pieces of mail. Optical character recognition reads the address and a classifier assigns a region. The task is clear, the volume is enormous, the stakes on any single item are low, and the value is high because the labor saved is immense. This is the shape of an ideal AI task.
Bad: eligibility determination. Someone applies for benefits and a system determines eligibility. The stakes are high, since it affects someone's income. It requires judgment about special circumstances. It carries political and legal exposure if a demographic group sees lower approval rates. And it leaves no real room for appeal, because "the algorithm said no" is not an explanation a person can answer.
The workable version keeps a human on every determination: the system can sort the queue so that clear-cut applications reach a caseworker faster and complex ones get more time, while a person confirms the outcome, signs it, and can explain it. Automate only decisions with clear rules and no exceptions, and note that an automatic denial does not meet that bar even when the rule looks simple, because a denial is exactly the decision a person has the right to appeal to another person.
Good: research assistance. A policy team researching housing costs across states has a system read the studies, extract key statistics, and identify trends. A human then writes the analysis and interprets what the numbers mean for policy. The AI handles the tedious reading of hundreds of studies; humans do the judgment about what it means and what to recommend.
Bad: policy recommendations. "Use AI to determine the best housing policy" fails because the task requires judgment, understanding of causality, weighing trade-offs, assessing political feasibility, and ethical decisions. Every one of those is human reasoning, and none of them is pattern completion.
Why the cost of getting it wrong is higher in government
In a business, a confident AI mistake might lose a sale. In government, the same mistake goes out under official letterhead. A wrong deadline, a fabricated regulation, an invented case citation, or a mistranslated legal notice does not just embarrass. It can mislead the public, trigger appeals, and damage the trust that lets government function at all. Priya's correction notice was cheap as these things go. The same error inside a benefits denial or a compliance instruction is not cheap, and the person who pays for it is rarely the person who made it.
That is why the verify step is not optional caution. It is the job.
Anti-patterns
Expecting AI to reason about cause and effect. Using a system to answer "why" questions when it can only answer "what" questions. Policy makers ask why poverty is increasing in a region, the system reports that areas with more industrial closures have more poverty, and people read a correlation as a cause and recommend supporting manufacturing. That may be right. It may also be that the real driver is something else correlating with both. The system did not think; it found a pattern. Use AI for what, where, and when questions such as which regions, how many cases, what is trending. Use human analysis for why and what-should-we-do.
Automating decisions that require judgment. Removing humans from decisions that need context, fairness assessment, or exception handling. A system automatically denies benefit applications where income exceeds a threshold, and a person is denied despite severe medical expenses and heavy debt, because the system cannot consider fairness. It does not hold the concept. Automate only where rules are clear and exceptions do not exist; use AI to assist with judgment calls rather than to replace them.
Trusting AI predictions because they feel data-driven. Assuming a prediction is more reliable than human judgment because a machine produced it. A system flags a student as high dropout risk, the counselor does not intervene because the AI presumably knows better, and the student drops out. This is automation bias, and it inverts the intended safeguard: the prediction was meant to prompt human attention, not replace it. Understand what a prediction is based on, whether the data is recent and unbiased, and compare the system's output against expert judgment rather than deferring to it.
Treating "low stakes" as permission to skip the rules. The green quadrant governs how much you verify, not which tool you use or what data you paste into it. Summarizing your own meeting notes in an unapproved consumer tool is still an unapproved tool. Summarizing a case file in one is a disclosure, regardless of how routine the summarizing felt.
Practice prompts
- Matching tasks to capabilities. Identify five tasks in your agency. For each, ask whether AI's strength in pattern recognition helps, whether human reasoning is required, and whether exceptions exist that demand judgment.
- Fact versus pattern. Find a news article about an AI system that "discovered" something. Does the article confuse correlation with causation? Write one sentence stating what the evidence actually supports.
- Failure analysis. Pick a government AI system. What would happen if it made a wrong decision? Is that consequence acceptable, and at what point does human review become mandatory rather than advisable?
- Grey zone analysis. Identify a task in your agency that looks like a strong AI candidate but actually requires human judgment. Explain precisely which part requires the judgment.
- Hybrid design. For a process you know well, design the split: what would AI do, what would humans do, and how would responsibility be divided when something goes wrong?
- Quadrant sort. Take five tasks you did last week and place each one on the fit-and-verify grid. Any surprises are the point of the exercise.
Reflection
Think of a major pain point in your agency, something that consumes staff time or creates a bottleneck. Test it against the shape of a good AI task. Is the volume high? Are the rules clear? Does it require judgment, which makes it poor for AI alone and potentially good for AI assistance? Is it people-focused, which makes the answer depend entirely on the details?
Then draft a short proposal covering how AI could help and, in the same document, what humans would still need to do. The second half is the part that makes the first half safe, and a proposal missing it is not yet a proposal. If you cannot describe who verifies the output and what authority they have to reject it, you have not finished scoping the work.
Glossary
- Pattern recognition. AI's ability to find correlations and relationships in data that humans might miss, especially at scale.
- Correlation. When two variables tend to occur together. It does not imply causation.
- Causation. When one variable directly causes change in another. Much harder to establish than correlation.
- Contextual judgment. Decision-making that requires understanding unique circumstances and adjusting the application of rules accordingly.
- Hallucination. When an AI system, especially a language model, generates false information while sounding confident.
- Human-in-the-loop. A design in which AI provides input and humans make the final decision.
- Automation bias. The human tendency to favor decisions made by automated systems even when they should be questioned.
- Task-appropriate. Whether a tool is genuinely well suited to the problem it is being used on.
- Transform task. A task that reshapes information you supply, such as summarizing, sorting, drafting, or translating.
- Knowledge task. A task that requires the system to supply facts, reason carefully, or know context it was never given.
Related lessons
- What AI Is and Is Not sets the boundary this lesson turns into an operating test.
- How AI Actually Works explains why a system that predicts likely words is strong on transform tasks and unreliable on knowledge tasks.
- Types of AI Systems gives you the vocabulary for naming which kind of system a task actually calls for.
- AI Confidence and Hallucination goes further into why confident output feels trustworthy and how to catch invented content.
- The AI Decision Framework: When to Use and When Not To extends the fit-and-verify test into a fuller decision process for a whole program.
Closing
You now know AI's real strengths and real limits, and you can look at a government task and assess whether it is a good AI problem. The pattern is consistent across every example in this lesson. AI works best for high-volume, low-stakes assistance tasks where a human can check the result against material the human supplied. It is dangerous for high-stakes decisions, especially those requiring judgment or fairness assessment.
That is the reality check. AI is powerful and it is not magic. It does not understand, it does not reason, and it does not care about fairness, because caring is not the kind of thing it does. Use it where it is strong, refuse it where it is weak, and keep the verification step even on the days when the output looks perfect. Especially then, because the output that looks perfect is the one that gets forwarded without reading.
Key Takeaways
- Most AI failures are failures of fit, not technology. A good tool aimed at the wrong task produces confident, polished mistakes, so the core skill is knowing where the line falls.
- AI is fluent before it is correct. It predicts likely words, which makes it a brilliant pattern-completer and an unreliable source of facts.
- AI excels at transform tasks. Pattern recognition at scale, classification, summarization and extraction, translation, drafting, and prediction from history all work from material you supply and can verify.
- AI fails at knowledge tasks. Causal reasoning, factuality without sources, contextual judgment, common sense, and language nuance are where it produces confident, false answers.
- The grey zone is where the money goes. High-stakes decisions without clear rules, shifting definitions, thin or biased data, and catastrophic failure modes look algorithmic and are not.
- Run the fit-and-verify test first. Ask whether the task transforms information or demands knowledge, and what happens if the answer is confidently wrong. Approved tools and data-handling rules apply in every quadrant.
- The cost of error is higher in government. A fabricated date or citation goes out under official letterhead, so verifying every consequential AI output against an authoritative source is the job, not optional caution.
Frequently Asked Questions
If AI summarized 600 comments correctly, why could it not answer one question about the rule?
Because those are different kinds of task. Summarizing works on material you supplied, so the answer lives in the input and a human can check it against the source. Answering a question about a regulatory deadline requires the system to supply a fact from its own memory, which it does not have in any reliable form. It produced a plausible-looking date because plausible-looking dates are what follow that kind of question in the text it learned from.
Can we use AI for eligibility screening if a human still signs the decision?
Sorting the queue is a legitimate use: a system can help clear-cut applications reach a caseworker faster and give complex ones more time. What should not happen is an outcome, particularly a denial, being issued because a system placed the case in a bucket. A denial is precisely the decision a person has a right to have explained and to appeal, and "the algorithm said no" is not an explanation. Keep a named human confirming the determination and able to state the reasons.
The tool is right almost every time. Do I really have to check?
Yes, and being right almost every time is what makes checking necessary rather than optional. A tool that failed obviously would be safe, because you would catch it. A tool that is usually right trains reviewers to skim, which is automation bias, and the one wrong answer then travels furthest. Errors being catchable is not the same as errors being caught; that difference is entirely a matter of whether someone actually looks.
Is it safe to paste a document into an AI tool if the task is low stakes?
Stakes and data handling are separate questions. "Low stakes" describes what happens if the answer is wrong. It says nothing about what happens to the content you pasted in. Use only tools your agency has approved, and keep sensitive citizen information out of anything not cleared for it, regardless of how routine the task feels. A green-quadrant task in an unapproved tool is still an unapproved tool.
Where does translation sit on this line?
Usefully in the middle. Machine translation works reasonably well and better for high-resource language pairs such as English and Spanish than for low-resource ones, and it enables service delivery that otherwise would not happen. It also loses context. Treat it as assistance, not as a replacement for human translation on sensitive documents, and have a bilingual human check anything official before it reaches the public.
Skill.re