When Government AI Goes Wrong
In October 2013, a man named Brian got a letter from the Michigan Unemployment Insurance Agency. It said he had committed fraud, owed roughly $50,000 in penalties, and that the state would be garnishing his wages and seizing his tax refund. Brian had not committed fraud. Neither had tens of thousands of others who got nearly identical letters. The accusations came from an automated system called MiDAS, which had been turned loose to find fraud with almost no human checking its work. Its error rate, a later review found, was around 93 percent. That number is not a typo. The machine was wrong more than nine times out of ten, and the state collected on those errors for two years before anyone stopped it.
These are not hypotheticals. They are real people harmed by real AI systems in real government agencies. You are studying them not to memorize horror stories but to recognize the early warning signs in your own agency, before a system ships and before citizens get hurt. Learning from other agencies' failures is the cheapest form of due diligence available to you. Three cases show the same failure pattern from three angles, and the pattern is more useful than any one of them.
Michigan's MiDAS: automation without a human
Michigan replaced human fraud examiners with software. The system was trained on historical fraud cases and flagged claims that did not match historical patterns of "legitimate" claims, then automatically issued fraud determinations, sometimes with no human review at all. The penalties were quadruple the disputed amount. The system reversed the burden of proof: accused residents had to prove their innocence, often after their wages were already gone. Benefits stopped. People waited months with no income. Some lost their homes. Some went hungry. Lawsuits and a state review followed.
Five things went wrong, and none of them was the algorithm's arithmetic. Training data bias: historical fraud cases reflect investigator and policing biases, not actual fraud rates, and the system learned those biases, systematically over-flagging claims from certain demographic groups. Inadequate testing: the system was not tested for demographic disparities before deployment. Scale: it was deployed statewide at once, reaching hundreds of thousands of people. Opacity: people did not understand why their claims were flagged and had no way to contest the decision. The urgency excuse: a pressing operational crisis was treated as a reason to skip proper due diligence.
Underneath all five sits the design choice to remove people from a high-stakes decision. The state took a tool that should have been Yellow, useful for flagging cases for a human to investigate, and ran it as Red, letting it punish citizens directly. Michigan eventually refunded millions and overhauled the system. The lesson: the more a decision can ruin someone's life, the less you should let software make it alone. A crisis is not an excuse for skipping fairness testing. High-stakes decisions require human review. Transparency and appeals are not optional extras. Demographic testing must happen before deployment, not after the letters go out.
COMPAS: bias hiding inside "objective" scores
COMPAS is a risk-scoring tool used in some courts to estimate how likely a defendant is to reoffend. Judges use these scores in sentencing and parole decisions, and they saw a number and treated it as neutral. In 2016, journalists at ProPublica analyzed thousands of these scores against what actually happened. They found the tool was about as accurate overall for Black and white defendants, but it made different kinds of mistakes for each group. It falsely flagged Black defendants as high-risk at roughly twice the rate it falsely flagged white defendants. White defendants who did reoffend were more often labeled low-risk.
Researchers who audited the system described it as systematically overpredicting recidivism for Black defendants and underpredicting it for white defendants, with the consequence that Black defendants were given longer sentences on the basis of predictions that were wrong in a patterned way. The vendor disputed the framing, and the math here is genuinely subtle: you cannot satisfy every definition of fairness at once. But that is exactly the point. A score that looked objective carried a value judgment buried inside it, one no judge had voted on.
Four failures compounded each other. The tool learned from arrest and conviction data, and arrests and convictions reflect policing biases, not actual crime rates. That created a feedback loop: the system was trained on arrests, so it predicted higher risk for over-arrested groups, so those groups drew more police attention, so more arrests followed, so the bias was reinforced. There was a lack of transparency, with the system used for years before its biases were publicly understood. And there was an assumption of objectivity, the belief that an algorithm would be more impartial than a human judge, when in fact it learned human biases from its training data.
The lesson: a number is not neutral just because a computer produced it. Someone chose the data, the definition of "risk," and which errors mattered. A system touching fundamental rights such as freedom and sentence length was not audited for fairness before deployment, and algorithms do not eliminate human bias. They scale it, and they lend it the authority of arithmetic.
Child-welfare screening: good intentions, hard tradeoffs
From 2016 onward, multiple cities, counties, and states have used predictive tools to help screen calls to child-abuse hotlines, scoring which families to investigate so that limited investigative resources go where they are most needed. Allegheny County in Pennsylvania built one of the more transparent versions, with public documentation and ongoing review. Even there, researchers and advocates raised a hard concern: the tools lean on data like prior use of public benefits, which is recorded more often for poor families. A model can end up treating poverty as a proxy for risk.
The consequences are specific. Children in certain neighborhoods were investigated more frequently even after controlling for actual risk, and families ended up with investigations in their records even where no abuse was found. That record follows people. An investigation that finds nothing is not a neutral event in a family's life, and a system that generates more of them for one group has imposed a real cost on that group regardless of what it protected.
The deepest problem here is the absence of ground truth. "Risk" was measured by investigations rather than by actual abuse, so the systems learned to predict who would be investigated rather than who was actually at risk. Add proxy variables such as poverty and neighborhood, which correlate with race, and training data drawn from investigation history, which reflects investigation bias, and you have a tool that reproduces the pattern it was meant to correct. Child welfare is already high-stakes, so biased decisions carry serious consequences, and for years the biases were not publicly acknowledged.
This case is the most important because nobody here was a villain. The county wanted to protect children and was unusually open about its system. The failure mode is quieter: a well-meaning tool that systematically over-investigates the families with the least power to push back. The lesson: good intentions and transparency are necessary but not sufficient. You still have to test who the system burdens, be careful about what counts as success, and give vulnerable populations extra protection rather than less.
The pattern behind every failure
Across all three cases the algorithm was rarely the real culprit. The culprit was a human decision about how much power to hand the algorithm, and how little to check it. Strip the cases down and the same five gaps appear every time. The U.S. Government Accountability Office published an AI Accountability Framework built around the same idea: governance, data, performance, and monitoring. Here they are as warning signs you can spot from inside any program.
- No meaningful human in the loop. The system acts on people instead of advising a person who acts.
- Biased or proxy data. The training data reflects past inequities, or uses stand-ins like benefits history that track poverty rather than the thing you actually care about.
- No measured error rate by group. Nobody asked "how often is this wrong, and is it wrong more for some people than others?"
- Reversed burden and no appeal. The citizen must prove the machine wrong, with no fast, fair path to do so.
- No monitoring after launch. The system runs for months or years before anyone audits whether it is working.
Six cross-cutting lessons follow from the three cases together. Crisis does not justify skipping due diligence; if anything, high-stakes situations demand more caution. Historical data encodes historical biases, so be very careful about what you are training on. Transparency and explanation are essential, because people affected by AI decisions have a right to understand why. Human review is necessary for high-stakes decisions, and no algorithm is good enough to remove humans from the loop. Testing for bias before deployment is non-negotiable. Feedback loops amplify bias, so what starts as biased training data grows more biased as the system influences what happens next.
A pre-launch early-warning checklist
Before any AI system that touches citizens goes live, walk this list with the program team. Any "no" or "we don't know" is a reason to slow down, not speed up.
- Stakes: Could a wrong output cost someone money, liberty, custody, or benefits? If yes, this is high-stakes and the rest of the list is mandatory, not optional.
- Human authority: Does a named person make the final decision and have real power to overrule the system? Not a rubber stamp, an actual reviewer with time to review.
- Data origin: Where did the training data come from, and what bias might it carry? Are any inputs proxies for race, income, or neighborhood? Does the label you are predicting measure the thing itself, or only who got noticed?
- Error rates by group: Do we know the false-positive and false-negative rate, broken out by the populations we serve? If not, we are not ready.
- Appeal path: Can a citizen find out an AI was involved, see why, and challenge the result quickly? Who staffs that?
- Monitoring plan: Who audits performance after launch, how often, and what triggers an automatic pause?
- Sunset trigger: What result would make us turn this off? Decide it now, in writing, while no one is invested in keeping it running.
If Michigan had answered question four honestly, a 93 percent error rate would have been on the table before the letters went out. Be precise about what that buys you, though. Measurement makes a failure visible; it does not by itself stop a launch. Someone with authority has to be required to read the result and empowered to act on it, which is why the list ends with a monitoring plan and a sunset trigger rather than with the test. Most government AI disasters were not caused by exotic technology. They were caused by skipping a checklist like this one, or by running it and filing the answers.
A worked case: a proposal on your desk
Imagine you are reviewing an AI system your agency has proposed. The pitch reads: "We want to implement an AI system to identify which community members are at highest need for social services. We will train the system on historical service utilization data. The system will flag people most likely to benefit from services." The intent is generous and the sentence sounds reasonable. Run it against the three cases and three red flags surface immediately.
Training on service utilization data. This measures who received services, not who needs them. Vulnerable groups may have received fewer services because of barriers to access, so the system would learn to flag them as lower-need. This is the child-welfare ground-truth problem in a new costume. Bias risk. Historical service data reflects patterns shaped by outreach and access barriers, and the system will learn those patterns. High stakes. This determines who gets offered services, so biasing it against the most vulnerable populations is exactly the wrong direction of error.
What you would recommend instead is concrete. Train on needs assessment data, meaning direct evaluation of need, rather than on service utilization. Test the system for bias before deployment, disaggregated by demographic group. Build in human review so social workers can override the system's assessment based on what they know about the person. Make the system's reasoning transparent to clients. Provide an appeal process. None of that requires blocking the project. It requires answering the questions before the system starts making decisions about people rather than after.
Anti-patterns to watch for
- "It is better than what it replaces." An agency deploys a biased system and argues the old manual process was biased too. Risk: bias is scaled and legitimized. That humans were also biased does not justify deploying a biased algorithm across a whole population at once.
- "The bias is not that bad." An agency finds bias but argues that overall accuracy is still good and the disparity is small. Risk: any systematic bias that disadvantages vulnerable populations is unacceptable, and overall accuracy is precisely the number that hides it.
- "We will fix it later." An agency deploys with known fairness concerns and promises a post-launch fix. Risk: later never comes, and the system continues causing harm to real people in the meantime.
- "There is no better alternative." An agency keeps using a biased system because alternatives are not fully developed. Risk: using a biased system is worse than using a fair one, even if the fair one is less efficient.
- Treating pre-deployment testing as a clearance. A fairness test is run before launch and the result is filed. Risk: testing surfaces the disparities that exist in the cases you tested on the day you tested. It is not a warrant that covers the system's future behaviour, which is why monitoring and a sunset trigger belong on the same checklist.
- Measuring without an obligation to act. Error rates by group are produced but no named person is required to read them or empowered to halt anything. Risk: the agency now has documentation of a harm it did not stop, which is worse than not knowing.
Practice prompts
- Does your agency use AI systems similar to the ones in these case studies? If so, have they been tested for bias, and can you find the result?
- What would it mean if one of these systems was biased against your demographic group? How would you feel receiving Brian's letter?
- If you discovered an AI system in your agency had significant bias, what would you do first, who would you tell, and what would you write down?
- Take the proposal in the worked case and write the three questions you would ask its sponsor in the first meeting.
Reflection
For each of the three case studies, write down three things: what went wrong, how it could have been prevented, and whether it could happen in your agency. Be specific on the third one. Name the system, name the decision it influences, and name the person who would have to notice. These cases are cautionary tales, and they are also evidence that the problems discussed across this course are real and have real consequences for real people. If your answer to the third question is "no," write down what makes you confident, and then check whether that thing is actually in place.
Glossary
- Bias. Systematic error or unfairness in decision-making.
- Feedback loop. When predictions influence outcomes, which then become the data for future predictions.
- Ground truth. What actually happened, as opposed to what was recorded or predicted.
- Proxy variable. A feature that indirectly represents something else, such as neighborhood standing in for race.
- Sunset trigger. A result decided in advance that will cause a system to be turned off.
Related lessons
- Algorithmic Fairness in Government covers the fairness definitions these three systems failed, and why measuring by group is the only way to find out.
- The Human in the Loop covers the missing safeguard that turns a flagging tool into a punishing one.
- Transparency: Citizens' Right to Know covers the notice and appeal rights whose absence made all three failures last as long as they did.
- Understanding AI Bias covers the mechanisms, including feedback loops and proxy variables, that produced these outcomes.
Closing
As you move forward in your work with government AI, let these lessons guide you. Test for bias. Ensure transparency. Keep humans in the loop for anything that can take something from a person. Do not deploy systems without adequate due diligence, and decide in advance what result would make you switch one off. Scale is what turns a mistake into a tragedy: a biased decision made by one person harms one person, and the same bias in an algorithm reaches everyone at once. The work is hard. The alternative, harming citizens through biased AI, is unacceptable.
Key Takeaways
- Removing the human is the most common failure. Michigan's MiDAS punished citizens directly with a roughly 93 percent error rate because no person checked its fraud findings.
- "Objective" scores hide value judgments. COMPAS made different kinds of errors for different groups; someone chose the data and the definition of risk, even if no one voted on it.
- Feedback loops amplify bias. Predictions shape who gets policed, investigated, or served, and those outcomes become the next round of training data.
- Watch what your label actually measures. Investigations are not abuse and service use is not need; a system trained on the record of who got noticed learns to predict attention, not risk.
- Good intentions are not enough. Even transparent, well-meaning child-welfare tools can over-investigate poor families when inputs act as proxies for poverty.
- Always measure error rates by group. Overall accuracy can look fine while the harm lands unevenly on the people with the least power to appeal.
- Give citizens a real appeal. A system that reverses the burden of proof and offers no fast challenge will eventually punish the innocent.
- Monitor after launch and decide a sunset trigger now. Name in advance the result that makes you turn the system off, before anyone is invested in keeping it alive.
- The algorithm is rarely the villain. The failure is almost always a human decision about how much power to hand the system and how little to check it.
Frequently Asked Questions
Would pre-deployment bias testing have prevented all three failures?
It would have surfaced the disparities in all three, which is not the same as preventing them. Testing produces a finding. Preventing a launch requires someone with authority to be obliged to read that finding and empowered to act on it. That is why the checklist here ends with a monitoring plan and a sunset trigger rather than stopping at the test.
Our overall accuracy is high and the disparity between groups is small. Is that acceptable?
No. Any systematic bias that disadvantages vulnerable populations is unacceptable, and an overall accuracy figure is exactly the number that conceals it. COMPAS was about as accurate overall for Black and white defendants while making entirely different kinds of errors for each. Ask for the false-positive and false-negative rates broken out by the populations you serve.
The old manual process was biased too. Why is the algorithm worse?
Because scale changes the nature of the harm and because the machine lends bias the authority of arithmetic. A biased human decision affects the cases that person touches. The same bias in a statewide system reaches hundreds of thousands of people the day it launches, and it arrives wearing a number that makes challenging it harder.
Our system just flags cases for review. Is that safe?
Only if the review is real. Michigan's system was defensible as a flagging tool and catastrophic as a deciding one, and the difference was whether a person with time, information, and authority stood between the flag and the citizen. If your reviewers cannot reverse the flag quickly, you are running a deciding system with a flagging system's governance.
We are under enormous time pressure. Can the fairness work wait?
Crisis does not justify skipping due diligence; if anything, high-stakes situations demand more caution. The urgency excuse is one of the five identified causes of the Michigan failure, and the state spent two years collecting on erroneous determinations and then refunded millions and rebuilt the system. The fast path was the slow one.
Skill.re