←
AI for Government
Aware · M11 · lesson 11 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Understanding AI Bias
📖
now learning

Understanding AI Bias

10 min

Aisha Bello reviews hiring referrals for a state agency's HR office. Her team adopted an AI resume-screening tool to handle a flood of applicants for entry-level inspector jobs. For a quarter, it seemed to work; the shortlists came back fast. Then a colleague noticed something. The tool was scoring applicants higher if their resumes mentioned playing on a varsity sports team, listing a fraternity, or having attended one of three universities. None of those things measure whether someone can inspect a restaurant. But the model had learned them, because the agency's past hires, the data it trained on, happened to share those traits. The tool was not picking the best inspectors. It was picking people who looked like the people the agency had already hired, and quietly screening out everyone else.

That is AI bias. It is rarely a villain writing a discriminatory rule. It is usually a system faithfully learning patterns from the past, including the unfair ones, and then applying them at scale to thousands of people. Understand this first: bias is not a flaw in some AI systems, something that happens only with poorly built models or negligent engineers. It can exist in well-intentioned, carefully designed systems, because it comes from many sources and gets baked in at multiple stages. For government, where decisions touch rights and benefits, understanding how bias enters is not optional.

What "bias" means here

In everyday speech, bias means prejudice. In AI, it means something more specific and more sneaky: a system that produces systematically different, and unjustified, outcomes for different groups of people. The system has no opinions. It does not need to. It just reflects the data and choices that built it. That is why "was anyone acting in bad faith?" is the wrong first question, and why the answer being no tells you almost nothing about whether the system is discriminating.

The trap is that biased systems often look objective. A number on a screen feels neutral in a way a human's gut feeling does not. Aisha's tool produced clean, consistent scores. That consistency was exactly the problem: it consistently disadvantaged anyone who did not fit the historical mold. An algorithm does not remove human bias. It can launder it, turning a messy human prejudice into a clean, scalable number that feels objective, and it can do so at a volume no individual hiring manager could ever reach.

Why this is a constitutional question, not only a fairness one

In government, AI bias is not just a fairness issue. It is a constitutional issue. It is an issue of equal protection under law. The foundational principle is simple: government cannot systematically treat people differently based on protected characteristics like race, color, religion, sex, or national origin. This principle has been central to democratic governance for generations, and it does not lapse because the decision was rendered by software rather than by a person sitting at a desk.

Here is the tension that makes AI hard. An AI system can violate equal protection even when the algorithm itself makes no explicit reference to race, gender, or any other protected characteristic. A model trained on historical data can learn to discriminate through proxies for those characteristics. And the discrimination can be systematic and widespread, affecting thousands of people, before anyone notices. Civil-rights law applies to automated decisions exactly as it applies to human ones. "The algorithm did it" is not a defense.

  • Government decisions are binding. If a private company's AI gets your recommendation wrong, you go somewhere else. If a government AI denies you benefits or flags you for investigation, you are stuck with that decision until you appeal.
  • Government decisions are about essential goods. Employment, housing, food assistance, freedom, safety. These are not luxuries. This is survival.
  • Government has a duty of equal protection. This is not just ethically right; it is legally required.
  • Bias can be invisible. A human hiring manager might consciously choose the best candidate. An AI system can systematically choose people of one demographic and nobody notices for months.

The three doors bias walks through

Bias does not appear magically. It enters an AI system at three points, and knowing all three lets you ask the right question at the right time instead of assuming a tool is fair because the vendor is reputable. Each door has its own diagnostic questions, and a system can pass cleanly through one door and fail badly at another, which is why a single reassurance from a vendor about any one of them is not an answer about the system as a whole.

Door 1: Training data

This is the most common source. AI learns from examples: you show the model thousands of them and it learns patterns. The problem is that historical data often contains the biases, inequities and discriminatory patterns of the past. Sometimes it records discrimination that was explicit. More often it reflects systemic inequities baked into the collection process itself. Aisha's tool walked through this door. It learned from past hires who were not a fair sample of who could do the job.

Take the concrete case. Your agency wants AI to help decide which job applicants to interview, so you train the system on hiring decisions from the past 5 years. The system learns which types of candidates got hired and which did not. But those past decisions were made by humans who, consciously or unconsciously, carried biases. Maybe they hired more men than women. Maybe they favored certain backgrounds, or preferentially hired from certain universities. Train an AI on that and it reproduces the pattern. Deploy it and you have systematized and scaled those biases more consistently and more widely than humans ever managed alone.

A second case from actual government shows the loop this creates. A city police department wanted an AI system to predict where crime would occur so it could allocate patrol resources, and trained it on historical arrest data from the past 10 years. But arrests are not uniform across neighborhoods, and they are not even uniform with respect to actual crime. They reflect where police patrol. Neighborhoods with more police presence produce more arrests, not necessarily more crime. The system learned that crime happens in heavily policed neighborhoods and recommended sending more police there, producing more arrests, which fed the next model iteration.

Training data goes wrong in a few recognizable ways worth naming separately. Historical bias, where the data accurately records a world that was unfair, so past arrests, past lending and past hiring all carry the inequities of their time. Representation gaps, where some groups appear rarely in the data and the model performs worse for them; a speech tool trained mostly on one accent struggles with others. Proxy variables, where a neutral-looking feature stands in for a protected trait, as ZIP code can for race and as "varsity sports" did for a certain background.

To spot training data bias, ask: Where did the training data come from? Does it reflect the population you are making decisions about, or does it reflect historic discrimination? Are there demographic groups underrepresented in the training data? And were the historical decisions in the training data themselves fair and unbiased? That last question is the one people skip, because answering it honestly often means conceding something uncomfortable about the agency's own record.

Door 2: Design choices and feature selection

Even if your training data were perfectly representative, bias can still enter through the choices engineers and data scientists make. When you build a system you choose which variables, called features, to include. "We will use age, employment history, education level and credit score to predict loan repayment" sounds straightforward. But some of those features may be proxies for protected characteristics, and the model has no way to know the difference between a feature that is predictive because it measures the thing you care about and one that is predictive because it tracks who has historically been favored.

ZIP code is not itself a protected characteristic. But ZIP code is highly correlated with race in many countries because of segregation and redlining, the historic policies that prevented people of certain races from getting mortgages in certain neighborhoods. Train an AI to use ZIP code as a predictor and you are potentially using a proxy for race. Credit score works the same way. It is influenced by education, employment stability and access to credit, all of which are shaped by historic and ongoing discrimination, so using it as a predictor can indirectly discriminate against groups historically denied credit.

The other design choice that carries bias is how you define success. Building a system to predict which job applicants will succeed requires defining what success means. Tenure, meaning how long they stay? Promotion? Performance reviews? Each definition has assumptions baked in. If you define success as "gets promoted within 5 years," and your organization has a history of promoting men more than women, you have defined success in a way that biases the system toward characteristics of men. The bias is in the target, not the data, and no amount of cleaning the inputs will remove it.

Humans also decide what the system optimizes for and where the line falls between approve and deny. Tell a benefits-screening model to minimize improper payments above all else and it may learn to deny aggressively, with the people most often wrongly denied clustering in one community. The model did what it was told. The bias was in the goal. To spot design bias, ask: What features does the system use, and are any of them proxies for protected characteristics? How was success defined, and does that definition itself reflect bias? Did anyone from the affected community review these choices? Would you make the same choices if you were being audited?

Door 3: Deployment and context

A tool can be fair in the lab and unfair in the field, because the field is different from the data. A system trained in one context may not behave the same way in another, and deploying it uniformly across different populations can produce disparate impact. This is the door people forget, because it opens after launch, long after the procurement review that everyone treated as the moment of decision.

An AI hiring system trained on data from a tech company, where success is defined by engineering metrics, gets deployed across all divisions including customer service and HR. The features that predict success in engineering, such as certain educational backgrounds and specific technical skills, may not predict success in customer service. Worse, if those features correlate with demographics, the system can disadvantage certain groups in customer service even though it looked unbiased in engineering. A single agency system predicting welfare fraud across all benefit programs has the same shape of problem, because fraud in unemployment benefits looks different from fraud in food assistance and different again from housing assistance, so one system across all of them can be accurate for some populations and inaccurate for others.

Context bias also appears when you deploy without understanding local conditions. A system that works well in one region can work poorly in another where the population, economic conditions or social dynamics differ. Or staff use the tool in a way the designers never intended, applying a screening score meant for triage as though it were a final verdict. To spot deployment bias, ask: Is the system being deployed uniformly across different populations or contexts? Has anyone tested whether it works the same way across them? Are there reasons to believe it might work differently in some contexts? Who benefits from the current deployment, and who is harmed?

Where the stakes are highest

When a streaming service's recommender is biased, you get a worse movie night. When a government system is biased, people lose things that matter. Three domains carry the heaviest risk, and each has a documented case behind it rather than a hypothetical.

Benefit determination

Agencies determine eligibility for social programs: food assistance, housing assistance, disability benefits, unemployment insurance. In many countries hundreds of thousands of people depend on getting these determinations right, and bias in these systems decides who gets help. If a system is biased against a particular demographic, people in that demographic are less likely to receive benefits they are legally entitled to, and the harm falls hardest on those least able to fight it.

The state of Michigan used an AI system to detect fraudulent unemployment claims during COVID-19. The system was biased. It flagged far more claims from one demographic group than another, even after controlling for the actual characteristics of the claims. The result was that thousands of people, disproportionately from certain groups, had their benefits wrongly denied or delayed. Some lost their homes. Some went without food. The bias entered through the training data: the system was trained on past fraud cases, but past investigation and prosecution of fraud reflected investigator biases rather than actual fraud rates, and the model learned those biases and reproduced them at scale.

Law enforcement

Agencies use AI for predictive policing, which forecasts where crimes will occur; for risk assessment, which predicts whether someone will reoffend if released; and for facial recognition, which identifies suspects. Bias in these systems has profound consequences: wrongful investigation, wrongful arrest, wrongful conviction, wrongful incarceration. These are not hypothetical harms, because people's freedom is at stake, and a feedback loop of the kind described above sends more enforcement to already over-policed communities.

The best-known case is COMPAS, a recidivism prediction tool used in U.S. courts, which predicts whether someone is likely to reoffend if released and whose predictions judges use in sentencing decisions. Researchers audited the system and found it was biased: it systematically overpredicted recidivism for Black defendants and underpredicted it for white defendants. That meant Black defendants were more likely to receive longer sentences on the basis of AI predictions of future dangerousness that were systematically wrong. The bias came from the training data, which was historical arrest and conviction records, and arrests and convictions are themselves influenced by policing practices, prosecution decisions and systemic inequities in the criminal justice system.

Hiring and employment

Government agencies are major employers, and they use AI to screen applications, schedule interviews and recommend candidates. Bias in hiring systems affects people's access to employment: if a system is biased against certain demographics, people in those groups are less likely to be hired regardless of qualifications, which undermines both fairness and the agency's duty to serve a diverse public.

Amazon developed an AI recruiting tool that was biased against women. The system was trained on historical hiring data from a male-dominated tech industry, learned to prefer the characteristics of successful past hires, most of whom were men, and systematically downranked women applicants. Amazon eventually abandoned the system, but only after it had screened thousands of applications. Government hiring systems carry the identical risk. Train on your own historical hiring decisions, and if those decisions reflected bias, the system will reproduce and scale it. That is Aisha's case, arrived at from a different direction.

A usable artifact: the bias-watch checklist

You do not need to be a data scientist to ask the questions that catch most bias. Use this checklist whenever your office considers, buys, or runs an AI tool that affects people, and map each question to the door it guards. Bring it to the vendor demo rather than to the post-incident review, because every question below is cheap to ask before procurement and expensive to ask afterward.

DoorQuestion to askRed flag answer
Training dataWhat data was this trained on, and who is over- or under-represented in it?"We're not sure" or "all our past cases"
Training dataCould any input act as a proxy for race, sex, age, or disability?ZIP code, school, club, or name used as a feature
Design choicesWhat is the system optimizing for, and who could that goal disadvantage?A single narrow metric like "minimize payouts"
Design choicesHow was success defined, and did anyone from the affected community review that choice?Success defined by an outcome the agency itself distributed unequally
DeploymentDoes performance hold for the actual population we serve, by group?No performance breakdown by group exists
DeploymentHow will staff use the score, and is it being treated as final?A triage score used as a verdict
OngoingWho checks outcomes by group over time, and how often?"We tested it once at launch"

Aisha's team ran their tool through these questions and failed three of them on the first pass. The fix was not to abandon AI. They worked with the vendor to remove the proxy features, re-tested shortlists by group, and added a human review step before any candidate was screened out. The next quarter's shortlists were both more diverse and, by the agency's own measures, better qualified. Read that result correctly: it is evidence that the specific proxies they found are no longer driving the score. It is not a certificate that the tool is fair, which is why the checklist ends with a question about who keeps checking.

Working a scenario end to end

Here is how bias could enter a government AI system, and how to prevent it. Your agency manages a job training program and wants to use AI to match job seekers with the training tracks most likely to lead to employment. You have 10 years of historical data covering thousands of people who completed training, including what they did, what they were trained in, and whether they got jobs afterward. You build a system that looks at the characteristics of successful participants and recommends similar people to specific tracks. Walk the three doors.

Through the training data door: the historical record reflects who trained in the past, and who trained in the past may not represent all job seekers. Maybe certain demographics were steered toward certain tracks. Maybe certain demographics faced barriers to completing training at all. If the data overrepresents one group, the system will learn their characteristics as the shape of success and underrepresent everyone else. Through the design door: you choose "highest education level achieved" as a feature, but education access is influenced by race, income and family background, so you are using a proxy for socioeconomic status and, indirectly, for race. Weight it heavily and the system will preferentially recommend training to people who already have more education.

Through the deployment door: you deploy uniformly, but training works differently in different regions, and a track that leads to employment in an urban area may fail in a rural one. Same recommendations everywhere means biased results in some places. The preventions map one to one onto the doors. Examine your training data and ask whether all demographic groups had equal access to training in the past; if not, your data is biased. Choose features carefully and avoid proxies for protected characteristics, and if you use education level, weight it recognizing that it reflects past opportunity rather than merit alone.

Then test for bias by disaggregating results by demographic group after building the model, asking whether the system is equally accurate at predicting success for all groups, and if not, why. Deploy contextually, testing in a few regions before going statewide and monitoring whether it behaves the same way everywhere. And monitor over time, continuing to measure performance by demographic group after deployment and investigating and fixing disparities as they emerge. None of these steps requires a research team. All of them require someone to own the question after launch day.

Anti-Patterns to Avoid

  • Testing for overall accuracy only, never demographic accuracy. An organization measures its system and reports that it is 92% accurate. Disaggregate the same result and it might be 96% accurate for one demographic group and 78% for another. The overall metric hides the disparity, so the system discriminates while the dashboard stays green, and by the time someone audits by group, thousands of people have been affected.
  • Deploying without testing for bias at all. A team is excited about a new capability, wants to launch quickly, checks overall performance, sees that it looks good and ships. Nobody tests for disparate impact across demographics or looks for proxies for protected characteristics. The system turns out to be biased, you find out later, people have already been harmed, and the agency faces legal liability and damage to public trust.
  • Assuming historical data is objective. An organization reasons that training on historical data must be objective because it is based on what actually happened, and never asks whether what happened was fair. The result is an AI trained to reproduce historic discrimination, and a system that becomes a tool for perpetuating past inequities.
  • Using proxies for protected characteristics without recognizing it. "We are not using race directly, so it is not discriminatory." But ZIP code, certain names, or education from particular institutions can carry race indirectly. The discrimination is harder to spot because the proxy obscures it, and harder to defend or fix for the same reason.
  • Treating one clean quarter as a clearance. Aisha's team removed the proxies they found and the next round of shortlists improved. That is evidence the known problem is fixed, not proof the tool is fair. A system that is unbiased at launch can become biased as it is deployed, as the population it affects changes, or as the context shifts, which is why testing has to be ongoing rather than a gate you pass once.
  • Assuming measurement finds all the bias there is. Disaggregation is powerful, and it only ever surfaces disparities in the groups you thought to compare, on the metrics you thought to compute, for the population you happened to sample. Finding no disparity in the comparisons you ran is a smaller claim than "the system is fair," and it should be reported as the smaller claim.

Practice Prompts

Spend 15 minutes on this exercise. Pick an AI system, real or hypothetical, used by your agency or one you are familiar with, and answer these questions in writing. For any question where you do not have an answer, that gap is itself the finding, and it is your starting point for improvement.

  • Where did the training data come from? Who collected it? Who does it represent?
  • Are there any demographic groups underrepresented in the training data? If so, why?
  • What features does the system use? Are any of them proxies for protected characteristics?
  • How is success defined? Does that definition itself reflect bias?
  • Has anyone tested this system for bias? If so, what did they find? If not, why not?
  • If you disaggregate the system's performance by demographic group, are the results equal? If not, what explains the difference?

Reflection

  • Think about an AI system used by your agency, or one you have heard about. Where did its training data come from? Did anyone examine whether that data reflects historical bias? What would you need to know to be confident the data is unbiased?
  • If you were to audit an AI system for bias, which demographic groups would you compare? What metrics would you use? How would you know whether the disparities you found were statistically significant rather than random variation?
  • Have you experienced or witnessed a case where an automated decision seemed unfair or biased? What was the mechanism: training data bias, design choices, or deployment context? What should have been done differently?
  • In your own work, which decisions could be affected by AI in the future? Which of those decisions affect vulnerable populations? What would you want to know about an AI system before you trusted it with those decisions?

Glossary

  • Algorithmic bias. Systematic errors or unfairness in AI decision-making, particularly when performance differs across demographic groups.
  • Disparate impact. When a neutral policy or system produces unequal outcomes for different demographic groups, often disadvantaging protected classes.
  • Training data bias. Bias that exists in the historical data used to train an AI system, often reflecting past discrimination or systemic inequities.
  • Proxy variables. Features that are not themselves protected characteristics but are correlated with them, such as ZIP code standing in for race.
  • Performance disaggregation. Evaluating an AI system's accuracy separately for different demographic groups rather than only looking at overall accuracy.
  • Feedback loops. Situations where biased AI predictions influence real-world outcomes, which then become data for training the next iteration of the model, amplifying the bias.
  • Equitable deployment. Using AI systems in ways that account for different contexts and populations, ensuring the system works fairly across diverse groups.

Closing

AI bias is not new. Humans have been biased for a very long time. What is new is that AI lets us scale bias, applying biased decisions to thousands or millions of people automatically and consistently. The flip side is that AI also gives us tools to detect bias. We can measure. We can disaggregate. We can audit. We can look for where bias is hiding, and when we find it we can fix it, provided somebody is looking and provided the finding is allowed to change the system rather than merely be filed.

The key is to do the work. Do not assume your system is unbiased; test it. Look at who wins and who loses. Be honest about what you find, especially when what you find implicates the agency's own history rather than a vendor's model. Then improve. In government this is not optional. It is fundamental to equal protection, to democratic legitimacy, and to the trust citizens place in you, and the stakes are high enough that you cannot afford to deploy AI without testing it for bias. Do the work upfront. It is worth it.

Key Takeaways

  • AI bias is usually learned, not authored. It is a structural risk in any system trained on historical data, not a rare flaw in badly built ones, and no one has to write a discriminatory rule for it to appear.
  • A clean number can launder a messy prejudice. Algorithmic outputs feel objective, which makes biased ones more dangerous than a visible human bias, not less.
  • In government, bias is a constitutional matter. It engages equal protection, government decisions are binding and concern essential goods, and "the algorithm did it" is not a defense.
  • Bias enters through three doors. Training data that reflects historical discrimination, design choices that embed it in features and in the definition of success, and deployment that creates disparate impact across contexts.
  • Watch for proxy variables. ZIP code, school, club, name or credit score can stand in for protected traits, and using a proxy instead of the characteristic itself obscures the discrimination rather than removing it.
  • Aggregate accuracy hides disparity. Disaggregate performance by demographic group, because an impressive overall number can conceal a large gap between groups.
  • The highest-stakes domains are benefits, policing and hiring. Michigan's unemployment fraud flagging, COMPAS recidivism scoring and Amazon's abandoned recruiting tool all show the same mechanism producing very different harms.
  • Testing has to be ongoing. A system unbiased at launch can drift as deployment, population or context changes, so one clean audit is evidence rather than clearance.

Frequently Asked Questions

If the model never sees race or gender, can it still discriminate? Yes. An AI system can violate equal protection even when the algorithm makes no explicit reference to a protected characteristic, because a model trained on historical data learns proxies for those characteristics. ZIP code correlates with race through segregation and redlining; credit score is shaped by education, employment stability and access to credit, all of which reflect historic and ongoing discrimination. Removing the protected field does not remove the pathway.

Our vendor says the tool was tested and is accurate. Is that enough? No, and the reason is the difference between an overall number and a disaggregated one. A system reported as 92% accurate might be 96% accurate for one demographic group and 78% for another, with the aggregate metric concealing the gap. Ask for performance broken down by group for the population you actually serve, and treat "no performance breakdown by group exists" as a red flag answer.

What is the difference between bias in the data and bias in the goal? Data bias comes from examples that record an unfair world. Goal bias comes from how you define what the system is optimizing for. Define hiring success as "gets promoted within 5 years" in an organization that promoted men more often, and the target itself is biased no matter how clean the input records are. Cleaning the data will not fix a biased definition of success.

We fixed the proxy features and the results improved. Are we done? You have evidence that the specific problem you found is no longer driving the score. Testing for bias must be ongoing, because a system that is unbiased when it launches can become biased as it is deployed, as the population it affects changes, or as the context shifts. Keep the disaggregated measurement running and name the person responsible for reading it.

What should I do if I find a disparity but cannot explain it? Report it rather than resolve it privately. An unexplained disparity is exactly the condition the reporting channels exist for, and the questions that follow, whether it is statistically significant rather than random variation, which door it entered through, and who is being harmed, generally need people beyond the person who noticed. Documenting what you found and when you found it also matters, because the harm from bias accumulates over the period before anyone looks.