Bias Detection Tools and Methods
Sam Whitfield, a data analyst at a city housing authority, was asked a question he could not answer. The authority used an AI tool to rank applicants for a waitlist, and a council member wanted proof it was not disadvantaging any group. Sam's boss said, "Just confirm it's fair." Sam stared at his screen. He had the tool's predictions and the applicant records. What he did not have was any idea how to turn "is it fair?" into a number he could actually compute, defend, and show a skeptical council member. "Fair" felt like a feeling. His job was to make it a measurement.
That translation, from a fairness question into concrete metrics you can compute with real tools, is a practitioner skill and it is learnable. You do not need to be a research scientist. You need to understand a handful of statistical fairness measures, know the open-source toolkits that compute them for you, and follow a disciplined testing routine. This lesson walks through all three, using Sam's waitlist as the running example.
Why Detection Is a Requirement, Not a Courtesy
Government AI systems must comply with Title VII of the Civil Rights Act, with Section 508 accessibility requirements, and with dozens of other civil rights and equality statutes. Those laws prohibit discrimination. Biased AI systems that discriminate against protected groups face legal challenges, and systems that disadvantage certain populations can trigger investigations by the Equal Employment Opportunity Commission, by the Office for Civil Rights, or by state attorneys general. Detecting bias is not an optional refinement on a project plan. It is a governance and legal requirement, and the agency that cannot produce the measurement is in a worse position than the agency whose measurement showed a problem it then addressed.
The specific claims are worth stating as plainly as the source material does. A hiring system that systematically disadvantages women violates Title VII. A benefits system that systematically denies assistance to certain racial groups violates the Civil Rights Act. A criminal justice risk assessment that predicts higher risk for certain demographic groups may perpetuate historical discrimination. Whether any particular disparity in your system crosses a legal line is a question for lawyers rather than analysts, but the direction of travel is not ambiguous: government serves all citizens, and it has a duty to treat them fairly that exists independently of whether anyone sues.
There is an operational argument too, and it is not the important one but it is real. Biased systems that exclude qualified candidates, deny legitimate benefits or misdirect resources are inefficient. They waste potential and misallocate money. Detecting and fixing bias improves fairness and effectiveness together, which is why the framing of bias testing as a compliance cost imposed on a working system gets the causation backwards. Without detection, discrimination hides in plain sight, producing outcomes everyone can see and nobody can name.
Turning "Fair" Into Something You Can Measure
The first move is realizing there is no single definition of fairness. There are several, they measure different things, and they can disagree with each other on the same system. Your task as a practitioner is to compute the relevant ones and present them honestly, not to find the one number that says "fair" and stop. There is no single number that means fair, the measures can conflict, and your job is to report them all rather than cherry-picking the friendly one.
Start by setting up the comparison. Sam picked the groups to compare, and for a housing waitlist that meant factors like race and disability status. He defined the outcome he was checking: getting ranked in the top tier that leads to an offer. Everything that follows compares that outcome across those groups. Decide both of those things before you compute anything, because the alternative is choosing your comparison after seeing which comparison flatters the system.
The Core Statistical Metrics, in Plain Language
Bias detection rests on statistical comparison across demographic groups, and the underlying question is simple: does the system treat different groups similarly, or does it treat some differently from others? The measures below approach that question from different angles. Sam computed several of them, because each catches a kind of unfairness the others miss.
Demographic parity: who gets selected?
This asks whether the selection rate is similar across groups. In a perfectly even system the acceptance rate would be equal across groups; if 80% of applicants from one group are approved and only 60% from another, that 20-percentage-point disparity warrants investigation. A common way to express the same thing is the selection-rate ratio: the lower group's rate divided by the higher group's. A ratio near 1.0 means parity. In Sam's waitlist, 30% of one group landed in the top tier against 12% of another, a ratio of 0.4, and that is a loud signal. The measure is simple and interpretable, and it carries an assumption you should say out loud: it treats equal rates as the fair result, which may not hold if groups genuinely differ in the underlying need or qualification the system is supposed to measure.
Equalized odds: are the mistakes spread evenly?
Selection rates alone can mislead, so Sam also checked error rates. Equalized odds compares the system's accuracy across groups, specifically its true positive rates and its false positive rates. In an even system the tool is equally accurate for everyone. If it correctly identifies 90% of eligible candidates in one group but only 70% in another, it is more accurate for the first group, and that is a fairness problem even when overall approval rates match. Error-rate disparities are often where the real harm hides, because they are invisible to anyone looking only at how many people got in.
Disparate impact and the four-fifths rule
This is the screen used in civil rights enforcement, and it is the one your legal team will ask about. As commonly stated, if the approval rate for a protected group is less than 80% of the approval rate for the reference group, there is potential legal risk. Worked through: a group approved at 80% and a group approved at 60% give a ratio of 60 divided by 80, which is 75%, below the 80% screen. That triggers disparate impact concerns and may require the agency to justify the disparity. Treat the result as a trigger for legal analysis rather than as a verdict. Falling below the screen does not establish illegality, and clearing it does not establish that a system is lawful or fair; only counsel, looking at your specific decision and its justification, can make that call.
Calibration: does a score mean the same thing for everyone?
Calibration asks whether a given score predicts the same real-world outcome regardless of group. If the system reports "80% confidence" for one group and that corresponds to 80% accuracy, but the same "80% confidence" for another group corresponds to only 65% accuracy, the score means different things for different people, and every decision built on it inherits the distortion. If a score of 80 reliably means high need for one group but overstates need for another, the score is not equally trustworthy. Testing calibration separately by demographic group is what surfaces this, and it is easy to skip because a miscalibrated system can look fine on selection rates.
Error-rate disparity in high-stakes decisions
Where a decision touches liberty or safety, look hard at who bears the wrong answers. If a criminal justice risk assessment wrongly flags one group as high risk 5% of the time and another group 15% of the time, the second group absorbs three times the burden of false alarms even if the tool's overall accuracy looks acceptable. The same logic applies in the other direction to missed cases: a system that fails to identify eligible people at different rates across groups is denying help unevenly. Both directions matter, and which one matters more depends entirely on what the decision does to the person on the receiving end.
| Measure | Question it asks | What it catches | What it misses |
|---|---|---|---|
| Demographic parity | Are selection rates similar across groups? | Groups being selected at visibly different rates | Unequal errors behind equal rates; and it assumes equal rates are the fair target |
| Equalized odds | Is the system equally accurate for each group? | Errors falling more heavily on one group | Whether the score itself means the same thing for everyone |
| Disparate impact (four-fifths) | Does a group's rate fall below 80% of the reference group's? | The pattern civil rights enforcement screens for | Everything about justification; clearing the screen settles nothing |
| Calibration | Does a given score predict the same outcome for every group? | Risk or need systematically overstated for one group | Who gets selected, and how the errors are distributed |
Read that table as a set of blind spots rather than a menu. Each row is strong precisely where the row above it is weak, which is why a report built on one measure is not a fairness assessment. It is one view of a system presented as the whole picture, and that gap is where most bad fairness reporting lives.
The Toolkits
You do not compute these by hand. Several organizations have released open-source fairness toolkits that automate the calculations, and knowing your way around them is a core practitioner skill. They also supply something you should not try to eyeball: confidence intervals and statistical significance testing, because you need to know not just whether a disparity exists but whether it is distinguishable from noise.
- Fairlearn produces a dashboard comparing your chosen metrics across groups and includes methods for reducing disparities. It integrates with the scikit-learn ecosystem, and its output is straightforward to put in front of a non-technical audience like Sam's council member.
- AI Fairness 360 offers a large library of fairness metrics along with bias mitigation algorithms, as a documented Python library. Reach for it when you need a measure or a mitigation approach the first toolkit does not cover, or when you want to compare several approaches on the same data.
- The What-If Tool is an interactive visualization environment for probing model behavior across different inputs, letting you see how decisions change as you vary attributes. It is built for exploration rather than for producing a report, which makes it useful early, when you are still forming hypotheses about where a problem lives.
- SHAP (SHapley Additive exPlanations) is not a fairness tool at all. It is an explainability method that attributes a prediction to the features that drove it, which is how you find out that a model leans heavily on zip code. Feature attribution points you at suspicious variables; it does not by itself prove that a variable is functioning as a proxy.
For most government work, Sam's approach is the right default: start in a toolkit that gives a clear, presentable read across the core metrics, and reach for a broader library when you need a metric or mitigation the first does not provide. All of them take the same three inputs, your predictions, the true outcomes, and a group label, and return the comparisons. The tool does the math. Your judgment decides which metrics matter, and your legal team decides what the results mean.
A Disciplined Testing Routine
Tools without method produce confident nonsense. Follow a fixed routine so your results are sound and defensible, and write the routine down before you run it.
- Identify the sensitive attributes. Which characteristics are protected by law or otherwise important for fairness here? Common ones are race, gender, disability status, age, national origin and religion. The relevant set depends on the decision, so know your context rather than copying a list.
- Define groups and the favorable outcome first. Decide what you are comparing and what counts as a good result before you run anything, so you are not fishing for a flattering answer.
- Collect ground truth labels. You need accurate outcomes for a test set: actual hiring decisions, actual eligibility determinations. Not the system's own predictions, which would only tell you the system agrees with itself.
- Run the system on the test set and generate predictions for every case.
- Disaggregate by group and check your sample sizes. Break the results down by protected characteristic. A dramatic-looking disparity in a group of twelve people may be noise; note where groups are too small to support a firm conclusion rather than reporting the number as if it were solid.
- Compute several metric families. Selection rate, error rates, disparate impact and calibration. Never report just one, because they measure different things and can point in different directions.
- Test for statistical significance. Is the disparity real or random? Use chi-squared or another appropriate test, and let the toolkit give you confidence intervals alongside the point estimates.
- Look for the proxy. If you find a disparity, examine which features drive the score and whether any of them quietly track a protected characteristic.
- Make the legal and ethical assessment, with the right people. Once disparity is measured, someone has to decide what it means: whether it is legally problematic, whether it is ethically defensible, and what the response will be. Technical teams supply the numbers; legal interprets them.
- Document method and limits. Record what you tested, on what data, with what group definitions, and state plainly what the numbers can and cannot establish.
Proxy Variables: The Bias You Did Not Put There
A system can encode a protected characteristic without ever being given it. This is the most insidious form of bias, because the model's inputs look clean and its behavior is not. The usual suspects are well known: zip code, which often correlates strongly with race because of residential segregation; name, which can indicate ethnicity or gender; school attended, which tracks socioeconomic status; credit history, which correlates with income, which in turn correlates with race; employment history, which encodes whatever discrimination the labor market has already done; and address history, which carries the residue of historical housing discrimination.
The detection approach is straightforward to describe and uncomfortable to run. Remove the obvious protected attributes and test whether the unfairness goes away. If you remove race and gender from a hiring model and the model still exhibits racial and gender disparities, it is reaching those characteristics through something else. Then investigate feature importance: which variables contribute most to the predictions? If zip code and name sit near the top, you have a strong lead. Treat that as a lead rather than a conclusion, and confirm it by testing what happens to both the disparity and the model's genuine accuracy when the suspect variable is removed.
Remediation is where it gets genuinely difficult, and there is no formula. Some proxy variables are legitimate: school attended may predict job performance for reasons that have nothing to do with protected characteristics. Others are not: a name should never predict a hiring outcome. Deciding which is which requires domain expertise and legal analysis together, and it is not a call an analyst should make alone in a notebook. What you can do alone is surface the candidates, quantify their contribution, and put the decision in front of the people whose job it is to make it.
Monitoring After Deployment
Bias detection does not end at launch. Systems drift, data distributions change, and a system that was even-handed in development can become unfair in production as the population it serves shifts or as the world around it changes. A workable monitoring pattern is to calculate fairness metrics monthly for the first six months and quarterly thereafter, tracking the trend rather than each isolated reading, with alerts when metrics degrade, investigation when group-level performance moves, and a retraining trigger if bias worsens.
What a monitoring dashboard looks like in practice is unglamorous. For a benefits system it might show the overall approval rate at 75% this month against last month, one group at 78% and another at 72%, a disparity of 6 percentage points tracked over time, a significance result showing the gap is real rather than noise at p below 0.05, and an escalation rule: if the disparity exceeds 10 points, escalate. That 10-point line is a threshold the agency chose for itself and wrote down in advance. It is not a legal test, it does not come from a statute, and an agency that adopts it should be able to say why that number and not another. Its whole value is that it was set before anyone saw the results, which is what stops the conversation about whether a gap is "really" large enough to act on.
Two Worked Cases
A hiring screen that learned names
A federal agency implemented an AI system to screen job applications and recommend candidates for interview. Initial bias testing showed an overall pass rate of 15%, with the male group passing at 18% and the female group at 12%, a disparity of 6 percentage points. Applying the four-fifths screen, 12 divided by 18 gives 67%, below 80%, which is a disparate impact concern. The agency's options at that point were to justify the disparity if it was genuinely job-related or to mitigate it.
Investigation using feature attribution pointed at the cause: the model was weighting names associated with male candidates. That is a proxy variable problem in its purest form, since names have no bearing on job performance, but the system had learned to use them as a predictor because the historical data it trained on carried the pattern. The response was to remove names, retrain and retest. The retested rates were 16% for the first group and 15% for the second, so the gap narrowed sharply and the ratio moved above the four-fifths screen. The source material also claims the overall approval rate dropped after the fix, comparing 16% against an 18% "average," but 18% was the first group's original rate rather than any average, and the retested rates do not support the claim; that conclusion is dropped here and the rates are left for you to work with. The substantive lesson survives intact: removing the proxy narrowed the disparity, and any change in overall selection rate has to be measured rather than assumed.
A risk assessment that learned sentencing history
A state justice system deployed an AI tool to predict recidivism risk. Calibration testing by demographic group found that when the system said "70% risk" for one group, actual recidivism was 72%, which is well calibrated. When it said "70% risk" for another group, actual recidivism was 58%. The system was assessing risk correctly for the first group and systematically overstating it for the second, which in a sentencing context means harsher outcomes for people whose actual likelihood of reoffending did not warrant them.
The analysis of why is the part worth carrying into your own work. The system had been trained on historical data that reflected past sentencing disparities, so it had learned to predict who received harsh sentences rather than who actually reoffends. The second group had historically received harsher sentences, and the model reproduced that as a risk score. The response was to retrain on actual recidivism outcomes rather than historical sentences, retest, add fairness constraints during training so the objective included calibration across groups and not only overall accuracy, and put bias monitoring in place to catch a recurrence. Note what would have happened without calibration testing: selection-rate parity alone would never have surfaced this, because the problem was not who got scored high but what a score meant.
Producing Evidence of Fair Treatment
Sam's deliverable to the council member was not the word "fair." It was a short, honest report: the selection-rate ratio of 0.4 that revealed the problem, the error-rate breakdown showing that qualified applicants in one group were ranked low far too often, and the feature analysis that traced the gap to a variable standing in for neighborhood. He also named what he could not yet prove, because some groups were too small to say anything firm about. That honesty is what made the report credible, and it is what made the follow-up work possible.
This is the deeper point for government practitioners. Bias testing is not only about catching problems; it is about being able to show your work. When a citizen, a council, an auditor or a court asks whether your AI treats people fairly, "we computed these metrics on this population, here are the group-level numbers, here are the limits of what they establish" is a defensible answer. "It looked fine" is not. What the measurement gives you is evidence rather than a verdict, and evidence is what accountable government runs on.
Anti-Patterns
- Treating a clean test as proof of fairness. Disaggregation surfaces only the groups you thought to compare, on the metrics you thought to compute, in the data you happened to have. A passed test says those specific comparisons did not show a significant gap. It does not establish that the system treats everyone fairly, and a report that claims otherwise is overstating its own evidence. Say what you tested and what remains unknown.
- One-time auditing. Bias detection is not a launch gate you clear once. Systems drift, populations change, and what was even-handed in development can become unfair in production. Testing must be continuous, with regular recalibration and a defined cadence.
- Assuming equal rates are always the fair answer. If groups genuinely differ in characteristics that legitimately bear on the decision, identical approval rates may not be the right target. A system approving 20% of one group and 15% of another might be defensible or might be discriminatory; the question is whether the difference is explained by legitimate factors or by bias, and answering it requires more than the one metric.
- Reporting a single metric. Demographic parity is one view. Equalized odds, calibration and disparate impact each catch problems the others miss, and they can disagree. Build the comprehensive picture rather than the convenient one.
- Ignoring proxies. The most dangerous bias is the kind encoded indirectly. If your system does not use race but uses zip code, you have moved the bias from explicit to implicit; it is still bias and it is still illegal. Search for proxies actively rather than waiting for one to be pointed out.
- Analysts making the legal call. Fairness metrics are technical; whether a disparity is legally problematic is not. A 5-percentage-point gap might be tolerated if genuinely job-related, and a 20-point gap might be indefensible, but those are illustrations rather than rules and neither is a threshold you can apply yourself. Technical teams provide the data; legal teams interpret the implications.
- Inventing your own legal threshold. Alert lines like "escalate above a 10-point disparity" are operational triggers your agency sets and records in advance. Presenting one as a legal standard misleads everyone downstream who reads your report.
- Detection without response. If monitoring reveals bias, what happens next decides whether any of this mattered. Does the team try to bury it? Does it reach leadership? Is there a plan and an owner? Build a culture where detection triggers response rather than defensiveness. A finding that produces nothing is documented knowledge of a harm you did not act on.
- Confusing feature attribution with proof. An explainability method showing that zip code drives predictions is a strong lead, not a demonstration that zip code is functioning as a racial proxy. Confirm it by testing what removing the variable does to both the disparity and the model's legitimate accuracy.
Practice Prompts
- Design a full bias testing plan. You are responsible for testing a state benefits eligibility system that outputs approve and deny recommendations, with race, gender and disability status as the protected characteristics of concern. Specify which metrics you would calculate, how you would define your test set, which groups you would compare, and what you would count as acceptable versus problematic disparity. Write that last one down before you see any results.
- Hunt the proxies. You are auditing a hiring system and find an 8-percentage-point racial disparity in recommendations. The system uses education level, years of experience, previous salary, zip code and name. Which of these might be proxies for race? How would you test whether each is encoding racial bias rather than legitimately predicting job performance, and who else needs to be in that conversation?
- Build the monitoring system. For a deployed criminal justice risk assessment, design continuous bias monitoring: which metrics you would track, how often you would measure them, what would trigger an alert, and exactly what you would do if the system became less fair over time. Name an owner for each step.
- Run the four-fifths screen on real numbers. Take the approval rates from a system you have access to, compute the selection-rate ratio between two groups, and write the sentence you would put in a report explaining what that number does and does not establish.
- Test the small-sample problem. Find a group in your data that is too small for a confident conclusion. Draft how you would report it honestly without alarming a council over noise or quietly dropping the group.
Reflection
Pick a government AI system you know or are responsible for and work through five questions on paper. Which demographic groups might be affected by bias in it? Which fairness metrics would be most relevant to what it actually decides? How would you conduct the test, and where would the ground truth labels come from? Who would you involve, across the technical team, legal and leadership? And what would you do if the testing revealed bias? That last question is the one people skip, and it is the one that determines whether the first four were worth answering.
Then sit with the harder version of Sam's problem. His boss asked him to confirm the system was fair, which is a request to produce a conclusion rather than a measurement. Most fairness work in government arrives in that shape. Being able to say "I can measure these specific things, on these groups, and here is what the result will and will not tell you" is not obstruction. It is the difference between a report that survives a council hearing and one that quietly becomes the agency's problem eighteen months later.
Glossary
- Demographic parity. A fairness measure comparing selection or approval rates across groups. Simple to compute and interpret, and it assumes equal rates are the fair result, which is not always true.
- Equalized odds. A fairness measure requiring comparable accuracy across groups, looking at true positive and false positive rates rather than at how many people were selected.
- Disparate impact. The screen used in civil rights enforcement: where a protected group's approval rate falls below 80% of the reference group's, there is potential legal vulnerability requiring justification. A trigger for legal analysis, not a verdict.
- Calibration. Whether a given predicted score corresponds to the same real-world outcome rate for every group. A miscalibrated system can look even-handed on selection rates while systematically overstating risk or need for one group.
- Proxy variable. A feature that does not name a protected characteristic but correlates with it closely enough to reproduce its effect, such as zip code standing in for race because of residential segregation.
- Fairness metric. Any quantitative measure of whether a system treats groups similarly. Different metrics measure different things, and one system can pass on one and fail on another.
- Ground truth labels. The actual outcomes a test set is scored against, such as real hiring or eligibility decisions. Testing against the system's own predictions measures only self-consistency.
- Statistical significance. Whether an observed disparity is distinguishable from random variation. Small groups produce dramatic-looking gaps that mean nothing, which is why significance testing and confidence intervals belong in every report.
Related Lessons
- Understanding AI Bias covers where bias enters a system in the first place, which is the context that makes these measurements interpretable.
- Bias Detection and Mitigation at Scale takes the same techniques from a single system to an agency-wide program.
- Algorithmic Fairness in Government addresses the policy and ethical framing behind the metrics.
- Systematic AI Output Validation is the broader quality discipline that bias testing sits inside.
- Continuous Monitoring Fundamentals covers the production monitoring machinery this lesson's cadence depends on.
- Human-in-the-Loop: Design and Implementation matters because the review step is often what a fairness argument rests on, and it has to be real.
- AI Incident Documentation and Response is what you need the moment a monitoring alert turns into a finding that affected real people.
- Algorithmic Impact Assessments is where these measurements get written into the formal record for a high-risk system.
Closing
Bias detection is fundamental to responsible AI governance. Agencies have a legal obligation under civil rights law to ensure their systems do not discriminate, and an ethical one that exists whether or not anyone enforces the first. The insight that matters is that fairness requires active measurement. You cannot assume it. You test, measure the disparities, search for proxies, and keep monitoring after deployment. Organizations that do this catch problems early; the others find out from someone else.
Sam's story ends where most of this work actually starts. He turned a feeling into a measurement, found a real problem, traced it to a variable standing in for neighborhood, and told a council member the truth including the parts he could not yet establish. None of that made the waitlist fair by itself. What it did was move the question out of assertion, where the answer was only as good as the confidence of whoever gave it, and into evidence, where it can be checked, challenged and improved. That move is the whole skill.
Key Takeaways
- Bias detection is a legal and governance requirement. Civil rights statutes prohibit discrimination, and disadvantaging protected populations can draw investigations from the Equal Employment Opportunity Commission, the Office for Civil Rights or state attorneys general. Not testing is not a neutral choice.
- Fairness has to become a measurement. Translate the question into specific metrics computed across defined groups for a defined outcome, and define both before you compute anything. A feeling is not evidence.
- There is no single fairness number. Demographic parity, equalized odds, disparate impact and calibration measure different things and can disagree. Compute several and report them honestly rather than picking the friendly one.
- Selection rates alone mislead. Always check error rates, because unevenly distributed false positives and false negatives are often where the real harm hides, and check calibration, because a score can mean different things for different groups while the selection rates look fine.
- The four-fifths screen is a trigger, not a verdict. A protected group's approval rate below 80% of the reference group's signals potential legal risk and calls for justification. Whether a disparity is lawful is a question for counsel, and no threshold in a dashboard settles it.
- Hunt for proxies. Zip code, name, school, credit history, employment history and address history can all encode protected characteristics. Remove the obvious attributes and test whether the disparity persists; if it does, proxies are at work, and deciding which ones to remove needs domain and legal expertise together.
- Use the toolkits, and use their statistics. Open-source fairness libraries compute the metrics and supply the confidence intervals and significance tests that tell you whether a gap is real, which matters because small groups produce dramatic gaps that mean nothing.
- Monitor continuously. Fairness drifts. Measure monthly at first and quarterly thereafter, track trends, and set your escalation threshold in advance while recording that it is an operational choice rather than a legal line.
- A clean result is evidence, not a clearance. Testing surfaces only the groups you compared on the metrics you computed. Document the method and the limits, and never let a report imply more than the measurement supports.
- Detection only counts if it triggers action. Combine technical measurement with legal review, escalate findings rather than absorbing them, and respond with retraining, adjustment or mitigation.
Frequently Asked Questions
Which fairness metric should we use? Several, always. Demographic parity tells you who gets selected, equalized odds tells you who absorbs the errors, calibration tells you whether a score means the same thing for everyone, and the four-fifths screen tells you whether the pattern is one civil rights enforcement would look at. They can disagree on the same system, and that disagreement is information rather than a problem to resolve by picking a favorite. Choose which ones matter for your decision before you run them, and report all the ones you ran.
Our disparity is small. Can we conclude the system is fair? No, and be careful with the word conclude. A small measured gap on the comparisons you ran means those comparisons did not surface a large problem. It says nothing about groups you did not compare, characteristics you lacked data for, metrics you did not compute, or how the system behaves in six months. Report the result with its scope and sample sizes attached, and keep monitoring.
What do we do if a group is too small to analyze? Say so explicitly, and resist both temptations. Reporting a dramatic ratio from a group of twelve as a finding sends a council chasing noise. Dropping the group hides exactly the people most likely to be affected by a system nobody designed with them in mind. State the group, its sample size, that the result is not statistically reliable, and what data collection would answer the question properly.
Is falling below the four-fifths screen illegal? That is not a determination an analyst can make, and it is not what the screen does. It indicates potential disparate impact and typically means the agency needs to justify the disparity, usually by showing that the criterion producing it is genuinely related to the decision being made. Whether that justification holds depends on the specific facts and on legal analysis. Bring the number to counsel with the method that produced it.
We removed race and gender from the model. Are we protected? Not necessarily, and this is the most common mistake in the field. If the disparity persists after those attributes are gone, the model is reaching them through proxies, and a model using zip code instead of race has moved the bias from explicit to implicit rather than removing it. Test after removal rather than assuming, look at which features are actually driving predictions, and confirm any suspected proxy by measuring what its removal does to both the disparity and the model's legitimate accuracy.
What if we find bias in a system that is already live? Treat it as an incident rather than a research finding. Establish who has been affected and how, get it to leadership and legal rather than working it quietly inside the technical team, decide whether the system keeps running while you investigate, and put an owner and a date on the remediation. Guard against defensiveness: a team that treats a detection as an accusation will stop detecting.
Skill.re