←
AI for HR Certification
Strategic · M3 · lesson 3 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Auditing AI Tools for Adverse Impact and Disparate Treatment: The 4/5ths Rule and Statistical Methods
📖
now learning

Auditing AI Tools for Adverse Impact and Disparate Treatment: The 4/5ths Rule and Statistical Methods

15 min

Overview

You've deployed AI in recruiting. Three months in, you need to audit it: Is the AI discriminatory?

Discrimination in AI isn't always intentional. It's often subtle. The AI screens out women at slightly higher rates than men. 30% for women, 40% for men. That 10 percentage point difference, repeated across 500 candidates, means 50 women unfairly screened out. That's disparate impact. That's potentially illegal.

You need to know. And you need to know how to measure it using the legal standard (the 4/5ths rule). You need to understand statistical significance. And you need to know: If you find bias, what do you do?

This lesson teaches you how to audit AI for adverse impact using the EEOC 4/5ths rule, how to interpret results, what remediation looks like when you find problems. By the end, you'll be able to run an audit, interpret the results, and explain your findings to leadership and regulators with confidence.

Under Title VII of the Civil Rights Act, if a selection process has disparate impact (screens out protected groups at higher rates), it's potentially discriminatory.

The 4/5ths rule (EEOC guideline):
"A selection rate for any racial, ethnic, or sex subgroup that is less than four-fifths (80%) of the rate for the group with the highest rate will generally be regarded as evidence of adverse impact."

In plain English:

If your AI screens in men at 50% rate and women at 45% rate:
- Men's rate: 50%
- Women's rate: 45%
- Women's rate as % of men's rate: 45/50 = 90%
- 90% is above 80%, so no adverse impact by this rule

But if your AI screens in men at 50% and women at 38%:
- 38/50 = 76%
- 76% is below 80%, so potential adverse impact

Important: Finding 4/5ths violation doesn't automatically mean you're liable. It means you need to investigate why. Could be bias. Could be legitimate (women applicants were actually less qualified on average). Your job: investigate and explain.

How to Audit Your AI: Step-by-Step

Step 1: Get the Data

You need:
- All candidates screened by AI in past [time period] (e.g., 6 months)
- AI's decision for each (screened in/out, confidence score)
- Candidate demographics (self-reported during application)
- Actual hiring decision (did recruiter hire after AI screening?)

Make sure:
- Data is complete (not missing demographics for some candidates)
- Demographics are accurate (self-reported is okay; more reliable than inferred)
- Time period is meaningful (6 months minimum for statistical validity)

Step 2: Calculate Selection Rates

For each demographic group:
- % screened in by AI
- % screened out by AI

Example:

Group
Screened In
Screened Out
Total
Selection Rate

Men
150
250
400
37.5%

Women
120
280
400
30%

Asian
180
220
400
45%

Black
100
300
400
25%

Hispanic
110
290
400
27.5%

Step 3: Apply 4/5ths Rule

Find the group with the highest selection rate (in example above: Asian at 45%)

For each other group: [Group rate] / [Highest rate] = ?

Group
Selection Rate
% of Highest
4/5ths Test

Men
37.5%
83%
PASS

Women
30%
67%
FAIL

Asian
45%
100%
PASS

Black
25%
56%
FAIL

Hispanic
27.5%
61%
FAIL

Interpretation:
- Women, Black, and Hispanic candidates are being screened out at significantly higher rates
- Potential adverse impact
- Need to investigate why

Step 4: Investigate Why

Is it legitimate or biased?

Legitimate reasons for differences:
- Different applicant pools (recruited from different channels; different demographics apply)
- Different qualifications (one group on average has higher qualifications)
- Different job requirements (some requirements that happen to correlate with demographics)

Bias reasons:
- Model trained on historical biased data (historical hires were mostly one demographic)
- Model learned proxy discrimination (uses variables that correlate with protected characteristics)
- Model learned stereotypes (what a "good candidate" looks like in this role)

How to investigate:
- Look at actual qualifications (do women actually have lower qualifications, or does AI think they do?)
- Look at model training data (what candidates was the model trained on?)
- Look at what variables the model uses (is it using variables that correlate with protected characteristics?)
- Compare to human judgment (have recruiters manually reviewed some of these candidates? Do they agree with AI?)

Step 5: Fix the Issue

Option 1: Retrain the model
- Use different training data (more balanced demographic distribution)
- Retrain
- Re-audit

Option 2: Adjust the model
- Weight variables differently
- Don't let model over-rely on variables that correlate with protected characteristics

Option 3: Add human review
- Have humans review borderline AI decisions
- Humans can catch bias the model misses

Option 4: Change the process
- Instead of AI screening, use different approach (blind resume review, phone screening first)

Step 6: Re-Audit

After making changes, audit again. Did the fix work?

New audit should show:
- Selection rates more balanced across groups
- All groups above 80% of highest rate

Statistical Significance Testing

The 4/5ths rule is a rough guideline. For larger sample sizes, use statistical significance testing.

Chi-square test: Are the differences we're seeing statistically significant, or just random variation?

Example: If you screened 100 women and 100 men, and rates differ by 10%, that might be random variation. If you screened 5,000 women and 5,000 men, and rates differ by 10%, that's almost certainly real bias.

If you're not a statistician:
- Use a vendor's bias audit tool (they do the math)
- Hire a data scientist to audit
- Use EEOC's Excel audit tool (they provide it free)

What matters: You're auditing, you're taking it seriously, you can defend your process.

Audit Frequency & Documentation

Frequency: Quarterly at minimum (more often if high hiring volume)

Documentation:
- Audit date and scope (what time period? what candidates?)
- Selection rates by demographic group
- 4/5ths analysis (any adverse impact?)
- Investigation findings (why do disparities exist?)
- Remediation taken (what did you do about it?)
- Re-audit results (did the fix work?)

Why: Legal defense. If someone sues, you can show: "We audited. We found issues. We fixed them."

Template: Adverse Impact Audit Report

AI Bias Audit Report

System: [Tool name and use case]
Audit Date: [Date]
Period Covered: [Jan-Mar 2024]
Methodology: 4/5ths rule, comparison of selection rates by demographic group

Selection Rates:

Group
Screened In
Screened Out
Total
Selection Rate
% of Highest

Men
150
250
400
37.5%
83%

Women
120
280
400
30%
67% FAIL

Asian
180
220
400
45%
100%

Black
100
300
400
25%
56% FAIL

Hispanic
110
290
400
27.5%
61% FAIL

4/5ths Rule Analysis:
Women: 67% (below 80% threshold) - POTENTIAL ADVERSE IMPACT
Black: 56% (below 80% threshold) - POTENTIAL ADVERSE IMPACT
Hispanic: 61% (below 80% threshold) - POTENTIAL ADVERSE IMPACT

Statistical Significance:
Chi-square test: p < 0.05. Differences are statistically significant (not due to random chance).

Findings:
We found potential adverse impact against women, Black, and Hispanic candidates. Selection rates for these groups are significantly lower than for Asian candidates.

Investigation:
- Reviewed training data: Model trained on 5-year historical successful hires
- Found: 70% of historical successful hires were male; 15% were women; 10% were Black; 5% were Hispanic
- Conclusion: Model learned historical hiring bias; perpetuating it into current screening

Remediation:
- Retrain model using balanced training data (equal representation of all groups)
- Implement bias monitoring (quarterly audits)
- Add human review for borderline cases (confidence score 60-75%)
- Schedule re-audit in 30 days to verify fix

Next Steps:
- Paused screening until re-audit confirms fix
- Re-audit scheduled for [date]
- All candidates screened out during problematic period will be manually reviewed

Interpreting Results: What to Do If You Find Adverse Impact

If you find 4/5ths violation:


  • Don't panic. You found it. That's good. You can fix it.

  • Investigate immediately. What's causing the difference? Bias? Real difference in applicants?

  • Make decision: Fix and continue, or pause and investigate more?

  • Communicate: To leadership, to affected candidates (if decisions were made).

  • Remediate: Retrain model, adjust weights, add human review.

  • Re-audit: Verify fix worked.

  • Document: Write it all down for legal defense.

Real-World Audit Example: Walking Through the Process

Let's say you audited your AI resume screening tool for the past 6 months and got these results:

Raw data:
- 1,000 candidates submitted resumes
- 400 were screened in by AI
- 600 were screened out
- Demographics:
- Men: 500 candidates, 250 screened in (50% selection rate)
- Women: 500 candidates, 150 screened in (30% selection rate)

Step 1: Calculate selection rates
- Men: 50% (250/500)
- Women: 30% (150/500)

Step 2: Apply 4/5ths rule
- Highest selection rate: Men at 50%
- Women's selection rate as % of men's: 30/50 = 60%
- 60% is below 80%
- Result: Potential adverse impact against women

Step 3: Assess statistical significance
- Sample size is large (500 of each group)
- Difference is 20 percentage points
- Chi-square test: p < 0.001 (difference is statistically significant, not due to random chance)

Step 4: Investigate why
Possible explanations:
1. Real difference in qualifications: Did women actually have lower qualifications on average?
- Review a sample of screened-out women: Are they objectively less qualified?
- Review a sample of screened-in men: Are they objectively more qualified?
2. Model bias: Did the model learn bias from training data?
- Check: What data was the model trained on? Was it biased?
- Check: What variables does the model use? Do any correlate with gender?
3. Data quality issue: Is the input data biased?
- Check: How is experience/qualifications assessed? Consistently for all groups?

Step 5: Make a decision and remediate

If investigation shows the model is biased:
1. Option A: Retrain on more balanced data
2. Option B: Adjust weights to reduce bias
3. Option C: Add human review for borderline cases
4. Option D: Don't use the model; use different approach

>
CALLOUT BOX: When to Pause the Tool

Pause immediately if:
- Large adverse impact (differences >30%)
- Hiring decisions already made based on biased screening (and you're concerned about liability)
- Regulatory inquiry received
- Multiple demographic groups showing disparate impact (suggests systemic bias)

Take time to investigate if:
- Moderate adverse impact (differences 15-30%)
- Large applicant pool (big differences might be real)
- No hiring decisions made yet (can fix before harm occurs)
- Single demographic group affected (more likely to have legitimate explanation)

Continue with enhanced monitoring if:
- Minimal adverse impact (differences 5-15%)
- Legitimate explanation for difference (e.g., different applicant qualifications)
- Strong mitigation in place (human review, appeal process)

This is judgment call. When in doubt, pause. Better safe than sorry.

Audit Checklist: Before, During, and After

Before audit:
- [ ] Determine time period (6 months minimum)
- [ ] Gather data (who screened in/out, demographics)
- [ ] Verify data completeness (missing demographics?)

During audit:
- [ ] Calculate selection rates
- [ ] Apply 4/5ths rule
- [ ] Identify any groups below 80%
- [ ] Perform chi-square test (is difference statistically significant?)
- [ ] Investigate findings

After audit:
- [ ] Document findings
- [ ] Decide on remediation
- [ ] Implement fix
- [ ] Schedule re-audit
- [ ] Communicate to leadership
- [ ] Archive documentation (for legal defense)

Deliverable: Your Adverse Impact Audit Report (2 pages)

Create a template document you'll use for quarterly audits:

Page 1: Audit Results
- Selection rates by demographic group
- 4/5ths analysis (which groups below 80%?)
- Statistical significance (is difference real or random?)

Page 2: Investigation & Remediation
- Investigation findings (why do differences exist?)
- Remediation plan (what will you do?)
- Timeline (when will you re-audit?)

What to Do Monday Morning


  • Gather audit data. Pull last 6 months of screening: who screened in/out, demographics.

  • Calculate selection rates. For each demographic group.

  • Apply 4/5ths rule. Which groups are below 80%?

  • Investigate findings. If disparities found, why?

  • Document results. Create audit report using template above.

  • Brief leadership. "Here's what we found. Here's what we're doing."

  • Schedule next audit. Quarterly, on calendar.

Key Takeaways


  • 4/5ths rule: selection rates below 80% of highest group indicate potential adverse impact.

  • Audit quarterly. Don't wait for lawsuits. Catch bias early.

  • Document everything. For compliance, learning, and legal defense.

  • Investigate before you conclude bias. Legitimate explanations exist sometimes.

  • Fix the issue when found. Don't ignore it.

  • Statistical significance matters for large samples. Chi-square test for big data.

Advanced: Auditing for Proxy Discrimination

The 4/5ths rule catches obvious bias. But AI can also discriminate indirectly through "proxy variables", variables that correlate with protected characteristics.

Example of proxy discrimination:

A model uses "years of experience" as a variable. But if women take more career breaks, "years of experience" becomes a proxy for gender. The model doesn't directly use gender, but it uses a variable that's correlated with gender. Result: indirect discrimination.

Other proxy variables to watch:
- Educational pedigree (top X universities): May correlate with socioeconomic background, which correlates with race
- Work history (no gaps): May disadvantage people with caregiving responsibilities, who are disproportionately women
- Years in current role: May disadvantage people who had to switch jobs due to discrimination or lack of opportunity
- Salary history: Often perpetuates historical pay inequities (women paid less for same work historically)

How to audit for proxy discrimination:
1. List all variables the AI model uses
2. For each variable, ask: Does this correlate with a protected characteristic?
3. If yes, investigate: Is this variable essential to the job, or is there a less-discriminatory alternative?
4. If it's not essential, remove it or adjust the model's reliance on it

Example remediation:
- "Years of experience" was correlating with gender
- Remediation: Use "relevant experience in this type of role" instead (removes the penalty for career breaks)
- Result: Model is fairer and just as accurate (relevant experience is more predictive than raw years anyway)

FAQ

Q: If we find adverse impact, are we liable?

A: Not automatically. If you can show legitimate reason for the difference (e.g., different qualifications), you're okay. But you need to investigate and document. Documentation is your legal defense. "We found this, investigated it, determined there was a legitimate reason" is much better than "We never checked."

Q: How do we measure disparate impact on things other than race/gender?

A: Same method. Age, disability, national origin, any protected characteristic. Selection rates by group, 4/5ths rule. The math is the same; you just substitute the demographic group.

Q: Can we use AI differently for different demographic groups?

A: No. That's illegal discrimination. You must use one consistent model for everyone. You cannot have "model for women" and "model for men", that's direct discrimination.

Q: What if the AI is fair but human hirers are biased?

A: That's still a problem. You need to audit the whole process (AI + human). If humans are biasing against AI findings, that's a separate governance problem. You need to train humans or change the process (e.g., less discretion for humans if they're biased).

Q: How long do we keep audit documentation?

A: 3-5 years minimum. For legal defense, longer is better. If you're sued, you want to show: "Here are all our audits. Here's our thinking. Here's what we did about problems."

What's Next

You've audited AI for adverse impact. Now you need to build inclusive AI practices across the entire employee lifecycle, not just recruiting. Next lesson: Inclusive AI Practices Across the Employee Lifecycle.

Your bias audits detect problems. Your inclusive practices prevent them.