←
AI for Recruiters
Proficient · M27 · lesson 27 of 32 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Sources of Bias: Data, Algorithms, Humans, and Systemic Factors

15 min

Marcus runs talent acquisition for a 1,200-person regional health system, leading a team of eight recruiters who fill roughly 600 roles a year. Last quarter his team rolled out an AI resume-screening tool to handle a flood of applications for nursing and administrative positions. It worked beautifully on speed: a stack that used to take three days to triage cleared in an afternoon. Then his head of nursing asked a question Marcus could not answer: "Are we sure this thing is fair?" Marcus did not know. He could not see inside the model, he had not audited its outputs, and he had no idea whether the tool was quietly screening out qualified candidates from groups his health system was legally and ethically committed to hiring. This lesson is the framework Marcus needed: where bias actually comes from, and how to find it before it finds you.

The Four Sources of Bias in Hiring AI

Bias in a recruiting system is rarely one villain. It is four overlapping sources that compound: biased training data, algorithmic bias introduced by how the model is built and optimized, human bias in how people use and interpret the tool, and systemic factors baked into the broader hiring environment. Marcus's screening tool could be flawless on three of these and still produce discriminatory outcomes because of the fourth. Diagnosing bias means checking all four, not assuming the problem lives in the algorithm.

The frame earns its keep because it tells you where a fix belongs. A disparity caused by training data is not repaired by retraining recruiters, and one caused by uneven human interpretation is not repaired by swapping vendors. Teams that skip the diagnosis default to whichever source is easiest to act on, normally the tool, then discover a year later that the numbers did not move.

The reason this matters is legal as much as ethical. Under Title VII of the Civil Rights Act, an employer can be liable for discrimination even with no intent to discriminate, if a neutral-looking practice produces a disparate impact on a protected group. An AI screening tool is a "practice." If it disproportionately rejects candidates by race, sex, age, disability, or another protected characteristic, the fact that "the computer did it" is not a defense. Marcus, not the vendor, owns the hiring decision.

Source Where it lives What a fix looks like
Biased training data The historical record the model learned from, including proxy variables that stand in for protected characteristics Ask what the model was tuned on, and audit for proxies rather than assuming a removed field removes the bias
Algorithmic bias Design choices: what the model optimizes for, which features it weights, how it handles edge cases Treat the tool as a black box and audit its outputs, since you cannot inspect a vendor's internal weights
Human bias in the loop How recruiters read, accept, override, and interpret the tool's recommendations Structured criteria applied to every candidate, documented reasoning, and AI as one input among several
Systemic factors Sourcing channels, job descriptions, credential requirements, and pay or scheduling structures upstream of the tool Ask who is not in the applicant pool and why, then fix the structural cause

Source 1: Biased Training Data

Most hiring AI learns from historical data: who applied, who got interviewed, who got hired, who succeeded. The problem is that history encodes the past's discrimination. If a department historically hired mostly one demographic, a model trained to "find candidates like the ones we hired" will learn to prefer that demographic, even when the trait it latches onto has nothing to do with job performance. The model is not malfunctioning when it does this. It is executing its stated objective correctly, and the objective was the problem.

The well-documented historical example is a large technology company that built an experimental tool to score resumes and discovered it had taught itself to penalize applicants for women. Because the resumes it learned from came largely from men, the model treated signals correlated with being a woman as negative. It reportedly downgraded resumes that contained the word "women's" (as in "women's chess club captain") and graduates of certain women's colleges. The company scrapped the project. The lesson is not that the engineers were malicious; it is that a model faithfully reproduced the bias in its training data, and nobody noticed until they went looking. The exact internal scores were never made public, so treat the case as a cautionary pattern, not a source of numbers.

Two details of that case are worth dwelling on. Nobody would write a rule penalizing the word "women's," and no engineer would defend one; the model discovered the correlation on its own from the pattern of who had been hired. And the behavior existed for as long as the model existed, becoming known only when someone audited the outputs. That combination, an unwritable rule discovered automatically and invisible until measured, is the general shape of training data bias, and it is why "we would never build that in" is not reassurance.

For Marcus, the practical question is: what did his vendor's model learn from? If it was tuned on his health system's own ten years of hiring data, it may have absorbed whatever skew already existed in nursing or administrative hiring. This question is answerable and most recruiters never ask it. Was the model trained on the vendor's aggregate customer data, on a general labor market corpus, or on Marcus's own history? Each answer implies a different set of inherited patterns, and a vendor who cannot or will not say is telling you something useful about how much of the outcome you can explain if you are asked.

Proxy variables are the trap. The model does not need a "race" field to discriminate. A zip code, a name, a gap in employment history, or the specific school listed can all act as a stand-in for a protected characteristic. Removing the protected field does not remove the bias if the proxies remain. This is the most commonly misunderstood point in the whole topic, because stripping demographic fields feels like the complete fix and it is the visible, easy half of one. The correlated features are still present, and a model optimizing for resemblance to past hires will find them.

Source 2: Algorithmic Bias

Even with reasonable data, the algorithm itself introduces bias through design choices: what it optimizes for, which features it weights, and how it handles edge cases. A model optimized purely to predict "matches our top performers" will amplify whatever made past top performers look similar. A keyword-matching screen that rewards the exact phrasing a particular group tends to use on resumes will quietly disadvantage equally qualified candidates who phrase things differently.

The keyword example deserves a second look because it is the most common design in practice and the least examined. Resume vocabulary is not evenly distributed. It tracks the conventions of the schools, employers, and professional communities a candidate has passed through, so a screen tuned to one vocabulary is measuring exposure to that vocabulary alongside the underlying capability. For Marcus, a nurse who describes the same clinical work in the language of a different training program or a different country's system is doing the same job and failing the same screen, and nothing in the tool's output distinguishes those two explanations.

Edge case handling is the quieter design choice. Every screen has to decide what to do with an incomplete record, an unfamiliar credential, a non-linear career, or a qualification it cannot parse. Defaulting those cases to a low score is convenient and it systematically penalizes candidates whose histories do not fit the expected template, which is a population, not a random sample. Marcus cannot see how his tool makes that decision, which brings him to the only workable posture available.

Algorithmic bias is also where vendor opacity bites. Marcus cannot inspect his tool's internal weights. What he can do is treat the tool as a black box and audit its outputs, which is exactly what fairness law increasingly requires. New York City's Local Law 144 (effective July 2023) is the clearest current example. It requires employers using an "automated employment decision tool" (AEDT) on candidates for jobs in NYC to commission an independent bias audit within the prior year, publish a summary of the results, and notify candidates that the tool is being used. The audit must report selection rates and impact ratios by sex and by race/ethnicity category. The point of the law is precisely that you do not need to see inside the algorithm to hold it accountable; you measure what comes out.

Worked Example: The Four-Fifths Rule

The core method for measuring adverse impact is the four-fifths rule (also called the 80% rule), adopted in the EEOC's Uniform Guidelines on Employee Selection Procedures. The idea is simple: compare the selection rate of each group to the selection rate of the most-selected group. If any group's rate is less than 80% (four-fifths) of the highest group's rate, that is a flag for adverse impact warranting investigation.

Here is how Marcus would run it on his screening tool. The numbers below are illustrative, chosen only to show the math, not drawn from any study. Suppose over one quarter the AI advanced applicants to the interview stage as follows:

  • Group A: 500 applicants, 100 advanced. Selection rate = 100 / 500 = 20%.
  • Group B: 300 applicants, 42 advanced. Selection rate = 42 / 300 = 14%.

Group A has the highest selection rate (20%), so it becomes the benchmark. The impact ratio for Group B is its rate divided by Group A's rate: 14% / 20% = 0.70. Because 0.70 is below the 0.80 threshold, the tool fails the four-fifths rule for Group B. That is Marcus's signal to investigate, not proof of illegality, but a clear flag that the AI is advancing one group at a meaningfully lower rate.

Two features of the arithmetic are worth naming, because both are commonly botched. The comparison is between rates, not headcounts, so the fact that Group A supplied more applicants and more advances than Group B is irrelevant to the test. And the benchmark is the highest-rate group, not the overall average and not a reference group chosen in advance. Comparing every group to the pool average produces a different, smaller-looking number that no auditor will recognize. Run the check at every stage as well, not only at the final advance, because a step that produces adverse impact can be masked by a later step that partially offsets it.

Now suppose Group B had advanced 50 instead of 42: 50 / 300 = 16.7%, and 16.7% / 20% = 0.83. That clears the 0.80 threshold. The same eight-candidate difference is the line between "audit passes" and "audit fails," which is why small subgroups and small numbers need careful handling: with tiny applicant pools, a few decisions swing the ratio dramatically, and statistical significance matters alongside the raw ratio. The four-fifths rule is a screening test, not the final word, but it is the number Marcus's NYC Local Law 144 audit will turn on.

The sensitivity cuts both ways, and Marcus takes two operating rules from it. A passing ratio on a small pool is weak evidence of fairness, because the same handful of decisions that could have failed it happened to fall the other way. A failing ratio on a small pool is a reason to look at the criteria rather than to dismiss the finding as noise, since the cheapest way to learn whether the gap is real is to examine what the screen is weighting and ask whether it is job-related.

Source 3: Human Bias in the Loop

The most fairness-conscious tool in the world cannot save a hiring process if the humans using it reintroduce bias. This is the source recruiters most often overlook, because it feels like the part they control. Three patterns show up repeatedly on Marcus's team, and none of them involves anyone holding a prejudiced view. They are ordinary responses to a ranked list produced by a system nobody on the team can inspect.

The first is automation bias: trusting the AI's ranking more than it deserves. When the tool puts a candidate at the top of the list, recruiters assume the ranking is objective and stop scrutinizing. When it ranks someone low, they skip the resume entirely. The model's recommendation becomes a self-fulfilling prophecy, and any bias inside it is laundered into "the data said so." The laundering is the dangerous part. A judgment that would have been challenged if a colleague had voiced it goes unchallenged when a tool produces it, because the tool is presumed to have no stake and therefore no bias.

The second is selective override. Recruiters accept the AI's recommendation when it matches their gut and override it when it does not, which means human bias gets the final vote anyway. This pattern is worse than either extreme. A team that always defers at least inherits a consistent, auditable process; a team that always overrides is running human judgment openly. Selective override produces an outcome shaped by individual intuition while leaving a record that looks like a tool-assisted process, which is the worst combination for both fairness and defensibility.

The third is inconsistent interpretation: two recruiters reading the same AI-generated candidate summary reach different conclusions because they weight the AI's notes against their own assumptions. The same flagged gap reads as a risk to one reviewer and as a non-issue to another, and which reviewer a candidate happens to draw becomes part of the selection process. The fix is the same discipline that fights bias everywhere: structured, consistent criteria applied to every candidate, documented reasoning that can be audited, and treating the AI as one input among several rather than the verdict.

Documented reasoning does more work here than it appears to. Writing down why you overrode a recommendation, or why you accepted one, converts an invisible pattern into a reviewable record, which is the only way selective override ever becomes visible. It is also the artifact that answers the question Marcus could not answer for his head of nursing. A process whose decisions can be reconstructed can be defended and corrected; a process whose decisions live only in the recruiters' heads can be neither.

Source 4: Systemic Factors

The fourth source sits outside any single tool or person. Systemic factors are the structural conditions that shape who even enters the pipeline. If Marcus's health system only sources from networks that skew toward one demographic, the fairest screening algorithm in the world will still produce a homogeneous slate, because it can only choose from who applied. Job descriptions that use exclusionary language, credential requirements that screen out capable candidates without a clear job-related justification, and pay or scheduling structures that deter certain groups all shape the applicant pool before the AI ever runs.

Credential requirements deserve particular attention in a health system, where some credentials are genuine licensure requirements and others are accumulated habit. The distinction is not whether the requirement sounds professional; it is whether it is tied to an essential function of the job. A requirement that survives only because it has always been on the posting narrows the pool without narrowing it toward better candidates, and it does so before any measurement Marcus runs on his tool would ever see it.

Scheduling and pay structures work the same way and are almost never audited as part of a fairness review. A shift pattern that assumes no caregiving responsibilities, or a posting that requires immediate availability, filters the pool at the point of application. Nobody records those candidates, because they never applied. That is the defining property of systemic bias: its effects are missing from your data rather than visible in it, which means no output audit can detect them.

This is why fairness cannot be solved at the screening step alone. A model can pass its four-fifths audit on the applicants it received and still leave Marcus with an unrepresentative workforce, because the bias happened upstream at sourcing. Systemic analysis means asking who is not in the pool and why, then fixing the structural cause rather than blaming the algorithm for a problem it inherited. The diagnostic question is comparative: how does the composition of your applicant pool compare to the composition of the qualified labor market you are recruiting from, and what in your channels, postings, or requirements would explain a difference.

The Compliance Landscape Marcus Operates In

Several laws govern Marcus's use of hiring AI, and they overlap. Title VII (and parallel EEOC enforcement) prohibits employment discrimination and the disparate-impact theory it supports is what makes the four-fifths rule legally meaningful. The Americans with Disabilities Act (ADA) is a sharp concern with AI tools: a screen that filters on traits unrelated to essential job functions, or a video-based assessment that penalizes candidates with disabilities, can violate the ADA, and the EEOC has issued specific guidance warning that AI tools can unlawfully "screen out" individuals with disabilities. The Age Discrimination in Employment Act protects applicants 40 and older, and tools that learn to prefer "recent graduate" signals can create age-based adverse impact.

On the data side, if Marcus's health system recruits candidates in the EU, the General Data Protection Regulation (GDPR) applies: it grants candidates rights over their personal data and, under Article 22, restricts decisions based solely on automated processing, generally requiring a human in the loop and the ability to contest the decision. And NYC Local Law 144, covered above, imposes the concrete obligations of an independent annual bias audit, public results, and candidate notice for tools used on NYC roles. The unifying theme: Marcus owns the outcome, the law expects him to measure and document fairness, and "the vendor's tool did it" protects no one.

Anti-Patterns

Blaming the algorithm for a four-source problem. This is treating any measured disparity as a tool defect, which usually leads to a vendor conversation or a vendor swap. It happens because the tool is the newest and most visible component, and because "the AI is biased" is a sentence everyone understands. What goes wrong is that three of the four sources survive the change untouched: the historical patterns still shape whatever the new tool is tuned on, the recruiters still practice automation bias and selective override, and the sourcing channels still determine who applied. The counter is to diagnose the layer before choosing the fix, and to ask specifically whether the gap could have been produced upstream of the tool entirely.

Stripping demographic fields and calling it debiased. This is removing race, sex, age, or name from the data the model sees and concluding the model can no longer discriminate. It happens because the reasoning feels airtight and because the change is easy to make and easy to describe to a stakeholder. What goes wrong is that proxies remain: a zip code, a name, a gap in employment history, or the specific school listed can each stand in for a protected characteristic, and a model optimizing for resemblance to past hires will find whichever ones are available. The documented case of a resume tool that learned to penalize the word "women's" is precisely this failure. The counter is to audit outputs by group rather than to reason about inputs.

Auditing the tool and never the pool. This is running a clean four-fifths check at the screening stage, filing the result, and treating fairness as demonstrated. It happens because output auditing is the part that is measurable, required, and satisfying to complete. What goes wrong is that systemic bias is invisible in exactly this data: candidates deterred by a shift pattern, an unnecessary credential, or a sourcing channel that never reached them are not in the pool to be selected at any rate, so the audit measures fairness among the survivors of an upstream filter. The counter is to compare the composition of your applicant pool against the qualified labor market you are recruiting from, and to treat a gap there as a finding rather than as background.

Practice

These exercises follow the four sources in order, and the first one is a conversation with your vendor rather than an analysis of your data.

  • Establish what your tool learned from. Ask your vendor whether the model was tuned on your organization's hiring history, on aggregate customer data, or on a general corpus. If it was tuned on your history, describe the composition of the hires it learned from and write down what the model would have learned to treat as the target.
  • Hunt for proxies. List every input your screen consumes. For each one, ask whether it could correlate with a protected characteristic: zip code, name, employment gaps, school, phrasing conventions. Note which ones you could not justify as job-related if asked, and what job-relevant quality each was meant to approximate.
  • Run the four-fifths check at every stage. Compute selection rates by group at screening, interview, and offer rather than only at the final advance. Benchmark each stage against its own highest-rate group. Where a pool is small, record the count as well as the ratio so you know how sensitive the result is.
  • Audit your own overrides. Pull a sample of decisions where a recruiter departed from the tool's recommendation and a sample where they accepted it. Is there a documented reason in each case? Do the overrides cluster in one direction or around particular candidate backgrounds?
  • Ask who is not in the pool. Compare your applicant composition against the qualified labor market for the role. Then examine your sourcing channels, job description language, credential requirements, and shift or pay structures for anything that would explain a difference. Fix one structural cause and watch the pool rather than the screen.

Reflection

  • Could you answer the head of nursing's question today? What evidence would you produce, and how long would it take you to assemble it?
  • Which of the four sources is your organization currently least equipped to detect, and what would you need in order to see it?
  • What does your screening tool consume that you would struggle to justify as job-related if a candidate asked?
  • When your recruiters override the AI, what is the record of that decision, and would a pattern in those overrides be visible to anyone?
  • Which requirement on your current postings exists because it is essential to the job, and which exists because it has always been there?

Glossary

  • Disparate impact. Liability under Title VII arising when a neutral-looking practice produces substantially different outcomes for a protected group, with no intent to discriminate required.
  • Training data bias. The pattern a model learns when the historical record it was tuned on encodes past discrimination, so that resemblance to past hires becomes the target regardless of job relevance.
  • Proxy variable. An input that does not name a protected characteristic but correlates with one, such as a zip code, a name, an employment gap, or a specific school.
  • Algorithmic bias. Bias introduced by design choices rather than data: what the model optimizes for, which features it weights, and how it handles edge cases such as unfamiliar credentials or non-linear careers.
  • Automation bias. Trusting an AI ranking more than it deserves, scrutinizing top-ranked candidates less and skipping low-ranked ones entirely, so that any bias inside the model is laundered into "the data said so."
  • Selective override. Accepting the tool's recommendation when it matches your instinct and overriding it when it does not, which gives human bias the final vote while leaving a record that looks tool-assisted.
  • Systemic factors. Structural conditions upstream of any tool, including sourcing channels, job description language, credential requirements, and pay or scheduling structures, that shape who enters the applicant pool at all.
  • Four-fifths rule. The EEOC Uniform Guidelines screen for adverse impact: divide each group's selection rate by the highest group's rate and treat any ratio below 0.80 as a flag warranting investigation.
  • Selection rate. A group's advances divided by that group's applicants at a given stage, which is what the four-fifths rule compares rather than raw headcounts.
  • Automated employment decision tool (AEDT). The category of tool covered by NYC Local Law 144, which for jobs in NYC requires an independent bias audit within the prior year, a published summary, and candidate notice.
  • GDPR Article 22. The provision restricting decisions based solely on automated processing, generally requiring a human in the loop and the ability for the candidate to contest the decision.

Closing

The question Marcus could not answer was not unreasonable and it was not hostile. His head of nursing asked whether the tool was fair, and the honest answer at that moment was that nobody knew, because nobody had measured. That is the ordinary condition of a recruiting function that has adopted AI for speed, and speed was real: three days of triage became an afternoon. Nothing in this lesson argues for giving that back. It argues that a tool fast enough to change the shape of your funnel is a tool whose outputs you now have an obligation to watch.

Watching means all four layers. Ask what the model learned from and audit for proxies rather than trusting a removed field. Treat the algorithm as a black box and measure what comes out of it, because that is both what you can do and what the law increasingly requires. Discipline the human step with structured criteria and documented reasoning, since automation bias and selective override quietly return the final vote to intuition. And look upstream at who never applied, because a clean audit on the people who reached your screen says nothing about the people your sourcing, postings, and requirements filtered out first. Marcus owns the outcome either way. The only choice is whether he can describe it.

Key Takeaways

  • Bias comes from four compounding sources, not one. Biased training data, algorithmic design choices, human bias in the loop, and systemic factors upstream of the tool. Diagnosing fairness means checking all four, because a tool can be clean on three and still discriminate through the fourth, and because the layer determines where the fix belongs.
  • History in the data becomes bias in the model. A model trained to find "candidates like the ones we hired" reproduces past discrimination. The documented case of a resume-screening tool that learned to penalize women's resumes is the cautionary pattern: nobody intended it, nobody could have written the rule down, and nobody caught it until they audited.
  • Removing the protected field does not remove the bias. Proxy variables such as zip code, name, employment gaps, or the specific school listed stand in for protected characteristics, and a model optimizing for resemblance to past hires will find whichever ones remain. Audit outputs by group rather than reasoning about inputs.
  • You do not need to see inside the algorithm to hold it accountable. Treat the tool as a black box and audit its outputs. That is exactly what NYC Local Law 144 requires: an independent annual bias audit reporting selection rates and impact ratios by sex and race/ethnicity, published results, and candidate notice.
  • The four-fifths rule is the number that matters. Compare each group's selection rate to the highest group's. If the ratio falls below 0.80, that is a flag to investigate. In the worked example, a 14% versus 20% advance rate gives a 0.70 ratio and fails; eight more advances would have cleared it, showing how sensitive small pools are and why the ratio needs a count beside it.
  • Humans reintroduce bias the tool removed. Automation bias, selective override, and inconsistent interpretation let human bias get the final vote while leaving a record that looks tool-assisted. Structured criteria, documented reasoning, and treating AI as one input among several are the discipline that counters it.
  • Systemic bias happens before the AI runs. If sourcing, job descriptions, credential requirements, or shift and pay structures skew the applicant pool, a perfectly fair screen still yields an unrepresentative slate, and the affected candidates are missing from your data rather than visible in it. Ask who is not in the pool and why, and fix the upstream cause.
  • You own the outcome, legally and ethically. Title VII disparate impact, the ADA, the ADEA, and GDPR Article 22 all point to the same conclusion: measure fairness, document it, keep a human in the loop, and never treat "the vendor's tool decided" as a defense.

Frequently Asked Questions

Our vendor says the tool is bias-tested. Is that enough? No, for two reasons. First, the liability does not transfer: under Title VII an employer can be liable for a neutral practice that produces disparate impact, and "the computer did it" is not a defense, so Marcus and not the vendor owns the hiring decision. Second, a vendor's testing was performed on the vendor's population, not on your applicants, your requisitions, or your recruiters' override behavior. If the tool is an automated employment decision tool used for jobs in NYC, Local Law 144 requires an independent bias audit within the prior year, published results, and candidate notice, which is the standard worth applying regardless of where you operate.

We removed race, sex, age, and name from the data. Are we safe? Not on that alone. The model does not need a protected field to discriminate, because proxies carry the same information: a zip code, an employment gap, or the specific school listed can each stand in for a protected characteristic. The documented resume tool that learned to downgrade the word "women's" never had a gender field to consult; it inferred the pattern from correlated signals in the resumes it learned from. Stripping fields is a reasonable step and it is not a fairness result. The result comes from auditing outputs by group.

Our applicant pools are small. Does the four-fifths rule still work? Use it, and interpret it with the count in view. In the worked example, moving eight candidates changed a failing 0.70 into a passing 0.83, so with tiny pools a few decisions swing the ratio dramatically, and statistical significance matters alongside the raw ratio. Practically, treat a passing ratio on a small pool as weak evidence rather than a clean bill, and treat a failing one as a reason to examine what the screen weights rather than as noise to dismiss. The rule is a screening test, not the final word.

Our screen passed its audit but our hires are still not representative. What did we miss? Almost certainly systemic factors. A model can pass its four-fifths audit on the applicants it received and still leave you with an unrepresentative workforce, because the audit only measures selection among the people who applied. Sourcing channels that reach one demographic, exclusionary job description language, credential requirements without a clear job-related justification, and pay or scheduling structures that deter certain groups all filter the pool before the tool ever runs, and those candidates never appear in your data. Compare your applicant composition against the qualified labor market and look upstream.

Where should a recruiter start if all four sources are in play? Start with the measurement, because it tells you where to look next. Run the four-fifths check by group at each stage rather than only at the final advance; a stage that fails points you at a specific set of criteria to interrogate for proxies. In parallel, ask your vendor what the model was tuned on, since that answer alone often explains a training data pattern. Then examine your override records for direction and consistency, and compare your pool against the labor market. The order matters less than the discipline of identifying the layer before choosing the fix.