Hands-On Project: Audit a Recruiting Workflow for Bias
Camille leads recruiting for a 600-person retail-tech company, and she would have told you, honestly, that her hiring process was fair. Her team used scorecards. They cared about diversity. They had never knowingly screened anyone out for who they were. Then her company adopted an automated resume-screening tool, the kind New York City regulates as an automated employment decision tool, and Camille learned she would owe a published bias audit before she could keep using it. She could not certify a process she had never measured. So she did the thing the regulation actually requires: she pulled her own funnel data and looked. What she found was not a villain. It was a phone-screen stage that quietly advanced one group at a noticeably higher rate than another, with no job-relevant reason she could point to.
Why an Audit Beats Intuition
You can read about bias, understand it completely in the abstract, and still run a biased process, because intuition does not surface patterns that only appear in aggregate data. Bias in a recruiting workflow rarely looks like bias from the inside. No single decision feels unfair; each rejection has a plausible reason attached. The bias lives in the pattern across hundreds of decisions, and a pattern is invisible to anyone standing inside it. An audit grounds fairness in counted outcomes instead of felt fairness, points to the specific stage where disparity enters, and converts a vague aspiration to "be fairer" into a targeted fix you can implement and verify.
An audit is also how bias awareness becomes concrete action. It forces you to look at actual data, actual decisions, and actual outcomes rather than at your intentions, and the result is both a baseline and a roadmap. This project walks a complete audit end to end: application screening, interview process, evaluation rubrics, decision-making, and outcomes. Camille worked it in seven phases, and they are sequential because each one narrows the question the next has to answer.
Two legal frameworks give the audit its structure. The EEOC's four-fifths rule, which also underpins the bias-audit requirement of New York City's Local Law 144 for automated employment decision tools, provides the headline test: compare selection rates across groups, and if the disadvantaged group's rate falls below four-fifths, or 0.80, of the highest group's rate, that is evidence of adverse impact worth investigating. Title VII and the ADA frame what counts as a protected characteristic and what makes a criterion defensible. Camille's audit is built around the four-fifths test applied stage by stage, because a funnel-wide pass can still hide a single stage that fails.
Phase 1: Define the Workflow End to End
Camille writes out her funnel before she touches any numbers, because you cannot locate bias in a process you have not described. Her stages run: applications received, resume screen against job requirements, phone screen, technical interview, hiring-manager interview, offer decision, and offer acceptance or decline. Yours may differ, and the point of the exercise is to capture yours accurately rather than to match hers.
For each stage she records five things: how many candidates enter, how many advance, who evaluates, what criteria they apply, and how the decision is actually made. The last three are where fairness risk concentrates. A stage where one person advances candidates on instinct is a different risk than a stage with a shared written rubric, and the map has to capture that rather than flattening both into "phone screen." Camille also notes which criteria are written down and which are improvised, because an improvised criterion cannot be applied consistently by definition, and every stage she marked improvised became a candidate for investigation later.
Phase 2: Collect the Outcome Data
Camille pulls her last 200 candidates for a single high-volume role. A sample of roughly the last 100 hiring decisions is a workable floor, and mixing multiple roles muddies the analysis because different roles have genuinely different applicant pools. A reasonable sample size also keeps the selection rates from swinging on one or two individual people, which is the failure mode that makes small-sample audits worse than useless.
For every candidate in the sample she records the furthest stage reached or the stage at which they exited, plus whatever demographic data is legitimately available: race or ethnicity where the candidate disclosed it, gender where it is known, age where available, location, and educational background. Some of her candidates self-identified through the voluntary EEO survey, which is the cleanest source. For the rest she is honest about the limits. You cannot compel candidates to share demographics, and while name-based inference, location, and educational institution are sometimes used as proxy variables, they are weak proxies that produce real error, particularly on race and national origin.
Camille treats those proxies with caution rather than certainty and documents the limitation directly in the analysis rather than in a footnote nobody reads. An audit that overstates the confidence of its inputs is worse than no audit, because it produces a number people will act on without knowing how soft it is. Stating plainly that a portion of the demographic assignment is inferred, and roughly what portion, is what makes the finding usable by legal, by a regulator, or by the hiring managers whose process is about to change.
Phase 3: Run the Four-Fifths Analysis, Stage by Stage
This is the analytical core, so Camille works a real stage in full. At the phone-screen stage, 120 candidates advanced from the resume screen and entered the phone screen. She compares two groups. Group A: 60 candidates entered the phone screen and 30 advanced to the technical interview, a selection rate of 30 divided by 60, which is 50 percent. Group B: 50 candidates entered and 15 advanced, a selection rate of 15 divided by 50, which is 30 percent.
The four-fifths test compares the lower rate to the higher one. The highest group rate is Group A at 50 percent. The adverse-impact ratio is Group B's rate divided by Group A's rate: 0.30 divided by 0.50, which equals 0.60. The threshold is 0.80. Because 0.60 falls below 0.80, the phone-screen stage fails the four-fifths rule, and that is a flag, evidence of adverse impact that demands investigation. To make the contrast concrete, Camille runs the same test on her technical-interview stage, where Group A advances at 45 percent and Group B at 40 percent. There the ratio is 0.40 divided by 0.45, which equals roughly 0.89, comfortably above 0.80 and passing. Side by side, the two results do exactly what a stage-by-stage audit is supposed to do: they isolate the phone screen as the stage carrying the disparity, while clearing the technical interview.
The shape of the results across all stages tells you something before you investigate any single one. If every stage shows a similar disparity, that pattern suggests something systematic, entering early and propagating, perhaps in sourcing or in the applicant pool itself. If the disparity appears sharply at one stage and not the others, that stage is where the mechanism lives and that is where the investigation goes. Camille's funnel showed the second pattern, which is why she could focus her time rather than boiling the whole ocean.
One discipline matters here above all others. A failed ratio is a flag, not a verdict. It says investigate this stage, not "this stage is biased." A finding that 50 percent of one group and 30 percent of another advanced from the phone screen is a finding, and the very next question is whether a genuine job-relevant factor explains it. The four-fifths rule surfaces adverse impact; it does not by itself prove the cause is unlawful. The next phase is where Camille tests whether a job-relevant factor explains the gap or whether bias does.
Phase 4: Audit the Process at the Flagged Stage
Camille focuses her investigation on the phone screen, the stage that failed. She asks who runs it, what criteria they use, how decisions get made, and whether those criteria are applied consistently across candidates. Then she reads the evidence directly rather than reasoning about it in the abstract: she pulls the screening notes for ten Group B candidates who were cut and ten Group A candidates who advanced, and looks for whether similar behavior was interpreted differently across the two sets.
The pattern she finds is subtle and familiar. Phone screens at her company were unstructured, every recruiter asked different questions in a different order, and "communication" was being graded partly on a conversational style that tracked with background rather than with the job. The same hesitation read as "thoughtful" in one candidate's notes and "unsure" in another's. That is the mechanism behind the failed ratio: not animus, but an unstructured stage where a non-job-relevant signal was quietly doing the sorting. Reading paired notes produced this finding, and no amount of staring at the ratio would have, which is why outcome and process analysis have to be done together.
Phase 5: Locate the Source of Bias
Camille tests each plausible source against the evidence rather than guessing, working through a standard list. Is the disparity in sourcing, with different groups arriving through different channels of differing quality? In screening, where criteria correlate with demographics even though they look neutral? In interviews, where the questions asked or the interpretation applied vary by candidate? In evaluation, where rating scales are applied unevenly? Or in decision-making, where certain groups are held to a different bar than others?
For each candidate source she looks for evidence rather than reasoning from plausibility, because several of these will always sound plausible. For her flagged stage the evidence points squarely at the interview and evaluation sources: an unstructured conversation scored on an inconsistent, partly stylistic notion of communication. Naming the source precisely makes the fix precise, because sourcing bias and evaluation bias call for completely different remedies, and a team that misdiagnoses will spend a quarter fixing something that was never broken while the real mechanism keeps running.
Phase 6: Develop Targeted Recommendations
A good recommendation names the mechanism and the change, not a sentiment. "Reduce bias in the phone screen" is useless because nobody can act on it and nobody can tell whether it was done. Camille writes implementable fixes instead. Replace the unstructured phone screen with a structured one: the same job-relevant questions in the same order for every candidate, scored against a written rubric tied to the role. Define "communication" by observable, job-relevant behaviors rather than conversational warmth. Add a brief calibration session so recruiters score the same answer the same way. Train the interviewers on confirmation bias specifically, since the paired-notes finding showed exactly that mechanism operating.
For the earlier resume screen she pilots blind first-stage review that masks name and institution, removing the weak proxies before judgment begins. Each recommendation gets a named owner and a date, because an audit ending in recommendations no one owns changes nothing, and "the recruiting team" is not an owner. The test is whether someone reading it a month later could tell you whether it happened. "Implement structured interviews with the same questions for all candidates, owned by the phone-screen lead, live by the start of next quarter" passes. "Improve fairness at the phone screen" does not.
Phase 7: Document and Re-Audit
Camille writes the audit report so it stands on its own: the process map, the demographic makeup of the candidate pool at each stage, the advancement rates and adverse-impact ratios stage by stage, the disparity identified at the phone screen, the hypothesized source with the supporting evidence from the paired notes, and the recommendations with owners and dates. For an automated employment decision tool under Local Law 144, this kind of documented, published bias audit is not optional, but the discipline is worth it for any workflow.
Documentation does double duty. It is legal protection, because it demonstrates that you took fairness seriously, measured it, and acted on what you found, a materially different posture from having never looked. And it is an operational roadmap, because it tells you and your successor what was broken, what was tried, and what remains. Publishing findings internally adds a third effect: when an organization audits, measures, and shares results, fairness becomes tracked and managed rather than assumed, and that accountability is often what makes the recommendations actually get implemented.
Then she commits to the part that closes the loop. After the structured phone screen has run for a quarter, she re-runs the same four-fifths analysis on fresh data to see whether the ratio moved back above 0.80. Audits are iterative by nature: audit, improve, audit again, and find out whether the intervention worked. An audit is a baseline, and an intervention is only proven by the next measurement. If the ratio has not moved, the diagnosis was wrong or the fix was not implemented as written, and both of those are worth knowing before another year of hiring passes.
Turning Findings Into Change
The gap between a finished audit and a changed process is where most of this work dies, so Camille runs the implementation as its own sequence. It starts with a current-state assessment broader than the single failing stage. She maps sourcing practices, the screening process, the interview approach, the evaluation method, the decision framework, documentation practices, and candidate communication, and for each asks four questions: is it systematic, is it fair, is it effective, and is it compliant? The audit answered those questions for one stage; the map tells her which other stages she is running on assumption.
Next comes priority setting, because attempting everything guarantees nothing lands. She ranks by three criteria. Impact: what would most improve outcomes, which for her is fairer evaluation, since that is where the disparity appeared. Feasibility: what is realistic now, which is often better documentation, because it requires a process change rather than new technology or budget. Leverage: what is a prerequisite for other improvements, usually the systems support that makes new practices sustainable rather than heroic.
Then she pilots rather than transforms. The structured phone screen goes live with one team and one role first, results get measured, the rubric gets refined on what actually happened, and only then does it expand. Piloting produces learning that planning cannot: you discover what works in your context, refine on real experience, and build buy-in through a success story rather than a mandate. Measurement runs alongside, and its purpose is to learn rather than to prove she was right.
The final step is sustainability. Once a practice works in pilot it has to be scaled into the ordinary machinery of the organization, built into systems, training, and job descriptions, so it becomes standard practice rather than a special project that decays when its champion changes roles. Scaling needs four things: training, because people need to know how to do this; systems, because the technology has to support it rather than fight it; accountability, because expectations have to be explicit and owned; and ongoing attention, because leadership focus is what keeps a practice alive after the novelty wears off.
Fairness Is a System, Not a Practice
Recruiting practices do not operate in isolation, and the fixes from a bias audit work considerably better in combination than alone. Structured interviews, which are an evaluation practice, work best combined with blind resume screening, which is a fairness practice, and with diverse hiring panels, which is an inclusion practice. Documentation, a process practice, is what makes fairness audits possible in the first place, which is why Camille's inability to audit before she started was really a documentation failure wearing a fairness costume. When you implement any single practice, the useful question is how it fits with the others rather than whether it is good in isolation.
That systems view also explains where bias can enter, which is everywhere. Sourcing should identify diverse candidates, screening should treat them consistently, interviews should assess against the same criteria, evaluation should be structured enough to prevent drift, and decision-making should be documented and auditable. Each element contributes to fairness, and each can introduce bias if done carelessly, which is why a single fix at a single stage is a beginning rather than a conclusion. Camille's phone-screen remedy closed one door; the current-state map told her which doors she had not yet checked.
The last principle is that none of this stays finished. Research changes what is known to work, new tools enable new approaches, legal requirements shift, and candidate expectations move. Effective recruiting requires continuous learning built into the practice: staying current on relevant research and legal requirements, experimenting deliberately with new tools, and gathering candidate feedback on their actual experience. This does not mean constant change, which is its own dysfunction. It means systematic attention to what is working and what could be better, on a cadence you can sustain.
Hiring Managers, Resistance, and Scale
Hiring managers determine whether a bias remedy survives contact with reality. They drive sourcing, participate in interviews, and make final decisions, so a structured phone screen they consider bureaucratic overhead will quietly erode within two quarters. Camille explains four things directly: why the change is happening, what the benefit to them is, how it improves their own ability to identify strong candidates rather than merely constraining them, and what specifically is expected of them. She also involves them in designing the rubric rather than presenting a finished one, addresses their objections on the merits, and shows them the measurement once it exists.
Resistance is normal and worth planning for rather than resenting. People like how things currently work, they are skeptical of new approaches, and they are reasonably worried about additional work landing on them. Five things address it: start small so there is a success to point at, involve people in the design so they have ownership, show data that the change works, make the practice easy by building it into systems they already use, and secure visible leadership support. Four of the five are things you do before the resistance appears.
Scaling beyond one team introduces variation you have to design for. Different regions may need different approaches, different roles genuinely have different requirements, and different hiring managers need different amounts of support. Hold the base practice consistent, structured questions scored against a written rubric, while allowing implementation to flex by context. A framework rigid enough to be identical everywhere tends to be abandoned everywhere, which is a worse fairness outcome than a flexible framework people actually follow.
Anti-Patterns That Undermine the Audit
Three habits defeat an otherwise sound bias audit. The first is reading data without context: you compute a disparity and declare bias without investigating whether a genuine job-relevant factor explains it. It happens because running numbers is easier than investigating, and what goes wrong is that you conclude bias where qualifications actually explain the gap, destroying your credibility for the findings that are real. The defense is that every disparity gets investigated for job-relevant explanation, and every failed ratio is a flag rather than a verdict.
The second is auditing outcomes without auditing process: you look at who got hired without examining how the decisions were made. It happens for the same reason, because outcomes are easier to count than decisions are to read. What goes wrong is that you identify a problem without understanding its cause, leaving you with a number and nothing to fix. The defense is to audit both, outcome for the disparity and process for the mechanism, which is what Camille's paired-notes review supplied.
The third is the audit without action: you measure, you write the report, the report feels like progress, and nothing changes, so the disparity persists into the next cycle untouched. This is the most common of the three precisely because completing an audit feels like having done something. The defense is to plan implementation alongside measurement, with owners and dates attached before the report circulates, so the audit ends in a change rather than a document.
Practice
Work these in order on your own funnel. Each one produces an input the next one needs.
- Map your process. Document your workflow stage by stage: how many candidates enter and advance, who evaluates, against what criteria, and how the decision is actually made. Mark which criteria are written and which are improvised.
- Collect the data. Pull your last 100 hiring decisions for a single role, recording the stage reached or exited, the demographic data available, and the notes if they exist. Write down which demographic assignments are self-reported and which are inferred.
- Run the disparity analysis. Calculate the advancement rate for each group at each stage, then divide the lower rate by the highest. Flag every stage below 0.80 and note whether the pattern is spread across stages or concentrated in one.
- Investigate the source. For each flagged stage work the list: sourcing, screening, interviews, evaluation, decision-making. Pull the notes for ten candidates who advanced and ten who did not, and look for the same behavior read differently.
- Write the recommendations. For each identified source, write a specific implementable change with a named owner and a date. Test each by asking whether a reader could tell a month later whether it happened.
- Schedule the re-audit. Put the follow-up measurement on the calendar now, one quarter after the intervention goes live, using the same method on fresh data.
Reflection
- What would a bias audit of your recruiting actually reveal, if you had to guess before running it?
- Where in your process do you already suspect bias appears, and what makes you suspect it?
- What data would you need to test that suspicion, and how much of it can you obtain today?
- If your audit found disparities, what would the most likely sources be given how your stages are actually run?
- What is one change you would make on the strength of audit findings, and who would own it?
- Which of your stages are unstructured enough that two evaluators could reach opposite conclusions about the same candidate?
Glossary
- Advancement rate. The percentage of candidates who move from one stage of the funnel to the next, calculated separately for each group being compared.
- Adverse-impact ratio. The lower group's advancement rate divided by the highest group's rate. Under the four-fifths rule, a result below 0.80 is evidence of potential adverse impact requiring investigation.
- Demographic parity. Similar advancement rates across demographic groups at a given stage.
- Disparate impact. When facially neutral criteria produce materially different outcomes by demographic group, whether or not anyone intended that result.
- Process audit. Examining how decisions are made at each stage, including who decides and against what criteria, rather than only counting outcomes.
- Source analysis. Identifying where in the process bias is entering: sourcing, screening, interviews, evaluation, or decision-making.
- Automated employment decision tool. The category of tool regulated by NYC Local Law 144, which attaches independent bias audit, public posting, and candidate notice obligations to its use.
Related Lessons
This project draws on several lessons and feeds directly into several others.
- Analyzing Your Own Workflows: Where Could Bias Hide? is the natural precursor, since it develops the process map that Phase 1 of this audit depends on.
- Fairness Metrics: Defining and Measuring Bias in Outcomes goes deeper on the measurement side, including which metrics complement the four-fifths ratio when a single number is not enough.
- Root Cause Analysis: Understanding Why Bias or Errors Occurred extends the Phase 5 source investigation with a more rigorous method for testing hypotheses against evidence.
- Fairness Interventions: Blind Reviews, Structured Processes, Diverse Panels covers the remedies Phase 6 recommends, including how to implement blind first-stage screening without breaking your ATS workflow.
- Structured Evaluation: Avoiding Halo Effects and Confirming Bias addresses the specific mechanism Camille found in her paired-notes review, where the same behavior was read differently depending on the candidate.
- Sources of Bias: Data, Algorithms, Humans, and Systemic Factors supplies the fuller taxonomy behind the five-source list used in Phase 5.
Closing
Bias audits are how fairness stops being an intention and becomes something measured and improved. Camille's process was fair in every way she could observe from the inside, and it still failed the four-fifths test at a stage she had never thought to examine, because unstructured phone screens were quietly grading conversational style. She did not find a villain and did not need one. She found a mechanism, and mechanisms can be fixed, which is why the audit is more useful than the conviction that your team means well.
The discipline that makes it work is unglamorous. Map the process before you count anything. Collect a real sample and be honest about the limits of your demographic data. Run the ratio stage by stage rather than across the funnel. Treat every failure as a flag and go read the actual decisions. Name the source precisely, write fixes with owners and dates, document everything, then measure again to find out whether you were right. Audit, improve, audit again.
Key Takeaways
- Bias hides in aggregate patterns, not individual decisions. Each rejection has a plausible reason, so intuition will not surface a disparity that only appears across hundreds of outcomes. Your sense that the process is fair is unvalidated until data validates it.
- Map the process before you measure it. Record how many enter and advance at each stage, who evaluates, what criteria apply, and how the decision is made. Improvised criteria cannot be applied consistently by definition.
- Be honest about your demographic data. Use voluntary self-identification where you have it, treat name, location, and institution as weak proxies rather than facts, and document that limitation in the analysis. An audit that overstates its confidence is worse than none.
- Apply the four-fifths rule stage by stage. Divide the lower group's selection rate by the highest group's rate; a ratio under 0.80 is a flag for adverse impact. A funnel-wide pass can still hide a single failing stage.
- Get the math right. When Group A advances at 50 percent and Group B at 30 percent, the ratio is 0.30 divided by 0.50, which equals 0.60, below the 0.80 threshold and therefore flagged, while a stage at 0.89 passes.
- A failed ratio is a flag, not a verdict. The four-fifths rule, which anchors the bias-audit requirement under NYC Local Law 144 for automated employment decision tools, surfaces adverse impact; you still have to investigate whether a job-relevant factor explains it. Not every demographic difference indicates bias, but every one requires investigation.
- Read the decisions, not just the outcomes. Pull the notes for ten candidates who advanced and ten who did not at the flagged stage and look for the same behavior interpreted differently. Outcome analysis finds the problem; process analysis finds the cause.
- Disparity at one stage points to the source. Test sourcing, screening, interviews, evaluation, and decision-making against evidence. An unstructured phone screen scored on a partly stylistic notion of communication is an evaluation-and-interview source, and naming the source precisely makes the remedy precise.
- Fix mechanisms, not sentiments. Same job-relevant questions in the same order, a written rubric, communication defined by observable behaviors, calibration sessions, confirmation-bias training, and blind first-stage screening, each with a named owner and a date.
- Implement in stages. Assess the current state across every stage, prioritize by impact, feasibility, and leverage, pilot with one team or role, measure to learn rather than to be proven right, then scale with training, systems, accountability, and sustained leadership attention.
- Documentation is both legal protection and operational roadmap. It shows you took fairness seriously and measured it, and it tells your successor what was broken and what was tried. Publishing findings turns fairness into something tracked and managed rather than assumed.
- Re-audit to prove the fix. An audit is a baseline; an intervention is only validated when the next four-fifths analysis on fresh data shows the ratio back above 0.80. If it has not moved, either the diagnosis or the implementation was wrong.
Frequently Asked Questions
What if I do not have reliable demographic data on my candidates? Work with what you legitimately have and be explicit about the gaps. Voluntary EEO self-identification is the cleanest source, and it is worth improving your response rate on it before your next audit. Beyond that, proxies such as name-based inference, location, and educational institution are used in practice but they are weak and produce real error, particularly on race and national origin. Use them with caution, state which portion of your assignments are inferred, and treat a marginal ratio built on proxy data as a reason to gather better data rather than a finding you would defend publicly.
Does a ratio below 0.80 mean we are breaking the law? No. The four-fifths rule surfaces evidence of adverse impact; it does not establish that a practice is unlawful. The finding obligates you to investigate whether a genuine job-relevant factor explains the difference, and if one does, to be able to show that the criterion is actually job-related and consistent with business necessity. What you cannot do is compute the number, notice it is below the threshold, and proceed as though you never looked. That is the posture that turns a manageable finding into a serious problem.
How large does my sample need to be? Large enough that a single candidate does not move the rate meaningfully. Roughly the last 100 hiring decisions for one role is a workable floor, and Camille used 200. Keep the analysis to a single role or genuinely comparable roles, because mixing a high-volume retail role with a specialized engineering role blends two different applicant pools and produces a ratio describing neither. If a stage has so few candidates in one group that the rate swings on two people, report that limitation rather than the ratio.
We are using an automated screening tool. Does this project satisfy Local Law 144? It does not replace the legal requirement. Under NYC Local Law 144, an automated employment decision tool requires an independent bias audit conducted within the past year before use, public posting of a summary of the results, and notice to candidates that an automated tool is being used. Independent means not the vendor and not your own recruiting team. This project is the internal discipline that makes you ready for that audit and that keeps you honest between audits, and it also covers the human stages the legally mandated audit of the tool will not touch.
The disparity appears at every stage rather than one. What now? That pattern suggests something systematic entering early and propagating rather than a single broken stage, so look upstream first. Sourcing is the usual candidate: if different groups arrive through channels of differing quality or fit, every subsequent stage will inherit that difference and appear to fail independently. Check whether the composition of your applicant pool differs by channel, and whether your resume screen encodes criteria that correlate with the channel rather than with the job. Fixing an interview stage will not help if the disparity was already present when candidates entered it.
The hiring managers think the structured phone screen is bureaucratic. How do I get it to stick? Address it before it becomes resistance. Explain why the change is happening and what it does for them, which is a more reliable read on candidates rather than a constraint on their judgment. Involve them in writing the rubric so it reflects what they actually care about, pilot with one willing team so you have a success story rather than a mandate, build the questions into tools they already use, and get visible leadership support. Then show them the re-audit numbers. A practice people helped design and can see working is the one that survives the quarter.
Skill.re