Fairness Checks: Identifying Gender, Age, Disability, and Other Bias Signals
Marcus runs talent acquisition for a 220-person fintech in Austin, with a team of four recruiters filling roughly 40 roles a year. Six months into using AI to summarize phone-screen notes and draft candidate assessments, he noticed something while comparing two write-ups for the same engineering role. A 27-year-old bootcamp graduate came back as "eager and coachable, will grow into the role." A 54-year-old with 25 years of experience came back as "may be overqualified and looking for a stepping stone." Both had passed the same technical screen with nearly identical scores. The AI had not measured anything different about them - it had narrated his pipeline back to him with the field's oldest stereotypes baked in, and it had done so in calm, confident, professional-sounding prose. That is the danger. Biased AI output does not announce itself. It reads like a competent assessment, which is exactly why it slips into hiring decisions unchallenged.
Why AI Output Is Not Neutral
The most expensive assumption a recruiter can make is that an AI assessment is objective because a machine produced it. AI models learn from enormous bodies of human text, including decades of hiring documents, performance reviews, and online writing that encode the exact biases hiring law exists to prevent. When the model writes a candidate summary, it is predicting the words that usually follow in text like the text it trained on. If "older worker" was historically followed by "overqualified" and "set in their ways," the model has learned that association. It is not reasoning about Marcus's candidate; it is completing a pattern.
It is worth being precise about where your responsibility sits. Eliminating bias from the model itself is the vendor's problem, and no recruiter can solve it from the outside. Catching bias in the output before it touches a hiring decision is your problem, and it is entirely solvable. Bias also compounds in both directions: ask a biased question and the model will hand back a biased answer with more confidence than you put into the question, which is why the prompt and the review are two halves of the same discipline.
This matters legally, not just ethically. In the United States, Title VII, the Age Discrimination in Employment Act (ADEA, which protects workers 40 and over), and the Americans with Disabilities Act (ADA) prohibit hiring decisions that disadvantage protected groups. The EEOC has made clear that an employer is responsible for discriminatory outcomes even when a vendor's software produced them. The model is the vendor's problem to train. The decision is yours to defend. If a biased AI summary steers a rejection, "the AI wrote it" is not a defense - the recruiter who acted on it owns the outcome.
Marcus's job, then, is not to fix the model. It is to catch the model's bias before it reaches a hiring decision. That requires knowing what biased output looks like, because the patterns are consistent and learnable.
The Five Bias Signals to Watch For
1. Gender coding. The clearest tell is the same behavior described in different language depending on gender. The AI calls a man's directness "assertive and decisive" and a woman's identical directness "abrasive" or, flipped, calls a man "ambitious" and a woman "detail-oriented and thorough." Watch also for gendered language in the prose itself, the same vocabulary problem covered in the prompting chapter, and for unprompted assumptions about home or family that attach only to women. The subtler version is the same observation given two different interpretations: a man who "asked few questions in the interview" is described neutrally, while a woman who did exactly the same thing "did not seem engaged in the interview discussion." One reading implies thoughtfulness, the other implies disengagement, and nothing but the stereotype separates them. When Marcus had the model write parallel summaries for two candidates who both pushed back on a salary range, the man "negotiated confidently" and the woman "came across as difficult." Same action, two genders, two verdicts.
2. Age signals. Age bias hides in framing. Early-career candidates are perpetually "eager to learn and grow"; experienced candidates are perpetually "overqualified," "set in their ways," or "may not keep up with the pace." Underneath sits a cluster of assumptions about energy level, learning ability, and motivation that are pinned to career stage rather than to anything the candidate demonstrated, along with assumptions about whether someone will "fit in" that rest on generational markers. Career length, graduation years, and dated technologies become proxies for age. Under the ADEA, treating "overqualified" as a polite synonym for "too old" is exactly the inference the law forbids. The honest question is whether the candidate can do the job, not whether their resume implies a birth year.
3. Disability assumptions. This is the subtlest and the most legally loaded. AI flags an employment gap from 2019 to 2020 and speculates the candidate "may not be serious about returning to work," or treats a mention of an accommodation as a reliability concern, or raises a medical detail that has no bearing on the job. The ADA prohibits employers from making disability-related inquiries or medical assumptions before a conditional offer. A gap could be caregiving, layoff, travel, education, illness, or anything else, and none of it is the AI's business to diagnose. Any time the output turns a gap, a lifestyle factor, or a health signal into a job-relevant worry, that is a fairness failure and an ADA exposure.
4. Cultural and national-origin signals. Bias here keys off names, accents, and presumed background. The AI writes "candidate speaks with an accent, may struggle to communicate clearly" or downgrades someone whose phrasing reads as non-native, or infers a work approach and communication style from where a person is presumed to be from. An accent is not a communication deficit, and national origin is a protected characteristic under Title VII. Watch for any judgment that conflates how someone sounds with whether they can do the work.
5. Class and education bias. The output overweights pedigree - "didn't attend a top-tier school," "came through a coding bootcamp rather than a CS degree, may lack fundamentals" - and treats non-traditional paths as evidence of lower capability, lower ambition, or less seriousness. Fundamentals are learnable and demonstrable through a technical screen; the path someone took to learn them is not a measure of whether they have them. This bias is not always tied to a single protected class, but it routinely correlates with race and socioeconomic status, which is how it becomes a legal problem as well as a talent-pool problem.
Each signal has a matching detection habit, and they are quick enough to run in the margins of a review. For gender, read the model's description of two candidates of different genders doing the same thing and ask whether the language and its implications differ. For age, look for whether early-career people are always "eager" and senior people are always "seeking something else." For disability, look for judgments about employment gaps or lifestyle factors that are not job-relevant. For cultural signals, look for judgments about communication that conflate an accent with an ability. For class and education, look for capability being judged by where someone learned rather than by what they know.
A Worked Example: The Four-Fifths Rule and Local Law 144
Patterns in a single summary are easy to argue about. Patterns across a pipeline are measurable, and that is where the law gets specific. The EEOC's four-fifths rule (the 80% rule) is the standard test for adverse impact: compare the selection rate of each group to the selection rate of the most-selected group, and if any group is selected at less than 80% of that rate, you have a presumptive adverse-impact problem worth investigating.
Marcus runs the numbers on the AI-assisted screen his team used over the last two quarters. The tool helped advance candidates from phone screen to onsite.
- Candidates under 40: 60 screened, 30 advanced. Selection rate 50%.
- Candidates 40 and over: 40 screened, 12 advanced. Selection rate 30%.
The most-selected group is the under-40 group at 50%. Four-fifths of 50% is 40%. The 40-and-over group advanced at 30%, which is below that 40% threshold - their selection rate is 60% of the favored group's rate (30 divided by 50), well under the 80% floor. That is a textbook adverse-impact flag against a group protected by the ADEA, and the "overqualified" language Marcus kept seeing in the older candidates' AI summaries is a plausible mechanism for it.
If Marcus's fintech were hiring in New York City, this would be more than a best practice. NYC Local Law 144 requires that automated employment decision tools (AEDTs) used to screen candidates undergo an independent bias audit within the prior year, that the audit results be published, and that candidates be notified the tool is in use. The audit math is built on exactly this kind of selection-rate comparison. The lesson generalizes: individual summaries tell you where to look, but the four-fifths calculation tells you whether the bias is moving real outcomes. Both belong in the workflow.
The Fairness Checklist
Marcus turned the five signals into a checklist his team runs on every AI-generated assessment before it enters the applicant tracking system. It takes about a minute per summary and it forces the question that biased prose is designed to suppress: would this sentence survive if the candidate's protected characteristic were different?
- Gender: Is the language gendered? Are identical behaviors described differently for men and women? Are there unprompted assumptions about family or availability?
- Age: Is "overqualified" or "eager" doing the work that the technical evidence should be doing? Are graduation years or career length being used as proxies for fit?
- Ability: Is an employment gap or health signal being turned into a job-relevant concern? (Under the ADA, it should not be.)
- Cultural and origin: Is accent, name, or presumed background being conflated with capability or communication skill?
- Class and education: Is capability being judged by where someone learned rather than what the screen shows they know?
- Motivation and fit: Are claims about motivation or "fit" resting on a protected characteristic rather than on observed behavior?
- Personality and communication style: Is a personality difference being treated as a job concern? Quietness, formality, or a different conversational rhythm are differences, not deficits, unless the role genuinely requires otherwise and the evidence shows it.
- Other: What other bias is present that these categories do not name? Leave this line open on purpose, because the bias you have not learned to look for is the one that reaches the decision.
- Evidence test: For every conclusion, is there a specific, job-relevant observation behind it? If not, it is a guess, and guesses are where bias lives.
Three Ways to Fix Biased Output
Catching bias is half the job. Here are three responses, in order of how durably they solve the problem.
Option 1: Revise the output directly. The fastest fix for a single bad line. Strip the biased framing and keep only the job-relevant observation. Marcus's "candidate didn't ask many questions, seems disengaged" became "candidate asked focused follow-up questions about system design and deployment." Same interview, evidence instead of inference. Use this when the problem is local and you are about to file the summary anyway.
Option 2: Ask the AI to revise with guidance. When the whole summary is tilted, send it back with a specific instruction rather than editing line by line. Marcus uses a tool like ChatGPT or Anthropic's Claude and gives a directive prompt: "Rewrite this assessment using only job-relevant evidence. Remove any inference about age, gender, family status, health, accent, national origin, or educational pedigree. For each conclusion, cite the specific behavior or qualification it rests on. If there is no evidence for a claim, drop the claim." A capable model handles this well because you have told it precisely what to remove and what standard to meet.
Option 3: Set better constraints upfront. The durable fix. If the same bias keeps surfacing, the prompt that generates the summaries is the problem, not any individual output. Marcus added a standing constraint to his team's assessment prompt: "Assess only against the role's stated requirements. Do not infer motivation, personality, or fit from communication style, accent, employment gaps, career length, or educational path. Do not assess capability based on educational path or background. Cite specific evidence for every claim. Do not comment on anything not relevant to performing this job." Fixing the prompt fixes every future summary at once instead of one at a time.
Three Anti-Patterns to Avoid
Assuming AI is neutral. Because output is machine-generated, it feels objective, so it gets less scrutiny than a human reviewer's notes would. That is backwards. Apply at least the same skepticism to AI output that you would to a hiring manager's gut-feel email, because the model has absorbed more biased text than any single person ever could. Neutrality is something you verify, not something you assume from the source.
Normalizing biased language. Subtle bias is dangerous precisely because it is easy to read past. "Overqualified," "culture fit," "not a strong communicator" all sound professional, and after seeing them fifty times they stop registering as anything at all. Once you stop noticing, you have normalized discrimination and trained yourself to ship it. The checklist exists to make the familiar phrase pause-worthy again, and it works best when you are specific about which bias you are hunting for in a given pass.
Fixing individual biases without fixing the process. Marcus could edit every biased summary by hand forever and never improve anything, because the prompt would keep producing them and a busy week would let one through. Patching outputs is not the same as preventing them. Every time you catch a recurring bias, push the fix upstream into the prompt constraints so the next hundred summaries inherit the correction. Otherwise you are doing manual quality control on a process you have chosen not to repair.
Practice Exercises
These five exercises turn the checklist from something you have read into something you can run under time pressure.
- Identify bias in real output. Take an AI summary or assessment you have already filed and walk it through the fairness checklist line by line. Note every piece of biased language or assumption you find, including the ones you would have waved through last month.
- Compare parallel descriptions. Ask the model to describe two hypothetical candidates doing exactly the same thing, varying only one characteristic at a time: one man and one woman, one early-career and one senior, and so on. Read the two descriptions side by side and ask whether the language or the implications differ.
- Revise biased output. Take a biased passage and rewrite it so the bias is gone and the job-relevant information survives intact. The test of a good revision is that the hiring manager loses nothing they could legitimately have used.
- Improve your prompt. When the same bias keeps reappearing, add a specific constraint to the prompt that generates the summaries, then rerun an old set of notes through it and check whether the constraint held.
- Build your own checklist. Write down the biases you personally see most often in your tool's output and keep that list beside the general one. Run it on every AI-generated assessment until you no longer need to look at it.
The Vocabulary of Bias Checking
Naming a bias precisely is most of the work of catching it, and these five terms cover the patterns in this lesson.
- Gender coding. Language that is implicitly masculine or feminine, leading to gendered interpretations of the same behavior.
- Age bias. Assumptions about capability, motivation, or fit based on career stage rather than actual qualifications.
- Disability bias. Negative assumptions based on health, medical history, or accommodations a candidate might need.
- Cultural and origin bias. Assumptions based on accent, name, presumed background, or country of origin.
- Class and education bias. Assumptions about capability based on educational pedigree or a non-traditional path into the field.
Fairness as Practice, Not Aspiration
Fairness is not a nice-to-have in hiring. It is a legal requirement and a competitive advantage at the same time. When you systematically remove bias from AI-assisted recruiting, you widen the pool of people who can reach your interview stage, and a wider pool with the same bar produces better hiring outcomes rather than looser ones. The organizations that treat this as compliance overhead lose the candidates that the organizations treating it as a talent strategy will hire.
Make it concrete this week. Run the fairness checklist on five AI outputs, write down every bias you find rather than fixing it silently, and then look at the five sets of notes together. The pattern that emerges is your real finding, and it tells you which prompt constraint to add so the same bias does not survive into next month's summaries.
Reflection
- Which type of bias are you most likely to miss in AI output, and why is that the one that gets past you?
- If you reviewed your hiring from the past year with this fairness checklist in hand, what biases do you think you would find?
- How would your hiring be different if bias were systematically removed from every AI-generated assessment your team touches?
Related Lessons
- Verification Techniques: Spot-Checking Facts, Sources, and Candidates is the natural next step. Bias checking asks whether the output is fair; verification asks whether it is true, and an assessment has to clear both bars before it informs a decision.
- Avoiding Bias in Prompts: Language, Examples, and Assumptions covers the upstream half of this work. Gendered language and loaded framing in the prompt reliably produce gendered, loaded output, which is why the durable fix in this lesson lives in the prompt.
- How AI Can Perpetuate or Amplify Bias explains the mechanism behind the patterns you are checking for, including how historical hiring data becomes a model's default assumption about who succeeds.
- Flagging Red Flags and Concerns Without Bias applies the same discipline to the specific moment where bias does the most damage, the point where a concern is written into a candidate's record.
- Documenting Decisions: Clear Records for Legal and Fairness Review shows what to keep once you have caught and corrected bias, so that the correction itself is part of a defensible record.
Key Takeaways
- AI output is not neutral, and you own the decision. Models learn hiring's historical biases from their training data and narrate them back in confident prose. The EEOC holds the employer responsible for discriminatory outcomes even when a vendor's tool produced them - "the AI wrote it" is not a defense.
- Five signals cover most AI bias. Gender coding, age signals, disability assumptions, cultural and national-origin signals, and class and education bias. The common tell across all five is the same behavior or qualification being described differently depending on a protected characteristic.
- Measure adverse impact, not just vibes. The four-fifths (80%) rule turns a hunch into a number. When the 40-and-over group advances at 30% against the under-40 group's 50%, that 60% ratio is below the 80% floor and flags an ADEA adverse-impact problem worth investigating.
- Know the law that applies to your tools. Title VII, the ADEA, and the ADA govern outcomes; the ADA specifically bars turning employment gaps or health signals into pre-offer concerns. In NYC, Local Law 144 requires an independent bias audit and candidate notification for automated screening tools.
- Run a one-minute checklist on every assessment. For each conclusion, ask whether there is a specific, job-relevant observation behind it. No evidence means it is a guess, and guesses are where bias lives.
- Fix at the highest level you can. Revise a single line when it is local, ask the AI to rewrite with explicit guidance when the whole summary is tilted, and rewrite the prompt constraints when the same bias keeps recurring. Prompt-level fixes correct every future output at once.
- Do not normalize what you stop noticing. "Overqualified" and "not a culture fit" sound professional and become invisible with repetition. Patching individual outputs without fixing the process guarantees you repeat the mistake - push every correction upstream.
Skill.re