←
AI for Recruiters
Visionary · M20 · lesson 20 of 30 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Policies for AI Use: What Should Be Required, Prohibited, or Encouraged?

15 min

Marcus runs talent acquisition for a 600-person fintech company with a recruiting team of nine. Last quarter his team adopted four AI tools in eight weeks: a resume screener, an outreach drafting assistant, an interview-notes summarizer, and a sourcing research helper. Nobody asked permission, and nobody wrote anything down. When Legal asked Marcus a simple question, "What is our policy on AI in hiring?" he did not have an answer. He had habits. This lesson is the document Marcus wrote in response: a one-page AI-use policy that sorts every practice into three buckets, required, prohibited, and encouraged, and ties the hard rules to the laws his team actually has to follow.

Why Three Buckets Beat a Wall of Rules

Marcus's first instinct was to write a comprehensive twelve-page policy. He abandoned it after a day. A policy nobody reads protects nobody. The three-bucket structure, required, prohibited, and encouraged, works because it answers the only three questions a recruiter actually has in the moment: What do I have to do? What can I never do? What am I free to try? Everything else is commentary.

The buckets also map cleanly to risk. Required practices are the controls that keep the team compliant and defensible. Prohibited practices are the bright lines that, if crossed, create legal exposure or harm a candidate. Encouraged practices are the productivity wins that carry little downside. When Marcus sorted his team's four tools this way, the policy wrote itself, and it fit on one page. The goal is not to slow the team down. It is to make the safe path the obvious path.

The hard part of policy writing is balance, and it is a genuine trade-off rather than a matter of getting the wording right. Too restrictive, and people ignore the policy or route around it, which leaves you with underground practice and no visibility. Too permissive, and you are not managing risk at all, you are documenting that you chose not to. Many organizations solve this by writing something strict and then never enforcing it, which produces the worst outcome of the three: policy as theater, where the document exists, everyone knows it is not real, and the organization has a written standard it is visibly failing to meet. The target is policy with teeth, strict enough to manage the risks that actually exist in your process and realistic enough that people follow it because it makes sense.

Required: The Non-Negotiables

Required practices apply to any AI tool that touches a hiring decision, with no exceptions and no tool-by-tool negotiation. Marcus attached a reason to each one so his recruiters would internalize the rule rather than resent it.

Human review of every AI screening decision. No candidate is advanced or rejected on the tool's output alone. A recruiter reads the resume and the AI's reasoning, then makes the call and can override it. This is the single most important control on the page. Under New York City Local Law 144, an automated employment decision tool used to screen candidates in NYC must undergo an independent bias audit within the prior year, and candidates must receive notice at least ten business days before the tool is used. Keeping a human in the loop does not exempt a tool from the audit requirement, but it does prevent the team from drifting into fully automated decisions that draw the heaviest scrutiny. Human reviewers will not, and should not, override most of the time; what they provide is the chance to catch the anomaly, the error, or the edge case the model missed.

Fairness testing before deployment. Before any new tool is used in recruiting decisions, it is tested for fairness by comparing its output across demographic groups and looking for disparate impact. Marcus's team uses the four-fifths rule from the EEOC's Uniform Guidelines: if the selection rate for any group is less than 80 percent of the rate for the highest-selected group, that is evidence of adverse impact and triggers investigation before deployment rather than after. Where a gap shows up, the tool does not go live until it has been investigated and remediated.

Baseline fairness metrics agreed in advance. Testing is only meaningful if you decided beforehand what result would concern you. Before deployment, establish what fairness looks like for this tool and what threshold triggers investigation. If a tool advances candidates at a 30 percent rate for men and a 25 percent rate for women, that is a five-point gap, and the question of whether it is acceptable has to be answered upfront rather than negotiated in the moment when someone has already committed to the tool. Baselines also give you a comparison point later, which is how you tell drift from noise.

Ongoing monitoring on a schedule. Deployment is not the end of the obligation. Fairness metrics get re-checked monthly or quarterly, because a tool that was balanced at launch can drift as the applicant pool changes or as new data feeds into it. Marcus requires a quarterly re-check and treats a change in the metrics as a trigger for investigation rather than something to note and move past.

A clear escalation process. If anyone suspects bias in a tool, there is a defined route for raising it. The policy names where a concern gets reported, who investigates, what the timeline is, and how the outcome is communicated back. Escalation paths that exist only in principle do not get used, because the person with the concern does not know where to send it and assumes someone else will.

Documentation of how each tool works and is used. For every AI tool, the team records what inputs it takes, what it produces, how the output is generated to the extent the vendor discloses it, and who reviews that output. This documentation is what an investigation depends on if a problem emerges, and it is the difference between answering an auditor with a document and answering with a guess.

Candidate transparency and explainability. Where a tool materially affects a decision, candidates are told in plain language, for example, "We use an automated tool to help screen resumes for required skills, and a recruiter reviews every result." This does not mean disclosing proprietary algorithms. It means a candidate can understand how AI affected the decision about them: which criteria the tool was assessing and what role a human played. This satisfies the spirit of Local Law 144's notice rule and the reasonable expectation that people know how they are being evaluated.

A reasonable-accommodation path. Under the ADA, a candidate must be able to request an alternative if an AI assessment disadvantages them because of a disability, for example, a video tool that misreads speech or facial movement. The policy names who handles those requests so they do not fall through the cracks.

Prohibited: The Bright Lines

Prohibited practices are the things that, done once, can create real harm or liability. Marcus kept this list short so it would be memorable, on the theory that a red line nobody can recite is not a red line.

No fully automated rejections. An AI tool may rank, flag, or summarize, but it may never send a rejection or advance a finalist without a human reviewing the decision. Automated rejection at scale is exactly the pattern that bias-audit laws and EEOC enforcement target. The same applies at the other end of the funnel: offers and terminations are high-stakes decisions that require human involvement, and no policy should leave room for a model to make them.

No inference of protected characteristics. The team does not use AI to guess age, race, gender, disability, religion, national origin, sexual orientation, gender identity, or pregnancy, and none of these belong in a model's decision logic. Using a protected characteristic this way is unlawful under EEO law in many jurisdictions, and there is no legitimate recruiting purpose that requires it.

No proxies that stand in for protected characteristics. This is the rule that actually bites, because proxy discrimination is rarely deliberate. If you cannot use age directly, you cannot use graduation year, which correlates with it. If you cannot use race directly, you cannot use a neighborhood or geographic feature that correlates with it. Biased systems usually get that way not because someone encoded a protected trait but because the correlated data was sitting in the training set. Consider a company that built an AI tool for technical assessment and deliberately excluded gender from the model, then trained it on the profiles of past high performers. Because the historical population of high performers in that organization was mostly men, the model learned male-associated patterns and reproduced gender bias without ever seeing a gender field. That is proxy discrimination, and the remedy was to retrain on balanced data or impose fairness constraints on the model. Using a proxy to do what the law forbids directly is still discrimination.

No deploying a screening tool without governance review. Every tool goes through the review process before it touches a candidate. Low-risk tools can be fast-tracked, and they should be, because a process with no fast lane is a process people will bypass. But skipping review entirely is prohibited, because that is precisely how tools slip into production without fairness testing.

No ignoring a raised fairness concern. If someone raises a fairness concern, it gets investigated. You are not obliged to change the tool because of every concern, but you are obliged to look. Declining to investigate is itself the violation, and it is the one that reads worst in hindsight, because it establishes that the organization knew and chose not to know more.

No candidate personal information in consumer AI tools. Recruiters may not paste names, contact details, salary history, or full resumes into a free, consumer-grade chatbot account, because in many consumer tiers the provider may retain inputs and use them to improve its models. Candidate data goes only into approved enterprise tools with a contract that forbids training on the team's inputs.

Encouraged: Where the Team Should Lean In

Encouraged practices are low-risk, high-value uses that Marcus actively wants his recruiters reaching for, because the productivity gain is real and the downside is small. These are the uses where AI helps a recruiter rather than replacing judgment, and naming them explicitly matters: a policy that only ever says no teaches the team that the policy is an obstacle.

Drafting, summarizing, and research. Outreach messages, job-description first drafts, and rejection notes are all fair game, as long as a human edits before sending and no candidate personal data goes into a consumer tool. Condensing structured interview notes, synthesizing feedback from multiple interviewers into a comparison against the same rubric, and tightening long intake notes save time and, done well, improve consistency across candidates. Market salary ranges, competitor talent landscapes, and skill-trend scans are open-ended tasks where a research assistant saves real time and carries little risk, because the output informs the recruiter rather than deciding anything about a person.

Pilot before you scale. Rather than deploying a new tool across the whole process at once, run it with one team or one role family first. Monitor fairness during the pilot, gather feedback from the people using it, and expand only when both look right. A longer pilot costs weeks; a bad tool deployed everywhere costs a quarter and a remediation project.

Involve the people who will use the tool. Recruiters, coordinators, and hiring managers who work the process every day notice things an evaluation committee does not. Bringing them into tool selection and testing surfaces practical problems early, and it has a second effect worth naming: people use a tool they helped choose, and resist one that arrived by decree.

Go beyond the minimum on candidate transparency. Required disclosure is a floor. Explaining why you use AI and what it helps you evaluate builds trust that bare compliance does not: "We use AI for initial screening because it helps us review the 500 or more applications we receive for each role fairly and quickly, and our recruiters review every application the tool flags."

Collect feedback and share results. Ask your team routinely whether a tool is helpful, whether it has bugs, and whether anyone has fairness concerns, and build that loop into the calendar rather than waiting for complaints. Then share what the monitoring finds: "this month's fairness metrics for the screening tool show a 35 percent advancement rate for men and 34 percent for women, within our threshold." Transparency internally does the same work as transparency externally. It builds trust, and it gets the whole team invested in the fairness numbers rather than treating them as compliance's problem.

Celebrate what works. When tools perform, when fairness holds, and when efficiency improves without a quality cost, say so out loud: "our new screening tool has reduced time-to-fill by 20 percent while holding our fairness metrics steady." Recognition reinforces the behavior you want, and it counteracts the impression that governance exists only to catch people doing things wrong.

Build communities of practice. Bring together people across teams who are using AI tools to share what they have learned and what has gone wrong. Institutional knowledge about which prompts work, which tool outputs need the most correction, and which failure modes recur is expensive to rediscover team by team.

Worked Example: Marcus's One-Page Policy Table

Marcus turned the three buckets into a single table his team pinned to the recruiting wiki. Each row is a concrete rule paired with the obligation or reason behind it, because a rule without a reason is the kind people quietly decide does not apply to their situation.

Bucket Rule Why it is there
Required A recruiter reviews and signs off on every AI screening result before a candidate is advanced or rejected. Local Law 144 and EEOC adverse-impact expectations; keeps accountability with a named person.
Required Any screening tool passes a four-fifths-rule check before launch and every quarter after. Catches adverse impact before candidates are affected, and catches drift after.
Required Every tool's data flow, output, and named reviewer is documented. An investigation or audit needs a document, not a recollection.
Required Candidates receive plain-language notice when a tool materially affects screening. Local Law 144 notice rule and basic candidate trust.
Required An ADA accommodation request reaches a named owner within two business days. Accommodation obligations fail through inattention, not refusal.
Prohibited No automated rejection or offer without human review. Fully automated adverse decisions draw the heaviest legal scrutiny.
Prohibited No candidate personal data in consumer chatbot accounts. Consumer tiers may retain inputs and train on them.
Prohibited No AI inference of protected traits or their proxies. Unlawful directly, and equally unlawful through a correlated feature.
Prohibited No screening tool live without a current bias audit and governance review. Review is how tools with untested fairness get caught before deployment.
Encouraged Draft outreach, job descriptions, and rejection notes with AI, then edit. Real time savings, and the human edit is the guardrail.
Encouraged Summarize interview notes and synthesize feedback against a shared rubric. Saves time and improves consistency across candidates.
Encouraged Run market, competitor, and skills research. Informs the recruiter without deciding anything about a person.

The four-fifths row is the one recruiters ask about most, so the policy shows the arithmetic. In the team's last check, the resume screener advanced 42 percent of one group and 31 percent of another. Dividing 31 by 42 gives 74 percent, which falls below the 80 percent threshold, so the tool was paused and retuned before relaunch. Writing the calculation into the policy means nobody has to be talked through it during an argument about whether a gap is big enough to act on.

The table is one page. A new recruiter reads it in five minutes and knows exactly where the lines are. That is the test of a good policy: not how thorough it looks, but how fast someone can act on it correctly. Each encouraged row carries the same pair of guardrails, a human reviews the output and no candidate personal data enters an unapproved tool, which keeps the permissive bucket from becoming an exception to the restrictive ones.

Enforcement: Making the Lines Real

Policies are only as strong as their enforcement, and the difference between a policy and a preference is what happens when someone ignores it. Marcus defined consequences at three levels, scaled to both the severity of the violation and the intent behind it.

Minor violations are things like late documentation or incomplete fairness testing that gets corrected quickly. The response is coaching: here is the policy, here is why it exists, let us make sure this does not recur. Treating a paperwork lapse as a disciplinary matter teaches people to hide lapses rather than fix them.

Moderate violations are things like deploying a tool without going through governance review, or receiving a fairness concern and not investigating it. The response is escalation and remediation: the decision gets undone, the tool comes out of production, the proper process runs, and relaunch waits until it is complete. This takes longer than the shortcut would have, and that is the point; the cost has to land somewhere visible or the shortcut remains rational.

Serious violations are deliberate: knowingly hiding a bias finding, building prohibited characteristics into tool logic, or retaliating against someone who raised a concern. These carry termination or legal consequences, and the policy says so plainly.

Intent is the axis that separates the tiers as much as impact does. Consider two teams that both deployed an assessment tool without committee approval. The first did not know the review was required; the response is a coaching conversation, an undeployment, a run through the proper process, and a clarified expectation going forward. The second knew the policy, deployed anyway to save time, and then concealed it. Same act, entirely different situation, and the policy should not pretend otherwise. What matters most across all three tiers is consistency. If one team violates policy without consequence, everyone learns what the policy is really worth, and no amount of restating it recovers the credibility.

Keeping the Policy Alive

Marcus scheduled a quarterly review of the page itself. Tools change, laws change, and teams accumulate practices faster than they write them down, so a policy that is never revisited rots into the same uselessness as having none. He dated the document, assigned himself as owner, and put the next review on the calendar. A named owner is what prevents the document from becoming everyone's responsibility and therefore no one's.

The review does two jobs. It reconciles the policy against what the team is actually doing, which is how you catch the fifth tool that arrived without anyone noticing. And it consolidates: policies accumulate over years and start to conflict, at which point people stop trying to work out which one governs and simply do what seems reasonable in the moment. One current, coherent page beats three overlapping documents of unknown vintage, and keeping it to one page is what makes the quarterly reconciliation cheap enough to actually happen.

Anti-Patterns

Policies that are too strict to follow. The organization requires that every AI decision go through full committee review with legal approval. It reads well and sounds responsible. In practice, approval takes three months, the team needs results sooner, and tools get deployed without approval. The recruiters are not trying to be non-compliant; they are trying to be responsive, and the policy has made those two things mutually exclusive. What goes wrong is that unrealistic policies generate workarounds, workarounds become underground practice, and governance becomes theater precisely because it was designed to be maximally rigorous. The fix is a decision-authority matrix: categorize decisions by risk, fast-track the low-risk ones, and reserve the full process for the high-risk ones where the delay is worth paying.

Policies with no stated reasoning. The document says all AI tools require fairness testing. A recruiter asks why a scheduling assistant needs a fairness test, gets no clear answer, concludes the policy is bureaucracy, and skips the test. Later the tool turns out to have a bug that affects candidates, one that testing would have caught. It happens because policies get written as rules rather than as arguments. What goes wrong is that people who do not understand a rule do not internalize it; they either ignore it or resent it, and both erode compliance elsewhere. The fix is to attach the reason to every rule, ideally a concrete one: we require fairness testing because a tool of this kind produced disparate impact somewhere and testing would have caught it, and we would rather find that in a test than in a complaint.

Conflicting policies nobody has reconciled. An older policy says hiring managers must make final hire decisions. A newer one says the AI tool's output is the final decision at the screening stage. Hiring managers cannot tell which governs, so they do whatever feels right that day, and the result is inconsistency that no one intended and no one owns. It happens because policies accumulate without anyone being responsible for the whole set. What goes wrong is that confusion erodes compliance more efficiently than disagreement does, since people who are guessing are often guessing wrong. The fix is periodic consolidation: review the full set, resolve conflicts explicitly, retire what is superseded, and maintain one authoritative version.

Practice Prompts

  • Run the required-practices checklist. For each required practice, fairness testing, baseline metrics, ongoing monitoring, escalation, tool documentation, candidate transparency, human review, and accommodation, assess whether your team does it completely, partially, or not at all. For each gap, write down what implementation would take and who would own it.
  • Audit against the prohibited list. Go tool by tool and verify four things: does it use protected characteristics, does it use proxy features, did it go through governance review, and are there fairness concerns that were raised and never investigated? Anything you find becomes a remediation plan with a date.
  • Write one policy properly. Pick a single area, fairness testing, escalation, or monitoring, and draft it fully: what must happen, by when, by whom, and what the consequence is if it does not. Specific and enforceable, not aspirational.
  • Build a decision-authority matrix. List the decisions your team makes, such as deploying a tool, changing evaluation criteria, or adding a sourcing channel. Categorize each as high, medium, or low risk and assign an approval path to each category, including a fast lane.
  • Draft the communication plan. Write three messages: introducing the policy, explaining why it matters, and reminding the team of the rules and consequences. If you cannot make the second one convincing, the policy probably needs better reasons rather than better wording.

Reflection

  • What is one AI recruiting practice on your team that genuinely concerns you, and what would a policy have to say to address it?
  • If you sat on the governance committee and someone proposed a new tool, what would you ask, and what evidence would you want to see before approving it?
  • Which required practice do you consider most important, and why that one over the others?
  • Have you seen a policy that was written and never enforced? What happened to the rest of the rules once people noticed?
  • If fairness testing revealed bias in a tool your team depends on, would you fix it, ban it, or use it with caveats, and what would justify your answer?

Glossary

  • Fairness testing. Analysis of an AI tool's output to detect disparate impact or bias before deployment, typically by comparing outcomes across demographic groups.
  • Proxy feature. A feature that correlates with a protected characteristic without being one, such as graduation year for age. Using a proxy to achieve what the law forbids directly is still discrimination.
  • Disparate impact. Outcomes that disproportionately disadvantage a protected group. This is what fairness testing looks for.
  • Four-fifths rule. The EEOC Uniform Guidelines benchmark under which a selection rate below 80 percent of the highest group's rate is evidence of adverse impact warranting investigation.
  • Governance review. The process by which a cross-functional group evaluates a tool before deployment, covering fairness, compliance, and operational impact.
  • Baseline metrics. Fairness measurements taken before deployment, used as the comparison point for detecting whether a tool has drifted.
  • Automated employment decision tool. A tool that substantially assists or replaces discretionary hiring or promotion decisions, and the category NYC Local Law 144 regulates.

A policy is one layer of governance. These lessons cover the structures it depends on and the practices it regulates.

Closing

Policies are how you operationalize values. They turn an abstract commitment to fairness and responsibility into concrete practices that guide day-to-day decisions and make clear what is acceptable and what is not. The best ones are clear, consistent, enforced, and understood; when people see that a policy makes sense and is applied evenly, they follow it, and when they see it ignored or enforced selectively, they stop. As you design yours, start from your organization's actual risks rather than from a template: where in your process do decisions carry real consequences for candidates, where do you need tight control, and where can you afford flexibility? Then write the rules that address those risks, attach the reasoning, set proportionate consequences, name an owner, and put the review on the calendar. That is how governance stops being a document and becomes something your team can be proud of.

Key Takeaways

  • Sort every AI practice into required, prohibited, or encouraged. The three buckets answer the only questions a recruiter has in the moment, and they keep the whole policy on a single page that people will actually read.
  • Make human review of AI screening decisions a hard requirement. No candidate is advanced or rejected on a tool's output alone, and no offer or termination rests on a model. This is the central control that keeps the team out of fully automated decisions.
  • Tie the required bucket to real obligations. NYC Local Law 144 demands an independent bias audit and candidate notice for automated screening tools; the EEOC's four-fifths rule flags adverse impact; the ADA requires an accommodation path. Cite the reason next to each rule so people internalize it.
  • Test before launch, set the threshold first, and re-check on a schedule. Decide in advance what gap would concern you, because a threshold negotiated after you have committed to a tool is not a threshold. A tool balanced at launch can drift, so monitor quarterly.
  • Keep the prohibited list short and absolute. No automated rejections, no candidate personal data in consumer AI tools, no inference of protected traits, no proxies for them, no tool live without governance review, and no ignoring a raised fairness concern.
  • Proxies are where good intentions fail. A model trained on historical high performers can reproduce bias without ever seeing a protected field, because the correlated signal is in the data. Excluding the characteristic is not the same as excluding its effect.
  • Encourage the low-risk wins explicitly. Drafting, summarizing, research, piloting before scaling, involving the people who will use a tool, sharing metrics, and building communities of practice. A policy that only says no teaches the team to route around it.
  • Give the policy teeth, an owner, and a review date. Scale consequences to severity and intent, apply them consistently, keep one authoritative version rather than several conflicting ones, and revisit it quarterly. Unenforced policy is worse than none, because everyone can see the gap.

Frequently Asked Questions

Does every AI tool need the full required treatment, even a scheduling assistant? Scale the process, not the principle. A decision-authority matrix lets you fast-track low-risk tools through a lighter review while reserving the full process for anything that touches a screening or selection decision. What you should not do is exempt a tool from review entirely on the grounds that it seems harmless, because that judgment is being made by the person who wants to deploy it. A short review that confirms a tool is low-risk is cheap; discovering after the fact that it was affecting candidates is not.

What if fairness testing shows a gap but we cannot tell whether the tool caused it? That is the normal case, and it is why the policy requires investigation rather than automatic removal. A four-fifths result below 80 percent is evidence of adverse impact, not proof of a defective tool; the applicant pools may differ, or the criteria may be genuinely job-related and unevenly distributed. What the policy forbids is proceeding without looking. Investigate, document what you find, and if the difference cannot be explained on job-related grounds, the tool does not launch until it has been remediated.

How do we enforce a policy without making the team defensive? Separate honest mistakes from deliberate ones, out loud and in writing. Most violations are people trying to move fast under pressure, and treating those as misconduct guarantees the next lapse gets hidden rather than reported. Reserve serious consequences for concealment, retaliation, and knowing violation, coach the rest, and be visibly consistent about both. Consistency is what makes the line credible; selective enforcement teaches the team that the policy applies to whoever is least able to argue.