←
AI for Recruiters
Proficient · M20 · lesson 20 of 32 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Hands-On Project: Develop Decision Rules for Your Recruiting Process

15 min

Marcus is a recruiting manager at a regional healthcare network with about 4,000 employees and a TA team of nine recruiters. His team fills roughly 600 requisitions a year, and during peak hiring his recruiters each carry 25 to 30 open roles at once. Under that load, two recruiters looking at the same candidate often reach different conclusions, and Marcus cannot explain, after the fact, why one borderline applicant advanced and an almost identical one did not. He has started using an AI resume-screening assistant to triage the highest-volume roles, which makes the consistency problem more urgent, not less, because now the inconsistency is partly automated. This project is what he built in response, and it is the artifact you are going to build for your own process.

What You Are Building, and Who You Are Building It For

This hands-on project produces one concrete artifact: a written decision-rules document that says, for each major stage of the process, what score or signal advances a candidate, what holds them for a second look, what rejects them, and exactly when a human must step in regardless of what the score says. By the end you will have a one-page rules table your team can apply on Monday morning, plus the supporting logic that makes it defensible if a candidate, a regulator, or your own legal team ever asks how the decision was made.

A decision rule is simply a documented, repeatable answer to a recurring question: given this input, what happens next, and who decides. The point is not to remove human judgment. It is to spend that judgment where it matters, on the genuinely borderline cases, instead of re-litigating every routine one from scratch. That reframing is worth holding onto, because the objection you will hear first is that rules mechanize a human craft. The load Marcus's recruiters carry is the answer. A recruiter juggling 25 to 30 open roles is not applying careful judgment to every candidate; they are applying whatever heuristic is closest to hand, differently on a Tuesday than on a Friday. Rules do not replace judgment there. They replace improvisation, and they free the judgment for the cases that actually need it.

The defensibility half matters just as much and is easier to underrate while nothing has gone wrong. Marcus's current problem is not that his team makes bad calls; it is that he cannot reconstruct any call after the fact. A documented rule turns "we thought she was stronger" into "she scored above the advance threshold on a rubric applied to every candidate for this role," and those two sentences have entirely different value to a candidate asking why, to a hiring manager challenging a shortlist, and to counsel responding to a charge.

Step One: Map Your Decision Points Before You Write Any Rules

Before Marcus can write rules, he has to name the moments where a decision actually gets made. Rules attached to vague stages drift; rules attached to specific decision points hold. For his network, the decision points are: resume screen (advance to phone screen or not), phone screen (advance to skills assessment or not), assessment plus interview (advance to offer, hold for committee, or reject), and reference and background stage (clear to offer or escalate). Each of those is a real fork with a real owner, which is what distinguishes a decision point from a stage name on a pipeline diagram.

For each point he writes down three things: the inputs available at that moment, who currently owns the call, and whether AI produces any score or recommendation that feeds in. This matters because a decision rule is only as good as the inputs it can actually see. There is no use writing a rule that depends on a structured interview rubric if the interview is unstructured and the rubric does not exist yet. Marcus discovers that his phone screens produce only free-text notes, so before he can set a phone-screen threshold he has to define three to five scored dimensions for the screen itself.

That is normal, and it is worth expecting rather than treating as a setback. Writing decision rules usually exposes the places where your process was never really standardized. The exposure is itself a finding: a stage that produces only free-text notes is a stage where two recruiters can reach opposite conclusions with nothing on record to explain the difference, which is exactly the problem Marcus set out to solve. Mapping first is what prevents you from writing a threshold against an input that does not exist, discovering it three weeks later, and concluding that rules do not work in your process.

Step Two: Turn the Criteria Into a Scoring Model

A threshold needs a number, and a number needs a scoring model. Marcus builds a simple one for his highest-volume role, a registered nurse position. He defines four weighted dimensions, each scored 1 to 5 by the assessor, with the AI screen contributing only to the first dimension and only as a recommendation a recruiter confirms. That last constraint is doing structural work: the tool's output enters the model at one defined point, bounded by one weight, and only after a human has confirmed it, which is what keeps the AI assistive rather than deciding.

  • Required credentials and licensure (weight 30 percent). Active license, required certifications, minimum clinical experience for the unit. This is close to pass or fail and is verified by a human before it ever scores high.
  • Relevant clinical experience (weight 30 percent). Depth and recency of experience in the specific specialty.
  • Structured interview performance (weight 25 percent). Scored against a fixed rubric with the same questions for every candidate for the role.
  • Role-specific competencies (weight 15 percent). Communication under pressure, teamwork, and other competencies tied to the unit, scored from defined behavioral anchors.

A candidate scoring 5, 4, 4, and 3 lands at (5 x 0.30) + (4 x 0.30) + (4 x 0.25) + (3 x 0.15) = 1.50 + 1.20 + 1.00 + 0.45 = 4.15 on a 5-point scale. Work through that arithmetic once by hand, because doing so shows you what the weights are actually buying. The candidate's weakest dimension is the one carrying the least weight, so a 3 there costs relatively little, while the same 3 on credentials would have pulled the composite down much harder. That is the intended behavior, and seeing it in a worked number is how you check that your weights encode the priorities you meant them to.

The weights are Marcus's illustrative starting point for his network, not a universal formula, and he expects to tune them as he watches how scores correlate with actual hiring outcomes over the next two quarters. Behavioral anchors deserve one note of their own. A dimension scored from defined anchors, where a 4 means a described behavior rather than a general impression, is what makes the same rubric produce comparable numbers across nine different recruiters. Without anchors, weighting is arithmetic performed on opinions, and the composite inherits every inconsistency the model was meant to remove.

Step Three: Set Advance, Hold, and Reject Thresholds

With a score, Marcus can write the rule table. The principle he follows: thresholds should be wide enough that clear cases resolve automatically, with a deliberate hold band in the middle where human review is mandatory. A process with no hold band forces every score onto one side of a hard line, which is exactly where small scoring errors turn into bad and hard-to-defend decisions. Consider what a single hard cut through the middle of the range would do to the two candidates sitting immediately either side of it, separated by one assessor's judgment on one anchored dimension and sent to opposite outcomes with no review in either direction.

Composite score (1 to 5)Default decisionOwnerHuman-review trigger
4.0 and aboveAdvanceRecruiterSpot-check 1 in 10 to confirm the rubric was applied consistently
3.0 to 3.9Hold for reviewRecruiter plus hiring managerMandatory second reviewer before any advance or reject
Below 3.0RejectRecruiterMandatory review if AI screen and human assessor disagree by 2 or more points on any single dimension

The hold band from 3.0 to 3.9 is where the team's judgment is spent on purpose. These are the candidates who are neither obvious yes nor obvious no, and they are also where bias most often hides, because ambiguity is where unexamined preferences fill the gap. Routing every borderline case to a second reviewer with a shared rubric is both a quality control and a fairness control, and it is the same intervention serving both ends rather than two competing demands on the same time.

Notice that every row carries a review provision, including the ones where the default decision is automatic. The 1-in-10 spot-check above 4.0 is not there because advancing a strong candidate is risky; it is there because that is the only way to detect a rubric being applied loosely, and a rubric that drifts at the top drifts everywhere. The disagreement check below 3.0 covers the outcome with the least natural scrutiny: nobody appeals an advance, so rejections are where an uncaught scoring error becomes permanent and invisible.

Step Four: Define the Override and Escalation Triggers

Thresholds handle the routine. Triggers handle the exceptions, and they are the most important part of the document because they are what keeps the AI in an assistive role rather than a deciding one. Marcus writes explicit triggers that force a human decision regardless of the score:

  • Model disagreement trigger. When the AI screen's recommendation and the recruiter's confirmed score diverge by 2 or more points on any dimension, the case stops and a second human adjudicates. A systematic pattern of disagreement is a signal the tool is miscalibrated for the role.
  • Protected-attribute proximity trigger. Any time a rejection reason touches something that could correlate with a protected class, for example an employment gap, the rejection requires manager sign-off and a documented job-related justification.
  • Accommodation trigger. If a candidate has requested an accommodation under the ADA, their assessment is reviewed by a human who confirms the scoring did not penalize a disability-related factor, and that an alternative assessment path was offered where appropriate.
  • Adverse-impact trigger. When the monthly fairness check shows the selection rate for any group falling below four-fifths of the highest group's rate, the rules themselves are escalated for review, not just individual candidates.

The first trigger is worth reading twice, because it does two jobs on two timescales. In the individual case it stops a decision that two assessors see differently. In aggregate it is a calibration monitor: a tool that keeps disagreeing with confirmed human scores on the same dimension is telling you something about the tool, and that signal only exists because the disagreements are being recorded rather than resolved silently in whichever direction is faster.

The last trigger connects the individual rules to the EEOC's adverse-impact analysis. Under the four-fifths (80 percent) rule of thumb, if the selection rate for one group is less than 80 percent of the rate for the group with the highest selection rate, that disparity is a flag for further investigation. The four-fifths rule is a screening heuristic, not proof of discrimination, but it is the right tripwire to build into a decision-rules system so that a quietly accumulating disparity gets caught at the process level. Note what it escalates: the rules, not a candidate. A pattern across a population is not a defect in any single decision, and looking for the responsible individual case is how a process-level problem stays unfixed.

Step Five: Pressure-Test the Rules Against the Law

Because Marcus's screen uses an AI tool to score candidates, his decision rules sit inside a growing legal frame, and he writes a short compliance note alongside the table so the constraints are visible to whoever applies it. That placement is deliberate. A compliance memo filed separately is read once; a note attached to the rules table is read by everyone who uses the rules, which is where the obligations actually have to bite.

If the network hires in New York City, the AI screen is an automated employment decision tool under NYC Local Law 144. That law requires an independent bias audit of the tool within the prior year before it is used, publication of a summary of the most recent audit results, and notice to candidates at least 10 business days before the tool is used on them, including the job qualifications and characteristics the tool assesses. Marcus's rules document references the audit date and the candidate-notice step so the obligation is not just policy but an enforced part of the workflow. The 10-business-day notice in particular is the kind of requirement discovered to have been missed only after candidates have already been screened, which is why it belongs in the workflow rather than in a policy binder.

The EEOC's adverse-impact framework and the four-fifths rule govern whether the outcomes of his rules disproportionately exclude a protected group, which is why the adverse-impact trigger lives in the table. The ADA requires that an assessment not screen out qualified candidates with disabilities and that reasonable accommodations be available, which is the accommodation trigger. If the network recruits candidates in the EU, the GDPR gives candidates rights around solely automated decisions and the EU AI Act classifies recruitment and selection systems as high-risk, carrying obligations for human oversight, transparency, and record-keeping.

Read that list as a design constraint rather than a reading assignment, because each item has already been answered by a specific feature of the rules table. The hold band and the triggers are the human oversight the AI Act names. The mandatory second reviewer is what keeps a decision from being solely automated. The rubric and the logged scores are the record-keeping. Marcus does not need to be a lawyer to write the rules, but he does need the rules to leave room for these obligations, and he flags each one for review with counsel rather than guessing at the details.

Step Six: Validate the Rules Before You Trust Them

A decision-rules document is a hypothesis until it is tested. Marcus validates his by running it retroactively against the last 50 closed requisitions for the nurse role: he scores those candidates under the new rules and compares the rules' decision to what actually happened. Where the rules would have rejected someone who turned out to be a strong hire, or advanced someone who washed out, he looks at why. Usually the answer is a weight that is off or a rubric anchor that is ambiguous, not that rules are a bad idea.

Retroactive testing is the right first move because it is free and immediate. You already have the outcomes, so the rules can be wrong on paper before they are wrong on a real candidate, and every disagreement between rule and history is a question worth asking rather than a verdict. Do not treat every mismatch as a rule failure, either. Some will be cases where the historical decision was the inconsistent one, which is the problem you set out to fix, and telling those apart is where the exercise earns its value.

He also sets a review cadence: the thresholds and weights are revisited quarterly, the fairness check runs monthly, and any single trigger that fires more than a handful of times in a month prompts a look at whether the rule or the tool needs adjustment. The different frequencies are not arbitrary. Fairness is checked monthly because a disparity compounds with every requisition it touches, while weights are revisited quarterly because you need enough closed outcomes to tell signal from noise. A trigger firing constantly is its own diagnostic: it usually means a threshold is drawn in the wrong place or the tool is miscalibrated for that role, not that the exception is unusually common. Decision rules are not set once. They are a living instrument that gets more accurate as the team feeds real outcomes back into it.

Anti-Patterns

Writing thresholds against inputs that do not exist. This is setting a phone-screen cutoff when the phone screen produces only free-text notes. It happens because the stage exists on the pipeline diagram and looks ready to have a number attached, and because mapping inputs feels like preamble to the real work. What goes wrong is that the rule cannot be applied as written, so recruiters improvise a translation from notes to score, and the document creates the appearance of consistency over exactly the same inconsistency it was meant to remove. The counter is Step One in full: for every decision point, write the inputs available, the owner, and whether AI feeds in, and build the missing scored dimensions before writing any threshold against them.

Drawing a single hard line instead of a hold band. This is one cutoff that sends everything above it forward and everything below it out. It happens because a single number is simpler to communicate and because a hold band looks like extra work with no obvious owner. What goes wrong is that the decisions nearest the line are the ones a scoring model is least reliable about, so two candidates separated by one assessor's judgment on one dimension get opposite outcomes with no review, and the ambiguous middle is precisely where unexamined preferences fill the gap. The counter is a deliberate band, in Marcus's case 3.0 to 3.9, with a mandatory second reviewer working from the same rubric before any advance or reject.

Treating the AI score as the score. This is letting the tool's output flow into the composite unconfirmed, or reading its recommendation as the answer the human is checking rather than an input the human owns. It happens because the number arrives first, looks precise, and saves the most time on the highest-volume roles where the pressure is greatest. What goes wrong is that the tool moves from assisting a decision to making one, which is the exact status that triggers obligations around solely automated decisions and human oversight, and a systematically miscalibrated tool then propagates silently because nothing records where humans would have disagreed. The counter is the structural bound Marcus uses: the AI contributes to one dimension, only as a recommendation a recruiter confirms, with divergence of 2 or more points stopping the case for a second human.

Escalating the candidate when the rules are the problem. This is responding to a monthly fairness check that shows a group's selection rate below four-fifths of the highest group's rate by reviewing individual files. It happens because individual review is the familiar remedy and because a population-level pattern has no single decision to point at. What goes wrong is that no individual case will explain the disparity, the review concludes that each decision looks defensible, and the pattern continues accumulating with the additional problem that it has now been looked at and cleared. The counter is to escalate the rules themselves, treating the four-fifths signal as a flag for investigating thresholds, weights, and tool calibration at the process level.

Publishing the rules and never validating them. This is finishing the document, rolling it out, and treating the weights as settled. It happens because the document is the deliverable and because writing it involved real thought that feels like it should be conclusive. What goes wrong is that weights are a hypothesis about what predicts a good hire, and an untested hypothesis applied at scale converts one person's guess into every decision the team makes, with no mechanism to notice it was wrong. The counter is retroactive validation against closed requisitions before you trust the rules, then a standing cadence, quarterly for thresholds and weights and monthly for the fairness check.

Build Checklist

Work these in order; each one produces a piece of the finished document.

  • Draw your decision points and their inputs. List every moment a real fork occurs in your process. For each, write the inputs available at that moment, who owns the call, and whether AI produces a score or recommendation that feeds in. Circle any decision point whose only input is free text; that is your first build task, not a rule you can write yet.
  • Build one scoring model for your highest-volume role. Define four or five weighted dimensions with behavioral anchors, so a 4 describes an observable behavior rather than an impression. Then score two real past candidates by hand and check that the composite ranks them the way you would have.
  • Set your bands and name an owner for each. Write the advance, hold, and reject thresholds, and for every band state who owns the decision and what review provision applies, including the automatic ones. A band with no review provision is a band where a scoring error becomes permanent.
  • Write your triggers as sentences that stop a decision. Draft the model-disagreement, protected-attribute proximity, accommodation, and adverse-impact triggers in your own context, each specifying what fires it, what stops, who adjudicates, and what gets documented.
  • Attach the compliance note to the table, not to a separate file. For each obligation that applies to you, write one line naming it and one line naming the feature of your rules that answers it. Flag the specifics for review with counsel rather than guessing.
  • Validate retroactively before you go live. Score your last batch of closed requisitions for one role under the new rules and compare each rule decision to what actually happened. For every mismatch, decide whether the rule was wrong or the historical decision was, and record which.

Reflection

  • Where in your process do two recruiters most often reach different conclusions about the same candidate, and what input is missing at that point?
  • If a rejected candidate asked why, at which stage would you be least able to answer?
  • What is your current hold band, and if you do not have one, where does the hard line fall and who lands just below it?
  • How would you find out if your AI screen were systematically disagreeing with your recruiters on one dimension?
  • Which of your rules exists because someone tested it, and which exists because someone assumed it?
  • When your fairness check last flagged something, did you review candidates or review the rules?

Glossary

  • Decision rule. A documented, repeatable answer to a recurring question: given this input, what happens next, and who decides.
  • Decision point. A specific moment where a real fork occurs, with defined inputs and a named owner, as opposed to a stage name on a pipeline diagram.
  • Scoring model. A set of weighted dimensions, each scored on a fixed scale by an assessor, that produces the composite a threshold can be set against.
  • Behavioral anchor. A described, observable behavior that defines what a given score on a dimension means, which is what makes the same rubric produce comparable numbers across different assessors.
  • Composite score. The weighted sum of the dimension scores, such as the 4.15 produced by scores of 5, 4, 4, and 3 against weights of 30, 30, 25, and 15 percent.
  • Hold band. The deliberate middle range, in Marcus's case 3.0 to 3.9, where no automatic decision is taken and a second reviewer is mandatory before any advance or reject.
  • Human-review trigger. A condition that forces human involvement regardless of what the score says, which is what keeps an AI tool assistive rather than deciding.
  • Model disagreement. Divergence between the AI screen's recommendation and the recruiter's confirmed score, which stops the case at 2 or more points on any dimension and, in aggregate, indicates miscalibration for the role.
  • Four-fifths rule. The EEOC rule of thumb that a selection rate below 80 percent of the highest-selecting group's rate is a flag for further investigation. A screening heuristic, not proof of discrimination.
  • Automated employment decision tool. The category an AI screen falls into under NYC Local Law 144, carrying obligations for an independent bias audit within the prior year, publication of a summary of results, and candidate notice at least 10 business days before use, including the job qualifications and characteristics the tool assesses.
  • Retroactive validation. Scoring already-closed requisitions under new rules and comparing the rules' decision to what actually happened, in order to test weights and anchors before applying them to live candidates.

Closing

Marcus started with a problem he could feel but not describe: two recruiters, one candidate, two conclusions, and no way to explain either after the fact. The AI screen did not create that problem, but it changed its character, because an inconsistency that is partly automated scales at machine speed and arrives with a number attached that looks more objective than the judgment it replaced. Nothing about that is fixed by using the tool less. It is fixed by deciding, in writing and in advance, what the tool's output is allowed to do.

The document you finish this project with is short. Decision points with their real inputs, a weighted rubric with anchors, three bands with an owner and a review provision each, four triggers that stop a decision cold, a compliance note attached to the table, and a validation pass against requisitions you have already closed. What it buys is not speed; the routine cases were always fast. It is that the borderline cases now get the judgment they need, the rejections can be explained, the tool stays assistive, and a disparity accumulating across a population gets caught at the level where it can actually be fixed. Then you tune it, because a rule nobody revisits is just last quarter's guess applied to this quarter's candidates.

Key Takeaways

  • Map decision points before writing rules. Name the specific moments a call gets made, the inputs available there, and who owns it. Writing rules usually reveals the stages that were never actually standardized, and those gaps have to be closed first.
  • A threshold needs a scoring model behind it. Define weighted, anchored dimensions scored consistently for every candidate in a role. A number without a transparent model is just a gut feeling wearing a disguise.
  • Build a deliberate hold band. Advance, hold, and reject thresholds with a middle band route borderline cases to mandatory second review. That hold band is where both quality and fairness are protected, because ambiguity is where bias hides.
  • Triggers keep AI assistive, not deciding. Model-disagreement, protected-attribute, accommodation, and adverse-impact triggers force a human decision regardless of the score, which is what keeps the system defensible.
  • Build the law into the workflow. NYC Local Law 144 requires an annual independent bias audit, published results, and 10-business-day candidate notice for automated employment decision tools. The EEOC four-fifths rule, the ADA, GDPR, and the EU AI Act each constrain how automated scoring may be used. The rules document should reference these obligations, not assume them.
  • Escalate the rules, not the candidate. A four-fifths signal is a population-level finding, so investigating individual files will clear each one and leave the pattern intact. Send thresholds, weights, and tool calibration for review instead.
  • Validate against real outcomes and iterate. Test the rules retroactively against closed requisitions, then review weights quarterly and fairness monthly. Decision rules are a living instrument, not a one-time policy.

Frequently Asked Questions

Does writing decision rules just replace recruiter judgment with a formula? It relocates the judgment rather than removing it. A recruiter carrying 25 to 30 open roles at peak is not deliberating carefully over every candidate; they are applying whatever heuristic is nearest to hand, which is why two of them reach different conclusions about the same person. Rules resolve the routine cases the same way every time and route the genuinely ambiguous ones, the hold band, to a second reviewer working from a shared rubric. The judgment is still there, spent where it changes an outcome instead of on cases that were never close.

Where do the weights and thresholds come from if we have never scored anything? They start as an explicit, documented guess, which is already an improvement on an undocumented one. Marcus's four dimensions and their weights are his illustrative starting point for his network rather than a universal formula, and he expects to tune them as he watches how scores correlate with actual hiring outcomes over the next two quarters. The important discipline is not getting the first numbers right; it is validating retroactively against closed requisitions before you trust them, and holding a quarterly review so the guess becomes evidence over time.

What if the AI screen and our recruiters disagree constantly? Then the trigger is doing exactly what it was built for, and the finding is about the tool rather than about any individual candidate. A single divergence of 2 or more points on a dimension stops that case for a second human to adjudicate. A systematic pattern of disagreement is a signal the tool is miscalibrated for the role, which is a different problem requiring a look at the tool, the role's dimensions, or both. The reason the pattern is visible at all is that disagreements are recorded rather than silently resolved in whichever direction is quicker.

Our fairness check flagged a group below the four-fifths threshold. What do we actually do? Escalate the rules, not the candidates. The four-fifths rule is a screening heuristic and a flag for further investigation rather than proof of discrimination, and the thing under investigation is the process: the thresholds, the weights, the rubric anchors, and whether the tool is calibrated for the role. Reviewing individual files is the tempting response and the least useful one, because each decision will look defensible on its own while the population-level pattern that produced the flag remains untouched, now with a review on record that appears to have cleared it.