Decision Rules: Explicit Criteria for Escalation
Dana runs recruiting operations at a 1,800-person logistics company, and last quarter her team deployed an AI screening model across 5,000 applications for warehouse-lead and dispatch roles. The model was good. It surfaced strong candidates fast and saved her recruiters an estimated 200 hours of resume review. It was also, in one specific way, dangerous: it had no idea when to stop and ask a human. Left alone, it would auto-route candidates to a reject queue at a fit score of 38, including people who clearly met every posted minimum qualification, and it would quietly infer a veteran-status flag from a resume line about military logistics experience. None of that was malicious. It was the predictable result of a system running without explicit escalation rules. Dana's job was not to make the model smarter. It was to design, in advance, the exact conditions under which a human being must look at a decision before it becomes final.
Why Escalation Rules Are a Design Decision, Not a Reflex
The instinct most teams bring to AI screening is to handle escalation by feel: a recruiter notices something odd, pulls the file, takes a second look. That works at ten applications a week. It collapses at 5,000 a quarter. At Dana's volume, "notice something odd" is not a control, it is a hope. The candidates who most need a human review are precisely the ones an overloaded recruiter is least likely to catch, because the model has already filed them into a queue that nobody reopens.
The same problem exists on the human side of the process, and it predates AI entirely. Most organizations never explicitly define who decides what, which produces ambiguity, slows hiring, and introduces inconsistency that only shows up when someone compares two similar candidates who got different treatment. Without written rules you get one of two failures and usually both: over-escalation, where everything goes to a senior person because nobody is sure they are allowed to decide, and under-escalation, where people make calls they should not be making because no one told them the call was not theirs. Decision rules are silent until you need them. You are evaluating a candidate, something does not fit, and the question is whether to escalate or move forward. Whether that moment goes well was determined weeks earlier.
An escalation rule turns hope into infrastructure. Instead of trusting that someone will catch the borderline case, you define the borderline case in advance and route it to a human automatically. The design question is not "should a human be involved" but "for which specific, named situations is human review mandatory, and what happens to everything else." When you can answer that with a table rather than a shrug, you have a system. When you cannot, you have an unmonitored decision engine making final calls about people's careers. This matters legally as much as operationally. Under the EEOC's guidance on algorithmic decision tools and Title VII, an employer remains responsible for discriminatory outcomes produced by a vendor's software. "The algorithm did it" is not a defense. Designing the human-in-the-loop triggers is how you keep accountable judgment in the loop where the law and basic fairness require it.
What a Decision Rule Actually Contains
A decision rule is not a sentence of guidance; it is a small structured record, and the structure is what makes it enforceable. For every decision type in your process, five fields have to be filled in: who decides, what information they need in order to decide, what triggers escalation to a higher authority, what timeline applies, and what documentation is required. A rule missing any one of those fields fails in a predictable way. Missing the information field produces decisions made on whatever happened to be in the file. Missing the timeline field produces decisions that are technically owned and practically abandoned. Missing the documentation field produces a process you cannot audit, which means you cannot tell whether it is being applied consistently.
Here is the shape of Dana's rule set for the human decisions that sit around the model, before any score is involved.
| Situation | Rule |
|---|---|
| Standard rejection | A recruiter can reject a candidate who does not meet explicit posted role requirements, such as a required technical skill or experience level. No escalation needed. |
| Borderline case | A candidate is missing a key skill but shows strong potential to learn it. Escalate to the hiring manager for a judgment call. |
| Offer decision | Made by the hiring manager within the approved salary band. Offers at the band ceiling require director approval; offers above the band require VP approval. |
| Fairness or legal concern | If a candidate discloses information that raises a fairness or legal question, escalate to HR before proceeding. No decision is finalized first. |
| Rejection on behavioral grounds | If the rejection rests on a behavioral assessment or reference concern, require a two-person review before anything is communicated to the candidate. |
Notice what each row does. It names the decider, it names the condition, and where it escalates it says to whom. The last two rows are the ones that matter most legally, and they share a feature worth copying: the escalation happens before the decision is communicated, not after. A rule that says "escalate concerns" without specifying that the candidate is not told anything in the meantime produces the situation where a rejection goes out on Tuesday and HR reviews it on Thursday, which is a review of something that has already happened.
The Score-Band Rule Table: Three Buckets and One Hard Override
Dana's core rule set for the model sorts every scored application into one of three fit-score bands, with a single override that outranks the bands entirely. The score is the model's 0 to 100 estimate of fit against the posted requirements.
Band A, scores 0 to 39 (low fit). The default action is to route to the reject queue. The mandatory override: if the candidate meets every posted minimum qualification, the application does not auto-reject. It escalates to human review regardless of score. A low model score on a candidate who is objectively qualified is exactly the signal that the model is reading something it should not, and a final automated rejection of a qualified applicant is the outcome you most want to avoid.
Band B, scores 40 to 69 (borderline). Every application in this band goes to human review. No auto-advance, no auto-reject. This is the genuine judgment zone, where the model is least certain and a recruiter's context adds the most value. The rule is deliberately blunt: borderline means a person decides.
Band C, scores 70 to 100 (high fit). These advance to the next stage automatically, with a 10 percent spot-check sample pulled for human review. The spot-check is not about catching bad advances; it is about catching a model that has started advancing the wrong people. If the sampled files stop looking like strong candidates, the sample is your early warning that the model has drifted.
Running the Numbers on 5,000 Applications
Rules are abstract until you put volume through them. Take Dana's 5,000 applications and assume an illustrative distribution: 55 percent land in Band A, 25 percent in Band B, and 20 percent in Band C. Watch what the rule table actually generates in human-review workload.
Band A holds 2,750 applications. Most will auto-reject, but suppose 8 percent of them meet every posted minimum qualification despite the low score. That is 220 applications pulled back from the reject queue into mandatory human review. Band B holds 1,250 applications, and every one of them goes to a human, so that is 1,250 reviews. Band C holds 1,000 applications; the 10 percent spot-check sends 100 of them to a reviewer. Add the categorical escalations described in the next section, and the headline number is roughly 1,570 human reviews out of 5,000, or about 31 percent.
That figure is the whole point of the exercise. Before the rules existed, the honest answer to "how many of these get human eyes" was "whatever the recruiters happen to open," which in practice trended toward zero for rejected candidates. After the rules, it is a known, staffable number that concentrates human attention on the cases where it changes outcomes. Dana can now plan headcount against 1,570 reviews instead of discovering a backlog when a complaint arrives.
Categories That Always Escalate, No Matter the Score
Some situations must reach a human regardless of where the fit score lands, because the score is the wrong instrument for the decision entirely. These are categorical escalations, and they sit above the band table.
Accommodation requests. If an applicant indicates a disability or requests an accommodation in the application or assessment process, the matter escalates to a human immediately and the automated screen does not act on it. Under the Americans with Disabilities Act, an accommodation request triggers an individualized, interactive process between employer and applicant. That process is legally a human responsibility. An AI model cannot conduct the interactive process, and an automated rejection that fails to account for a requested accommodation is a serious ADA exposure.
Protected-status and veteran flags must never be inferred. The rule here is two-sided. The model must not infer protected characteristics, including veteran status, disability, age, or national origin, from resume content, and any place where the system appears to be doing so is a defect to fix, not a signal to use. Where protected-status information is legitimately collected, for example voluntary self-identification for affirmative-action reporting, it is firewalled from the screening decision and any decision touching it escalates to a human.
Ambiguous or conflicting data. When the source material contradicts itself, for instance a resume claiming a certification that the structured application denies, the model should flag and escalate rather than guess. Guessing on conflicting data is how a system manufactures false rejections.
Any adverse action, and any anomaly. A final rejection is an adverse action, and the system is configured so that no candidate who met minimum qualifications receives an automated final rejection. Anomalies, such as a sudden spike in rejections from one applicant source or a score distribution that shifts overnight, escalate to whoever owns the model, because an anomaly is often the first visible symptom of drift or a data problem.
Specific Beats Vague: Writing Criteria That Hold
The single most common defect in escalation criteria is that they sound prudent and mean nothing. Vague criteria look like this: "if you're not sure, ask." "If it's a senior role, escalate." "If there's any concern, escalate." Each reads as responsible, and each is unusable, because none can be applied the same way twice by two different people. Their practical effect is over-escalation, since everything can be described as uncertain if you look at it long enough, and once everything escalates the senior people who were supposed to be a safeguard become a queue.
Specific criteria are testable by a person who was not in the room when they were written. "If the candidate does not meet two or more explicit role requirements, escalate." "If the role is director-level or above, the hiring VP has final authority." "If there is a fairness or legal concern, escalate to HR before any action." The test for any criterion you draft is whether two reasonable people, applying it to the same candidate, reach the same answer. "Missing two required skills" passes that test. "If something feels off" does not, and the fact that it feels safer to write is exactly why it survives in so many process documents.
This is also where the fairness argument and the efficiency argument point the same direction, which is rare enough to be worth naming. A vague criterion is applied differently to different candidates by definition, because the only thing determining the outcome is who happened to read the file. That is an inconsistency problem before it is a bias problem, and it becomes a bias problem the moment the variation correlates with anything about the candidates. Writing the criterion precisely is how you make the same situation produce the same handling.
Speed Is a Design Output, Not a Trade-Off
More escalation means slower hiring, and teams often accept that as the price of control. It is not, or at least not at the rate most teams pay it. The slowness comes from undefined authority far more than from genuine oversight. If every offer requires VP approval, you are slow and the VP is reviewing decisions they have no particular insight into. If offers within the approved band are the hiring manager's call, offers at the band ceiling need director approval, and offers above the band need VP approval, you are faster and the control is actually tighter, because the escalations that do happen are the ones worth someone senior's attention.
The design goal is therefore rules that maintain quality while enabling speed, which in practice means pushing authority down to the lowest level that can hold it and reserving escalation for the cases that genuinely exceed that level. Well-drawn rules produce a specific set of effects: decisions get made faster because nobody is waiting to find out whether they are allowed to decide; they get made more consistently because the criteria are the same across candidates; accountability is clearer because a named person owns each decision type; fairness issues surface earlier because they are a defined escalation trigger rather than something someone eventually raises; and second-guessing drops, because a decision made inside documented authority does not need to be defended.
Mapping Your Own Decision Types
Building your own rule set starts with observation rather than design. Take your recent hires and rejections and ask, for each one, who actually made the decision, on what basis, and whether that was the right decision point. The gap between who decided and who should have decided is your entire agenda, and it is usually not where people expect: the surprises are typically decisions being made two levels too high rather than too low.
Then enumerate the decision types in your process, which for most recruiting functions means the initial screen pass or fail, the first interview pass or fail, the technical screen pass or fail, the debrief decision to hire, reject, or hold, the offer negotiation, and the separate category of decisions involving fairness, legal, or behavioral concerns. For each, write the five fields. Worked out for the debrief decision, Dana's version reads like this. Authority sits with the hiring manager, with recruiting team input. The information required is structured interview feedback from all panelists, assessment results, the reference check summary, and any background check findings. The escalation triggers are disagreement among panelists, meaning an all-in voice against hesitant ones, any behavioral or legal concern, an offer that would exceed budget parameters, and the case where the candidate is an external referral from an executive. The timeline is a debrief within two business days of the final interview, and if escalation is needed, a decision within five business days. The documentation required is debrief notes including the rationale, plus escalation notes if it escalated.
That last trigger, the executive referral, is the one people leave out and the one most worth keeping. It is not there because executive referrals are bad candidates, but because it is the decision most likely to be made on a basis nobody wants to write down, and naming it as a trigger converts an awkward social pressure into a routine process step.
Monitoring With the Four-Fifths Rule
Escalation rules govern individual cases. Monitoring governs the pattern across cases, and the standard tool is the four-fifths rule. As a practical screening heuristic drawn from the federal Uniform Guidelines on Employee Selection Procedures, if the selection rate for any protected group falls below four-fifths (80 percent) of the rate for the highest-selected group, that is a flag for potential adverse impact that warrants investigation.
Dana runs this comparison on the screen's pass-through rates by group every cycle. Concretely, if the highest-passing group advances at 60 percent, any group advancing below 48 percent, which is four-fifths of 60, triggers a review of whether the model is producing disparate impact. The four-fifths check does not prove discrimination, and it is not a safe harbor on its own, but as an ongoing monitor it catches a drifting rule before it has rejected thousands of people. Pairing per-case escalation with this aggregate monitor is what turns a screening tool into a governed system: one watches the individual decision, the other watches the rule.
Documenting the Rules and Honoring Local Law
A rule that lives in someone's head is not a control. Dana's table, override conditions, categorical escalations, and monitoring thresholds are written down, versioned, owned by a named person, and shared with the team rather than filed somewhere. That documentation does double duty: it makes the system auditable internally, and it is the artifact you produce when a regulator or candidate asks how a decision was made. The internal audit is not optional either. Once a quarter, check whether similar decisions actually got made at similar authority levels.
Jurisdiction shapes the specifics. New York City's Local Law 144 governs automated employment decision tools used on candidates for NYC jobs, requiring a bias audit and candidate notice, and it presumes meaningful human involvement rather than fully automated final decisions. Illinois and Maryland regulate AI in specific hiring contexts, and the EU AI Act treats recruitment AI as high-risk. The portable principle underneath all of them is the one Dana built her table around: a human must remain genuinely in the loop for consequential decisions, and "genuinely" means defined triggers, not a rubber stamp after the model has effectively decided.
Anti-Patterns
Default escalation. This is the pattern where every decision goes to a senior person, stripping authority from the people who should hold it. It happens for an understandable reason: senior approval feels safer than personal judgment, especially in a risk-averse culture or after something has gone wrong. What goes wrong is that hiring slows to the pace of the busiest calendar, capable recruiters are underused and stop developing judgment they are never asked to exercise, and a VP spends their week on decisions they have no special insight into, which means the genuinely hard case arrives in a queue behind twenty routine ones. The counter is to define authority explicitly at the lowest level that can hold it, and to escalate only when a written criterion is met.
Vague criteria. This is escalation guidance like "if you're not sure" or "if it seems risky." It survives because specific criteria feel limiting while vague ones feel prudent, and because nobody has ever been criticized for writing a cautious sentence. What goes wrong is that everything qualifies, so everything escalates and nothing is decided at normal levels, which reproduces default escalation by another route. It also destroys consistency, because the criterion is effectively "whatever the reader thinks today." The counter is to write criteria a stranger could apply: "if the candidate is missing two required skills" is testable, "if something feels off" is not.
Inconsistent authority. This is the pattern where a hiring manager sometimes decides independently and sometimes escalates, with no discernible rule. It usually reflects rules that exist but were never written down, so people apply their memory of them. What goes wrong is that similar decisions produce different outcomes, which is the operational definition of unfairness even when every individual decision was made in good faith, and it is the pattern that looks worst in an audit precisely because no explanation for the variation exists. The counter is documented rules, consistent application, and a quarterly audit that compares like decisions and asks why any diverged.
Practice
- Audit five recent decisions. Take five recent hiring decisions and write, for each, who decided, on what basis, and whether that was the right person to decide. Look for decisions made two levels higher than necessary as well as ones made too low.
- Map your decision types. List every decision type in your recruiting process, from the initial screen through the offer, and for each write who should decide, what information they need, and what escalates it. The list is usually longer than people expect, and the unlisted decisions are where inconsistency lives.
- Rewrite your vaguest criterion. Find the escalation criterion in your process that reads most like "use your judgment" and rewrite it so that two people applying it to the same candidate would reach the same answer. Then test it on a real case from last quarter.
- Draw the score-band table for one live screen. If you use an AI screen, define the auto-reject band, the mandatory-review band, the auto-advance band, the spot-check rate, and the qualification override. Then run your actual volume through it and calculate how many human reviews it generates, so you know whether it is staffable before you turn it on.
- Run a consistency audit. Compare recent decisions of the same type and check whether they were made at the same authority level. Where they were not, find out why, and treat the answer as a defect in the rule rather than in the person.
Reflection
- Which decision type creates the most uncertainty in your process right now, and is the uncertainty about the candidate or about who is allowed to decide?
- Which decisions are currently escalated that could safely be made at a lower level, and what would it cost you to push that authority down?
- Do similar hiring decisions in your team consistently get made at similar authority levels? How would you find out, rather than assume?
- What fairness concerns trigger escalation in your process today, and are they written down anywhere a new recruiter would find them?
- If your AI screen made a final rejection tomorrow on a candidate who met every posted minimum qualification, what in your rule set would have stopped it?
Glossary
- Decision rule. An explicit, documented statement of who decides a given decision type, what information they need, what triggers escalation, what timeline applies, and what documentation is required.
- Escalation criterion. The specific, testable condition that moves a decision to a higher authority level. Testable means two people applying it to the same case reach the same answer.
- Authority matrix. The documented map of who holds decision authority for each decision type.
- Standard decision. A decision that can be made independently within documented authority, with no escalation required.
- Escalation event. A case that meets a defined criterion and therefore moves to a higher authority level.
- Categorical escalation. A situation that must reach a human regardless of any score, because the score is the wrong instrument for the decision: accommodation requests, protected-status questions, conflicting data, anomalies, and adverse action against a qualified candidate.
- Fit score band. A range of model scores with a defined default action attached, such as auto-reject, mandatory human review, or auto-advance with spot-checking.
- Qualification override. The rule that a candidate meeting every posted minimum qualification cannot be auto-rejected regardless of model score.
- Four-fifths rule. The screen from the Uniform Guidelines on Employee Selection Procedures: a group's selection rate below 80 percent of the highest group's rate flags potential adverse impact warranting investigation. Not proof of discrimination, and not a safe harbor.
- Consistency. Applying the same decision rules across candidates and across time. It is a fairness requirement, not an administrative preference.
Related Lessons
- Escalation Paths: When to Involve Legal, Compliance, DEI, or Leadership takes over the moment one of your criteria fires on a fairness or legal question, covering which function owns which kind of decision and what evidence to bring each one.
- Escalation Processes: How Concerns Flow Up and Decisions Get Made is the machinery a triggered escalation travels through, giving it an owner, a clock, a documented decision, and an answer that comes back.
- Where Humans Remain Essential: Judgment, Context, and Nuance makes the case for what the mandatory-review band is actually buying you, which is the judgment a score cannot encode.
- Judgment Calibration: Building Intuition about AI Confidence is the skill that lets a reviewer use the borderline band well rather than deferring to the number in front of them.
- Auditing AI-Assisted Decisions: Sampling Methodology and Fairness Metrics develops the spot-check and the four-fifths monitor into a full audit design.
- Documentation Standards: What to Document and How Much Detail covers how much to record for each decision type, which is the fifth field in every rule you write here.
Closing
Decision rules are the scaffolding of fair, efficient hiring, and what makes them worth the afternoon they cost is that they are written while nothing is at stake. When the borderline candidate is in front of you, or the executive referral arrives, or the model quietly files a qualified applicant into the reject queue, the moment is too fast and too loaded for a good decision about who should be deciding. That question has to be settled in advance or it gets settled by whoever is most uncomfortable.
The practice is compact enough to hold in your head: map your decisions, define who should decide each one, write escalation criteria specific enough that a stranger could apply them, document the rules and share them, and audit for consistency. Dana's version of that work produced a table, three bands, one hard override, a short list of categories that always reach a human, and a quarterly four-fifths check. What it bought her was not caution. It was speed with a floor under it, and a known number of reviews she could staff, and the ability to answer the question every regulator and every rejected candidate eventually asks, which is not "was your model accurate" but "who decided, and how."
Key Takeaways
- Design the triggers in advance, do not leave them to chance. At 5,000 applications a quarter, "a recruiter will notice" is not a control. Define the borderline and mandatory-review cases up front and route them to a human automatically, so the cases that most need judgment are the ones that actually get it.
- Every rule needs five fields. Who decides, what information they need, what triggers escalation, what timeline applies, and what documentation is required. A rule missing any one of them fails in a predictable way, and the missing field is usually the timeline or the documentation.
- Bucket by score, but let qualifications override the score. Auto-reject only the clearly low band, send the borderline band entirely to humans, and spot-check the high band at 10 percent. Above all, never let a candidate who meets every posted minimum qualification receive an automated final rejection.
- Run the numbers so review is staffable. On an illustrative 5,000-application distribution the rules concentrate roughly 1,570 reviews, about 31 percent, on the decisions where human attention changes outcomes. A known number you can staff beats an invisible backlog you discover during a complaint.
- Some categories always escalate, regardless of score. ADA accommodation requests trigger a human-led interactive process. Protected and veteran status must never be inferred by the model. Conflicting data, anomalies, and any final rejection of a qualified candidate all require a person in the loop.
- Escalation criteria must be specific enough for a stranger to apply. "Missing two required skills" is testable; "if you're not sure" is not, and vague criteria escalate everything, which is the same failure as having no rules. Fairness and legal concerns always escalate to HR before any decision is communicated.
- Clear authority makes hiring faster, not slower. Push each decision to the lowest level that can hold it, with defined thresholds for the step up. The result is faster and more consistent decisions, clearer accountability, earlier detection of fairness issues, and less second-guessing.
- Monitor the pattern with the four-fifths rule. Per-case escalation watches the individual decision; a four-fifths selection-rate check across groups every cycle watches the rule itself, catching potential adverse impact before it compounds across thousands of candidates.
- Write it down, share it, and audit for consistency. Document the table, overrides, and thresholds with a named owner, and check quarterly that similar decisions were made at similar authority levels. Under EEOC guidance the employer owns the outcome even when a vendor built the tool, and laws like NYC Local Law 144 expect meaningful human involvement, not a rubber stamp.
Frequently Asked Questions
Will 31 percent human review defeat the point of automating the screen? No, and the comparison that makes it feel that way is the wrong one. The alternative is not zero human review, it is unplanned human review, where recruiters open whatever files they happen to open and rejected candidates get no eyes at all. The rules do not add work so much as redirect it, concentrating attention on the borderline band, the qualified-but-low-scored candidates, and a sample of the auto-advances, while the clearly unqualified and the clearly strong move without a queue. The number also becomes something you can plan against, which is the operational difference between a staffed process and a backlog you discover when a complaint arrives.
How specific does an escalation criterion have to be? Specific enough that two people applying it to the same candidate reach the same answer without discussing it. That is the whole test. "If the candidate does not meet two or more explicit role requirements, escalate" passes. "If there's any concern, escalate" fails, and it fails in the expensive direction, because everything can be described as a concern, so everything escalates and nothing gets decided at normal levels. If you cannot make a criterion testable, that is usually a sign the underlying decision has not been thought through rather than a sign the situation is inherently subjective.
What if my hiring managers resist giving up approval rights? Show them what the approvals are actually costing and what they are actually catching. The argument for pushing authority down is not that oversight is unnecessary; it is that undifferentiated oversight is weak oversight, because a senior reviewer looking at every offer is not looking closely at any of them. A band structure, where offers within the band are the manager's call, band-ceiling offers need director approval, and above-band offers need VP approval, keeps genuine control on the decisions that carry real exposure and returns the routine ones to the people closest to them. That usually reads as tighter control rather than looser once it is laid out.
Can the model just flag protected characteristics so we can monitor fairness? No, and the distinction matters. The model must not infer protected characteristics such as veteran status, disability, age, or national origin from resume content, and any sign that it is doing so is a defect to fix rather than a feature to use. Where protected-status information is legitimately collected, such as voluntary self-identification for affirmative-action reporting, it is firewalled from the screening decision, and any decision touching it escalates to a human. Aggregate fairness monitoring is a separate activity that runs on outcomes, which is what the four-fifths check does, rather than on individual-level inferences fed back into scoring.
How do I know whether my rules are actually being applied? Audit for consistency on a fixed cadence rather than waiting for a complaint. Take recent decisions of the same type and check whether they were made at the same authority level and with the same information; where they diverged, find out why. A documented rule applied inconsistently is worse than no rule at all, because you have created a written standard and then missed it. Treat every inconsistency you find as a defect in the rule, usually a criterion that was less specific than it looked, rather than as a failing of the person who applied it.
Skill.re