←
AI for Recruiters
Capable · M23 · lesson 23 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Scenario Practice: Review and Critique

15 min

Marcus is a senior recruiter on a four-person talent acquisition team at a 600-person fintech company, and he runs roughly 40 candidate assessments a week through an AI assistant that drafts summaries from his interview notes. The AI saves him hours. It also, about once a day, produces something that would embarrass the company or sink a candidate unfairly if it shipped unreviewed. This lesson is the workshop where Marcus stops trusting the polish of AI prose and starts critiquing it like an editor who can be held liable for what goes out under his name.

Why Critique Is the Real Skill

The hard part of using AI in recruiting is not writing the prompt. It is reading the output with enough suspicion. AI-drafted candidate summaries are fluent, confident, and grammatically clean, which is exactly what makes them dangerous. A summary that reads like a polished assessment can contain a hallucinated degree, an inflated tenure, or a coded bias that a tired reviewer skims past because the sentence "sounds right." Marcus learned this the hard way when a draft he nearly forwarded described a candidate as having "led a team of five for two years" when the source notes mentioned no direct reports at all.

Critique is a structured habit, not a vibe check. Marcus runs every draft through the same three moves: identify the issues, rate their severity, and decide what to do. The goal is not to catch everything in one heroic read. It is to make the review consistent enough that the same flawed sentence gets the same treatment whether it lands in his inbox on Monday morning or Friday at five.

Theory is one thing, and practising on real material is another. Everything that follows is a workshop rather than a lecture. You will run the full cycle four times: read the output, identify the issues, decide whether to use, escalate, or reject, and then write the feedback that makes the next draft better. The cycle is not complete until the feedback exists, because a rejection nobody explains teaches nobody anything.

The Three-Question Framework

Marcus reviews each draft against three questions in order. First, what issues are present? He scans for five categories: hallucinations (facts inflated or invented), bias (assumptions tied to protected characteristics), missing evidence (claims with no supporting example), vague language (assessments that could describe anyone), and omissions (information that should be there and is not). Second, how severe is each issue? He rates it critical (changes the hiring decision), major (notable but potentially fixable), or minor (does not affect the decision). Third, what is the verdict? Three outcomes: USE the output, ESCALATE it to the hiring team, or REJECT it.

The ordering matters. Severity only makes sense after the issues are named, and the verdict only makes sense after severity is assigned. Marcus found that recruiters who jump straight to "is this good enough?" accept biased output because the overall draft felt competent. Separating identification from judgment forces each problem to be seen on its own before it gets averaged into a gut feeling.

It is worth being precise about what each verdict means, because the whole framework collapses if the three outcomes blur together. REJECT is for output carrying multiple critical issues, unfixable bias, or factual claims you cannot trust. ESCALATE is for contradictions, unclear signals, fairness questions, and any decision that genuinely needs human judgment rather than a recruiter's private call. USE is reserved for clean output that carries clear information and no red flags. If you find yourself reaching for a fourth category, such as "fine, I suppose," you have found the gap where standards quietly slip.

Scenario One: The Glowing Summary That Says Nothing

Marcus opens a draft for a product manager candidate. It reads: "Sarah is an excellent product manager with 12 years of experience. She is passionate about building world-class products and has strong leadership skills. She led teams of up to eight people and has excellent communication skills. Strong cultural fit for our dynamic, fast-moving team."

It sounds great. That is the problem. Running the framework, Marcus flags six issues. "Excellent" and "world-class" are vague, they describe a verdict without the evidence behind it. "Passionate" is unsubstantiated, nothing in the notes supports it. The "12 years" and "teams of up to eight" are specific numbers that must be verified against the resume before they are repeated, because a confident number is the easiest hallucination to miss. "Excellent communication skills" has no cited example. And "strong cultural fit for our dynamic team" is the most dangerous phrase in the paragraph, because "cultural fit" used without a defined rubric routinely becomes a proxy for "people like us," which is precisely the mechanism by which subjective hiring criteria produce disparate impact.

Two further faults are worth naming because they are properties of the whole draft rather than of any single sentence. Not one specific example or piece of evidence is cited anywhere in it. And the register is wrong: the paragraph reads like a sales pitch rather than an objective assessment, which is a signal in itself. When a summary is working to persuade you, it has stopped working to inform you.

On severity, the vagueness is major and the cultural-fit framing is critical. The verdict is REJECT and rewrite. Marcus sends feedback rather than silently fixing it: rewrite each claim to cite a specific example from the notes, replace "cultural fit" with the actual values-alignment behaviors observed, and remove any adjective that is not anchored to evidence.

Scenario Two: Spotting Hallucinations Against the Source

The clearest hallucinations are invisible unless you read the draft beside the source. Marcus pulls up a draft that states: "Alex graduated from Stanford with a degree in Computer Science. He has eight years of backend engineering experience, with deep expertise in distributed systems and microservices. He led a team of five engineers for two years. His GitHub shows 150 repositories with significant open source contributions."

The actual resume says: "BS in Computer Science, State University. Six years backend engineering experience. Worked on distributed systems projects at a previous role. Contributions to open source frameworks." Side by side, the draft falls apart. Stanford was invented in place of State University, a fabricated prestige signal. Eight years became the real six. "Deep expertise" inflated a plain "worked on." The two-year team lead role appears nowhere in the source, it is fully hallucinated. And "150 repositories" is a precise number conjured from a vague "contributions to frameworks." Five of these are critical because they are verifiable factual claims that are simply false. The verdict is REJECT, and the lesson Marcus writes into his prompt notes is to instruct the AI to quote or cite the source line for any factual claim, so unsupported facts have nowhere to hide.

One issue in that draft is subtler than the rest and easy to skip past. "Significant open source contributions" is not a hallucination in the same sense, because the resume does mention contributions. It is a vague claim resting on a vague source. The word "significant" carries a weight the resume never supports, so it belongs on the list too. Flag it, and ask what evidence would make "significant" a defensible word before anyone repeats it to a hiring manager.

This is also a verification discipline worth naming on its own. Any draft that mentions a school, a tenure in years, a team size, a certification, or a metric is making a checkable claim. Marcus treats every one of those as unverified until he has matched it to the resume or the notes. The polish of the sentence is not evidence. The source is.

Scenario Three: Bias, and Where the Law Sits

A third draft reads: "Jessica is a strong candidate but seems like she might prioritize family over work. She mentioned wanting flexibility, which suggests she is not serious about a senior role with demanding hours." There is nothing to fix here. The whole inference is the problem. "Might prioritize family" is an assumption about caregiving that maps onto gender stereotypes. Treating a request for flexibility as evidence of low commitment penalizes a protected pattern of behavior. And "not serious about a senior role" is a conclusion with zero supporting evidence. The verdict is REJECT, with no salvage, because the reasoning itself is discriminatory.

Notice what the draft would have cost if it had shipped. Every substantive claim in it is an inference, and the inferences are all pointed the same way. A candidate the draft itself calls strong would have been disqualified on assumptions about family and flexibility rather than on anything she did or said about the work. That is the shape of the harm: bias does not usually arrive as an insult, it arrives as a reasonable-sounding paragraph that quietly converts a stereotype into a rejection.

The legal stakes are real, and Marcus keeps them in view. Decisions driven by assumptions about caregiving, sex, or family status run directly into Title VII and, where disability is implicated, the ADA and EEOC guidance. The exposure compounds when the AI tool is part of an automated screening pipeline. Suppose Marcus's team used an AI scorer that quietly downranked candidates who mentioned flexibility, and over a hiring cycle it advanced 50 of 100 men and 30 of 100 women. The selection rates are 50 percent and 30 percent; 30 divided by 50 is 0.60, below the four-fifths (80 percent) threshold the EEOC uses as a rule of thumb for adverse impact. That gap is the kind of pattern that triggers scrutiny. And in New York City, an automated employment decision tool like that scorer falls under Local Law 144, which requires an independent bias audit and candidate notification before the tool is used. A biased summary is not just unfair to one candidate, it is a thread that, pulled, can unravel into systemic liability. Bias is always a REJECT or, where a human needs to weigh genuine ambiguity, an ESCALATE, never a "close enough."

Scenario Four: The Borderline Call Worth Escalating

Not every imperfect draft is a reject. Marcus reviews: "Tech fit: strong in Python and backend systems. Moderate experience with our specific cloud platform (AWS). Communication: good in technical discussions, quieter in group settings. Cultural fit: unclear. Concerns: limited experience with our specific tech stack. May need two to three months to ramp up." Running the framework, the issues are not flaws so much as honest open questions. The technical assessment is specific and supported. The communication note is tied to observed behavior. Crucially, the AI marked cultural fit "unclear" rather than fabricating a verdict, which is exactly the restraint Marcus wants. The concerns are concrete and reasonable.

This draft has escalation flags, not reject flags. The ramp-up estimate and the team-dynamics question are real decisions that need a human conversation, not a recruiter quietly accepting or discarding. The verdict is ESCALATE: forward it to the hiring manager with the open questions highlighted. Learning to tell an honest "this needs discussion" from a defective "this is wrong" is what separates a critic from a censor.

The Verification Checklist and the Anti-Patterns

Marcus distilled his practice into a checklist he runs before any AI summary leaves his hands. Are all named facts, schools, tenures, team sizes, and metrics matched to the source? Is every assessment claim backed by a cited example? Is there any language that infers from a protected characteristic? Has "cultural fit" been replaced with defined, observable behaviors? Does the draft read like an objective assessment rather than a sales pitch? If any answer is no, the draft does not ship as written.

He also watches for three anti-patterns that creep in under deadline pressure. The first is accepting "good enough," letting an 80 percent draft through when the standard is higher, because small unverified claims compound into a body of work nobody can trust. The second is going soft on bias, telling himself a flagged assumption is "not that bad," when bias is never a matter of degree. The third is rejecting without feedback, which leaves the AI producing the same flaw tomorrow. When Marcus rejects, he always says specifically what was wrong, because that feedback is what improves both the next draft and the prompt behind it.

Each anti-pattern has the same cure, applied at a different point. Against "good enough," hold output to a consistent standard and let the reject and escalate criteria do the deciding instead of your mood. Against permissiveness on bias, treat any bias flag as automatically disqualifying: reject or escalate, never accept. Against silent rejection, write the specific fault down, because the alternative is receiving the same bad output indefinitely while wondering why the tool never improves.

A Short Glossary

Four terms carry most of the weight in this lesson, and it helps to keep them sharp. Critique is the systematic evaluation of AI output for accuracy, bias, completeness, and reliability, as opposed to a general impression of whether it reads well. A hallucination is AI-generated information that is not in the source material, typically inflated or simply false, like the Stanford degree that the resume never mentioned. Bias, in this context, is discriminatory language or an assumption resting on a protected characteristic, of the kind Scenario Three turns on. And missing evidence describes a claim made without a supporting example or citation, which is the most common defect you will find and the easiest to send back for a fix.

Practice: Run the Cycle Yourself

Start by working through the four scenarios above on your own before rereading Marcus's analysis. For each one, list the issues you can find, rate their severity, and commit to a verdict of use, escalate, or reject. Comparing your list against his is more useful than agreeing with it, because the gaps show you which of the five issue categories you tend to skim.

Then bring in your own material. Take a real AI-generated candidate assessment from your current work and put it through the identical framework. Working on live output is where the habit forms, because your own drafts carry the deadline pressure that makes shortcuts tempting.

Next, practise the half of the cycle that recruiters usually skip. Choose an output you rejected and write the specific improvement feedback you would send: what was wrong, what evidence was missing, and what the corrected version would have to contain. Vague feedback produces vague drafts.

Calibration comes after that. Share your critiques with colleagues and compare verdicts. Do you agree on what should be rejected versus escalated? Disagreement is not a failure of the exercise, it is the point of it, because a team that quietly holds four different standards is a team whose candidate decisions vary by whoever happened to pick up the file.

Finally, turn what you have learned into an artifact. Build a personal critique checklist from the issues you actually found, in your own words, ordered by how often they show up in your work. Marcus's checklist is a starting point, not a template to copy, and yours will be sharper because it is built from your own rejections.

Reflection Questions

Which types of issue do you find most often in AI output: hallucinations, bias, or missing evidence? The answer tells you something about the tool, and something about your prompts.

When you critique output, are you usually rejecting, escalating, or accepting? Look at the distribution rather than any single call, because a pattern of near-universal acceptance and a pattern of near-universal rejection both suggest the framework is not being applied.

How much time do you spend reviewing AI output compared with using it? Is that ratio right for the stakes of the decisions the output feeds?

And if you shared your critiques with the person or team generating those outputs, what would they learn? If the honest answer is "quite a lot," that is an argument for sharing them.

Putting It Into Practice

Critical review is a skill that develops through repetition. The more you critique, the faster you spot issues, and the faster you spot issues, the less time you waste on low-quality output. Speed here is a product of practice, not of lowering the bar.

Give yourself a concrete target for the coming week: critique ten AI-generated candidate assessments and document the feedback for each. Then look across the ten and ask what patterns emerge in what you rejected. If you keep striking out vague cultural-fit language, that is not ten separate problems, it is one prompt that needs a constraint. Feeding rejection patterns back into your prompts is how the review workload shrinks over time without the standard moving.

Red Flags and When to Reject or Escalate AI Output sets out the decision criteria this workshop applies, so if the boundary between reject and escalate still feels blurry after Scenario Four, that is the lesson to revisit.

Verification Techniques: Spot-Checking Facts, Sources, and Candidates is the discipline behind Scenario Two. It goes deeper on how to check a claimed school, tenure, or repository count against the source material without spending your whole week verifying.

Common AI Errors in Recruiting: Hallucinations, Misinterpretations, Omissions explains why the failure modes you are hunting for take the shapes they do, which makes them faster to recognise in a draft.

Fairness Checks: Identifying Gender, Age, Disability, and Other Bias Signals extends Scenario Three, where the bias was overt. Much of what you meet in practice is quieter, and that lesson trains the eye for it.

Feedback Loops: How to Report AI Errors and Improve System Performance picks up the anti-pattern of rejecting without explanation and turns individual feedback into something the wider system can act on.

Team Agreements: Building a Culture of Responsible AI Use is where the calibration exercise leads. Once you and your colleagues have compared verdicts, a written agreement is what keeps those standards steady.

Where This Goes Next

You are now equipped to review AI output critically and to make principled decisions about when to use it, when to send it up, and when to throw it out. That is an individual competence, and on its own it protects only the candidates whose files happen to cross your desk. The next step is scaling it: taking the frameworks in this lesson to the rest of your team and building the systems that support responsible AI use at scale, so the standard holds regardless of who is reviewing.

Key Takeaways

  • Critique is the skill, not prompting. AI drafts are fluent and confident, which is exactly what hides hallucinations, bias, and empty claims. Read every summary like an editor who is accountable for what ships under their name.
  • Run a fixed three-question framework. Identify the issues (hallucination, bias, missing evidence, vague language, omission), rate severity (critical, major, minor), then decide (USE, ESCALATE, REJECT). Naming issues before judging them stops a polished draft from masking a serious flaw.
  • Verify every checkable claim against the source. Schools, years of tenure, team sizes, and metrics are the easiest hallucinations to miss because a confident number reads as fact. Treat each as unverified until matched to the resume or notes.
  • Bias is a hard reject, and the legal stakes are real. Inferences tied to caregiving, sex, or disability implicate Title VII, the ADA, and EEOC guidance. In a screening pipeline, a 30-versus-50-percent advancement gap fails the four-fifths rule, and in NYC such tools fall under Local Law 144's bias-audit requirement.
  • Replace "cultural fit" with observable behavior. Used without a rubric, "fit" becomes a proxy for "people like us" and a driver of disparate impact. Demand the specific values-aligned behaviors that were actually observed.
  • Escalate honest ambiguity instead of rejecting it. A draft that openly flags an unclear area and raises concrete questions is doing its job. Forward it to the hiring team rather than quietly accepting or discarding the decision.
  • Always reject with feedback. Silent fixes leave the AI repeating the same flaw. Saying specifically what was wrong improves the next draft and the prompt that generated it.