Common AI Errors in Recruiting: Hallucinations, Misinterpretations, Omissions
Marcus runs candidate screening for a 250-person healthcare staffing firm, and on a typical week he reviews roughly 40 resumes and 12 phone-screen write-ups for travel-nurse and allied-health placements. He started using an AI assistant to summarize his interview notes, and for two months it felt like a gift: a clean paragraph in fifteen seconds instead of fifteen minutes of typing. Then a hiring manager at a client hospital called him, irritated. The summary Marcus had forwarded said a candidate held an active "BLS and ACLS certification." The candidate held neither. The notes Marcus had dictated said only that she was "comfortable in acute settings." The AI had not lied on purpose. It had filled a gap. That phone call is the entire reason this lesson exists: AI in recruiting does not usually fail loudly. It fails confidently, in ways that look exactly like good work, and the cost lands on a real candidate and a real hiring decision.
Why AI Errors Are Not Like Human Errors
When a tired recruiter makes a mistake, it usually looks like a mistake: a typo, a blank field, a note that trails off. When a large language model makes a mistake, it looks like competence. The grammar is clean, the tone is professional, and the fabricated detail sits in the same confident sentence as the true ones. This is why AI errors are dangerous in hiring specifically: they pass the eyeball test. A reviewer skimming for obvious problems finds nothing, because there is nothing obvious to find.
It helps to understand the mechanism. A language model predicts the most plausible next words given everything before them. It is not a database that returns "not found" when a fact is missing. When the source material is thin, it does what it was trained to do everywhere else: produce fluent, plausible text. Plausible text about a nurse includes certifications; plausible text about an engineer includes a team size, so the model supplies one. This is not a bug you can patch away. Reviewing AI output critically means designing around it rather than wishing it away.
Notice what that implies about who does this well. The strongest recruiters using these tools are not the ones who trust the output blindly, and they are not the ones who refuse to use it at all. They are the ones who know precisely where the technology breaks down, build prompts and processes that prevent those specific breakdowns, and catch what slips through before it reaches a hiring decision. Error awareness is not skepticism for its own sake. It is the first step of error prevention, and it is learnable because the failures are predictable.
AI in recruiting fails in three recognizable shapes: hallucinations (invented detail that was never there), misinterpretations (the right words read with the wrong meaning), and omissions (information that mattered, silently dropped). Marcus learned to scan for all three. We take them one at a time.
Hallucinations: Invented Detail That Reads as Fact
A hallucination is information in the output that was never in the input. It is the error that bit Marcus, and it is the most legally dangerous of the three because it manufactures facts that a hiring decision can rest on.
Models hallucinate most reliably when you ask a specific question the source does not answer. Ask "how many years of experience does this candidate have?" against notes that say "several years," and a model will frequently return a number: six, eight, ten. It is invented, but it arrives in a confident sentence. The same happens with team sizes ("led a team of five"), credentials ("holds a PMP certification"), and enthusiasm ("passionate about machine learning") the candidate never claimed. The thinner your source notes, the more the model pads, because there is more empty space for plausible filler to occupy.
Consider a worked version of what Marcus hit. His dictated notes read: "Candidate has acute-care background. Comfortable in fast-paced units. Mentioned wanting flexible scheduling." He prompted: "Summarize this candidate for the ICU travel role." The AI returned: "Experienced ICU nurse with 8 years in acute care, holds active BLS and ACLS certifications, seeking flexible scheduling." Three fabrications in one sentence. "ICU" was never stated, only "acute care." "8 years" was invented from "background." The two certifications appeared from nothing. The only true facts were the acute-care background and the scheduling preference. Every concrete, checkable detail the hiring manager would actually weigh was made up.
The pattern is not specific to clinical roles. Run the same failure on a technical req and it looks like this. The notes from a first-round screen read: "Has worked in backend engineering. Machine learning came up in a recent project. Has worked independently for most of their career." The prompt was nothing more than "summarize this interview for an engineer role." What came back: "This engineer has 15 years of experience, with deep expertise in machine learning, and led a team of 10 people." Three inventions again, and the third one inverts the source outright, turning a candidate who worked independently into a manager of ten. A hiring manager reading that summary would form an entirely wrong picture of the person, and would be annoyed for the wrong reasons when the interview did not match.
You catch hallucinations by treating every specific fact in the output as a claim to verify against the source. Years of experience, team sizes, certifications, titles, named technologies: each must trace back to a real line in the notes. If the summary says "8 years" and the notes say "background," that gap is a hallucination, full stop. You prevent them by instructing the model to stay inside the source: "Use only information explicitly stated in the notes below. Do not infer years of experience, certifications, or seniority. Where a fact is absent, write 'not stated' rather than estimating." Adding "mark any inferred information as inferred" gives you a second layer, converting the model's guesses into labeled guesses. That last instruction gives the model a sanctioned way to leave a blank instead of filling it.
Misinterpretations: The Right Words, the Wrong Meaning
A misinterpretation is subtler than a hallucination. The model reads the words accurately, but it assigns a meaning the words do not carry. This is where AI errors cross directly into fairness and discrimination risk, because misinterpretations tend to land on personality, communication style, and motivation, and those readings often encode stereotypes rather than evidence.
The pattern is consistent. A note says "candidate was quiet in the group exercise," and the AI summary reads "poor communication skills" or "may struggle in collaborative settings." A note says "asked detailed questions about the compensation structure," and the summary reads "primarily motivated by money." A note says "mentioned valuing work-life balance," and the summary reads "may not be fully committed to the role." In each case the model took a neutral observation and converted it into a negative character judgment that the recruiter never made. It does this because, statistically, those words co-occur with those judgments in its training data, not because it understood this candidate.
Here is one in full, because seeing the two texts side by side is what makes the error stick. The interviewer's notes read: "Candidate is introverted. Asked thoughtful questions when they did engage." The summary came back: "Candidate seems disengaged. Asked few questions." Nothing was fabricated in the sense of the previous section, and yet the summary is false in every way that matters. "Introverted" became "disengaged," a judgment about motivation drawn from a description of style. "Asked thoughtful questions" became "asked few questions," a quality observation flattened into a quantity complaint. A hiring manager reading only the summary would decline the candidate, and would never know that the reason for declining came from the tool rather than from the interview.
The stakes here are real. If your AI consistently reads "quiet" or "soft-spoken" as a deficiency, and those readings influence who advances, you can build a pattern that disadvantages candidates whose communication style correlates with neurodivergence, introversion, English as a second language, or cultural background. The EEOC has been explicit that employers remain responsible for discriminatory outcomes from automated and AI-assisted hiring tools, and the ADA framework treats screening practices that disadvantage protected groups as the employer's liability regardless of whether a vendor's algorithm produced them. An AI that turns "introverted" into "disengaged" is not just inaccurate; it is the kind of inaccuracy that creates legal exposure if it shapes decisions at scale.
You catch misinterpretations by checking whether each judgment is an observation or an interpretation, and whether the interpretation matches what you actually saw in the room. Your own impression of the conversation is the reference standard here, and if the summary's read of a candidate does not match yours, the summary is wrong until proven otherwise. You prevent them by forbidding the leap: "Report what was observed in neutral terms. Do not infer motivation, engagement, personality, or cultural fit from communication style, and do not characterize a candidate as disengaged, uncommitted, or money-motivated unless the notes state it directly. Distinguish observation from interpretation, and label any interpretation as inferred. For any judgment, cite the specific note it rests on." Requiring a citation for every judgment is the strongest single safeguard, because a model that has to point at evidence cannot quietly substitute a stereotype.
Omissions: The Error You Cannot See
Omissions are the hardest of the three to catch because there is nothing on the page to flag. A hallucination is a wrong fact you can confront; an omission is a right fact that is simply gone. The summary looks complete, reads well, and quietly leaves out the thing that mattered most.
Omissions cluster around length limits and prioritization. Ask for a 200-word summary of a 45-minute interview that covered a technical screen and a team-fit round, and the model will compress. Compression means choices, and the model makes them to optimize for fluent prose, not for your decision. It may drop the team round entirely, flatten a candidate's junior-to-senior growth into one job title, or replace specific evidence for a skill with a vague rating. A particularly corrosive variant is the half-reported concern: the flag survives the compression but the context that explained it does not, so a hiring manager reads "interviewer raised a concern about pace" with no trace of the sentence that said the candidate had been covering two open roles at the time. A concern without its context is not information; it is an accusation. The most dangerous omission is the dropped disqualifier: a flagged employment gap, an expired license, a "would not rehire" reference comment, a relocation restriction. When the model omits a disqualifier, you are deciding on information you do not know is incomplete, and acting on an omitted gap or lapsed credential carries the same downstream risk as acting on a hallucinated one.
Suppose Marcus's notes for one candidate included a single line: "License lapsed in 2023, renewal in progress, cannot start until cleared." He asked for a tight three-sentence summary. The model produced a clean, positive blurb about acute-care experience and availability and never mentioned the license. Nothing in the output was false. The output was simply missing the one fact that determined whether this candidate could be placed at all. A reviewer reading that summary has no signal that anything is absent.
You catch omissions by comparing the summary against the source for completeness, not just accuracy, with special attention to anything that could disqualify or constrain. A fast version of that check is to ask yourself what the interviewers emphasized, the points they returned to or wrote in capital letters, and then confirm each one survived into the summary. You prevent them by removing the pressure to compress and naming what must survive: "Do not use a word limit. Summarize completely, prioritizing technical depth, then team fit, then concerns. You must include, if present in the notes: any licensing or certification gaps, employment gaps, availability or relocation constraints, and any concern raised by an interviewer along with the context that explains it. If something is missing that would help the hiring team decide, include it. Never drop a stated disqualifier to save space." Telling the model that disqualifiers are non-negotiable turns a silent drop into a preserved fact.
Prevention by Design: Prompts That Catch Their Own Errors
Detecting errors after the fact is necessary but exhausting at 40 resumes a week. The better leverage is designing prompts so common failure modes are blocked before they happen. Three layered instructions handle most of it.
The first is the explicit-versus-inferred distinction: require the model to label each claim as stated directly in the source or inferred, with reasoning attached to anything inferred. This converts hallucinations from invisible filler into visible, labeled guesses you can accept or reject. The second is the evidence requirement: every meaningful claim about a skill, a trait, or a fit must cite the specific line it rests on, which cures misinterpretation because a claim that must point at evidence cannot smuggle in a stereotype. The third is scope clarity: state which dimensions must be covered and that completeness beats brevity, which cures omission because the model can no longer drop a dimension to save space.
A prompt carrying all three reads roughly like this: "Summarize the notes below for a hiring decision. Label each claim as 'stated' or 'inferred'; for inferred claims, give your reasoning. Cite the specific note behind any judgment about skills, fit, or character. Cover technical depth, team fit, and any concerns or constraints completely; do not omit a stated disqualifier to save space; do not impose a length limit. Where a requested detail is not in the notes, write 'not stated.'" It is longer than "summarize this interview." It is also the difference between output Marcus forwards with confidence and output that produces an angry phone call.
Three Anti-Patterns
Each error type has a corresponding bad habit, and the habits are what actually cause the damage. Naming them makes them easier to catch in yourself and in your team.
- Ignoring hallucinations. This is accepting an AI summary without verifying its specific facts, usually because the prose reads well and the day is long. It fails because hallucinations compound: one invented certification becomes a shortlist decision, becomes a client submission, becomes a phone call from a hiring manager. The fix is mechanical rather than clever, which is to spot-check every specific fact against the source before the summary leaves your hands.
- Assuming good interpretation. This is trusting the model's read on soft skills, motivation, or fit without asking what it rests on. It fails because those interpretations are frequently drawn from stereotype rather than evidence, and stereotype is precisely the thing that turns an accuracy problem into a discrimination problem. The fix is to mark every interpretive claim as interpretive and require the evidence behind it.
- Accepting incomplete information. This is making a decision from a summary that has silently dropped a whole dimension, most often soft skills or team fit, sometimes a constraint that changes everything. It fails for the obvious reason: you are deciding without complete information and you do not know it. The fix is to specify up front which dimensions must appear and then review the output for completeness rather than only for accuracy.
Keeping a Human in the Loop
No prompt is airtight, and the goal is not to make one. The goal is a human review step where the cost of an error is highest, made fast by knowing exactly what to look for. AI handles the volume; the recruiter owns the judgment and the accountability. That division is not just good practice, it is the legal posture regulators expect: the employer, not the tool, answers for the decision.
Marcus built a thirty-second checklist he runs on every AI summary before it leaves his hands. Verify each specific fact, especially numbers and credentials, against the notes. Confirm that every character or fit judgment is tied to a real observation, not a stereotype dressed as analysis. Scan the source once more for anything missing, with a hard look for disqualifiers and constraints. Treat anything the AI labeled "inferred" as a question for the candidate, not a finding. Over a few weeks the checklist became reflex, and the summaries he forwards now carry his judgment, not just the model's prose. The AI drafts, the human decides, and the line between the two never blurs.
The Vocabulary, in One Place
Four terms carry this lesson, and being able to name an error precisely is what lets you report it usefully to a colleague or a vendor.
- Hallucination. The model generating information that was not in the source and may not be true.
- Misinterpretation. The model reading the words correctly but assigning them the wrong meaning through pattern matching.
- Omission. The model leaving out information that mattered for the decision.
- Explicit versus inferred. The distinction between information stated directly in the source and information the model derived from context. Making the model declare which is which is the single most useful instruction in this lesson.
Practice
Work these on real output from your own pipeline, not on invented examples. The point is to discover which error type your particular workflow produces most.
- Identify hallucinations. Take one AI summary and, for every specific fact in it, years of experience, team size, credentials, named achievements, trace the fact back to the source material. List anything that does not trace.
- Catch misinterpretations. Take another summary and mark every interpretive claim: anything about motivation, personality, or engagement. For each one, check whether the interpretation actually matches the evidence in the notes.
- Check for omissions. Review a summary against the full source and ask what important information did not survive. Technical detail? Soft skills? The context behind a concern? A constraint on availability?
- Improve a prompt. Take a prompt that has produced errors and add the three safeguards: the explicit and inferred distinction, the evidence requirement, and scope clarity. Run the improved prompt on the same source and compare the two outputs.
- Build your own error checklist. Using the errors you actually found in the first three exercises, write a short review checklist tailored to the mistakes your tool and your workflow produce most often. A checklist built from your own findings gets used; a generic one does not.
Reflection
- Which of the three error types have you already experienced? When did it happen, and how did you catch it, or how did you find out you had not?
- If you reviewed the last month of your AI-assisted recruiting with this framework, what would you expect to find?
- Which error type is most dangerous in your specific hiring context, and why? For a compliance-heavy placement the answer is often omissions; for client-facing submissions it is often hallucinations. Your answer should shape where you spend your review time.
Putting It to Work This Week
Understanding these three error types is what lets you design prompts and processes that prevent them, and prevention is what makes the tool genuinely trustworthy rather than merely fast. An AI you have engineered against its own failure modes is a different instrument from one you are simply hoping about.
So pick one. This week, identify the single error type that most affects your hiring, build a specific prevention strategy for it, whether that is a prompt change, a checklist step, or a review handoff, and then test it on live work. One error type fully prevented is worth more than three partially understood.
Related Lessons
- Hallucinations, Accuracy Errors, and Information Loss goes deeper on why models fabricate and lose detail, filling in the mechanism behind the first and third error types here.
- Verification Techniques: Spot-Checking Facts, Sources, and Candidates turns the fact-checking habit in this lesson into a repeatable method, including how much to check and when sampling is enough.
- Hands-On Practice: Review AI Summaries, Identify Errors, Revise is the applied version of the practice exercises above, working through real summaries end to end.
- Red Flags and When to Reject or Escalate AI Output picks up where detection stops: what to do once you have found an error that matters, and when a problem needs to leave your desk.
- Feedback Loops: How to Report AI Errors and Improve System Performance covers reporting what you find, so the errors you catch improve the system rather than dying in your inbox.
- Designing Guardrails: What AI Can Do, What Requires Human Approval sets the boundaries that decide which outputs need the human review step described here, and which can move without it.
Key Takeaways
- AI errors in recruiting fail confidently, not obviously. Fabricated certifications and invented years of experience arrive in clean, professional sentences that pass a quick skim. Treat fluency as no evidence of accuracy, because the model produces plausible text whether or not the facts behind it are real.
- Hallucinations are invented detail; verify every specific fact. Numbers, credentials, titles, and team sizes must trace to a real line in the source. Models pad most when notes are thin, so the sparser your input, the harder you check, and the more you instruct the model to write "not stated" instead of estimating.
- Misinterpretations are a fairness risk, not just an accuracy one. When AI turns "quiet" into "disengaged" or "asked about pay" into "money-motivated," it encodes stereotypes that can disadvantage protected groups. The EEOC and ADA frameworks hold the employer responsible for discriminatory outcomes from AI tools, so require evidence for every character judgment.
- Omissions are the error you cannot see on the page. A dropped disqualifier, an omitted employment gap, or a lapsed license can sink a placement, and a clean-looking summary gives no signal it is incomplete. Forbid word limits, name the dimensions that must survive, and tell the model never to drop a stated disqualifier.
- Design prompts that block errors before they happen. Layer three instructions: label claims as stated or inferred, cite evidence for every judgment, and require complete coverage over brevity. A longer prompt that prevents fabrication is worth far more than a short one that invites it.
- Keep a human at the highest-cost decision point. AI drafts at volume; the recruiter verifies facts, tests judgments against evidence, and scans for what is missing before anything reaches a hiring manager. The accountability for the decision stays with the human, which is exactly where the law places it.
Skill.re