Verifying Candidate Information: Spotting Hallucinations and Inaccuracies
Reuben sources senior data scientists for Aperture Analytics, a 250-person firm where he works largely alone and largely passively, building candidate profiles from public signals before anyone ever applies. AI accelerates that work enormously, and that is exactly the danger. When you synthesize a candidate from scraped fragments, the AI fills the gaps with fluent invention, and there is no resume on file to catch it. The week Reuben presented a "Kaggle Grandmaster who led ML platform at a unicorn" and the hiring manager found a Kaggle Expert who had been one of forty engineers on a platform team, he stopped treating AI synthesis as research and started treating it as a draft that has not yet been verified.
Why Synthesized Candidate Profiles Hallucinate
A hallucination is a confident, fluent claim with no basis in fact. In sourcing it is especially insidious because the inputs are sparse and the model is rewarded for coherence. Given a name, a job title, and a few conference mentions, an AI will produce a tidy narrative: years of experience, scope of impact, seniority of role. Each detail is plausible, the prose is clean, and none of it was verified, because for a passively sourced candidate there is no application packet to verify against. Reuben's job is to treat every synthesized profile as a set of claims to be confirmed, not a set of facts to be forwarded.
It helps to see why sourcing is a worse case than screening. When a candidate applies, the resume is an artifact the candidate produced and stands behind; if it overstates something, the overstatement is theirs, and the interview is the natural place to test it. A sourced profile has no such author. The model assembled it from fragments that were never meant to be read together, and the finished paragraph does not mark which sentences came from a source and which were bridged across a gap. Fluency is uniform across both kinds, which is precisely why reading the output more carefully does not help.
The second mechanism is that a language model has no reliable way to report the absence of information. Asked for a candidate summary, it produces a candidate summary. Where the record is thin, the most fluent available continuation is a reasonable-sounding specific, and reasonable-sounding specifics are what a summary is made of. Reuben noticed that the gaps in the public record were exactly where the most confident sentences appeared, because those were the sentences the model had the most freedom to construct. Sparse input does not produce a hedged output; it produces an invented one delivered in the same voice as the verified parts.
There is a third mechanism, and it belongs to the recruiter rather than the tool. An inflated profile is easier to sell. A hiring manager who is triaging twenty sourcing notes responds to the one with a leadership title and a named credential, so the inflated version is the version that gets a reply, gets a screen, and survives. Nobody has to consciously prefer the exaggeration for the exaggeration to be selected for. This is why verification has to be a step in the process rather than a matter of personal skepticism, since the incentive runs against skepticism at exactly the moment it matters.
The Three Shapes a Hallucination Takes
Reuben learned to recognize three recurring shapes. Naming them matters because a shape tells you where to look: each one fails in a characteristic way and each one has a characteristic source that settles it. Once he could sort a synthesized paragraph into these three buckets in under a minute, verification stopped being a vague instruction to be careful and became a short list of specific things to open.
The Inflated Number
The first shape is the inflated number: "8 years of experience" when the public record supports 5, or "led a team of 20" when the person was one contributor among twenty. Specific quantities are the most fabricated category because precision is what makes a summary sound authoritative. A range invites a question; a number closes one. Watch particularly for years of experience, team size, budget or scale figures, and any percentage of improvement attributed to the candidate's work, because those are the four quantities a hiring manager will repeat out loud in a debrief and therefore the four that do the most damage when they are wrong.
The tell is that the number is rarely traceable to a single visible statement. A profile that says a candidate has eight years of experience is almost never quoting a sentence where the candidate said so; it is arithmetic performed on partial dates, and arithmetic performed on partial dates is where the inflation enters. Reuben's habit is to reconstruct the count himself from the dated entries on the candidate's own profile, and to notice specifically whether the model has counted adjacent experience as if it were the experience the role actually requires.
The Upgraded Role
The second shape is the upgraded role: "led the ML platform" when the truth is "contributed to the ML platform team." The model rounds participation up to leadership because leadership reads better. The verbs are where this lives. "Worked on" becomes "led." "Supported" becomes "owned." "Was part of the team that shipped" becomes "shipped." None of these substitutions is a lie about which project the person touched, which is what makes the upgrade so hard to see. The project is right, the company is right, the dates are right, and the only false element is the size of the person's role inside it.
The upgraded role is also the most consequential shape, because seniority is usually the thing the requisition is actually screening for. A senior req that opens on a candidate presented as having led a platform will run an interview designed to probe leadership decisions the candidate never made. The candidate looks weaker than they are, the hiring manager concludes the sourcing is poor, and the only party who behaved badly is the summary. Reuben checks role claims against the candidate's own wording rather than the model's paraphrase of it, because the candidate's own wording is usually careful about exactly this distinction.
The Invented Attribute
The third shape is the invented attribute: a certification, a publication, or an open-source project that simply does not exist, stitched together from adjacent facts the model found nearby. Someone who spoke at a conference acquires a paper at that conference. Someone who works in a field where a certification is common acquires the certification. Someone whose colleague maintains a well-known repository acquires a contribution to it. The invented attribute is drawn from the neighborhood of true facts, which is why it survives a plausibility check: it is exactly what a person like this one would plausibly have.
Invented attributes are the easiest of the three to kill, because a named artifact either exists at a findable address or it does not. That is also why they are worth checking first. A single failed lookup on a claimed credential tells you the profile was constructed rather than compiled, and that finding should raise your suspicion about every number and every verb in the same paragraph, since they came out of the same process.
| Shape | What it looks like | Where it is settled |
|---|---|---|
| Inflated number | "8 years of experience" against a record supporting 5; "led a team of 20" when the person was one of twenty | Rebuild the count yourself from dated entries on the candidate's own profile |
| Upgraded role | "Led the ML platform" when the record says "contributed to the ML platform team" | The candidate's own description in their own words, not the model's paraphrase of it |
| Invented attribute | A certification, publication, or repository assembled from nearby true facts | The issuing body, the venue's own site, or the repository itself, opened directly |
A Verification Protocol for Sourced Candidates
Because Reuben sources at volume, he runs a protocol rather than improvising. First, he separates the verifiable from the unverifiable: every specific number, role, and named artifact goes on a list. Second, he checks each against an independent public source. A claimed GitHub is opened and read, not counted; a claimed publication is found on the actual venue's site; a claimed role is cross-checked against the person's own LinkedIn description rather than the AI's paraphrase of it. Third, anything that cannot be confirmed from a source is demoted from "fact" to "question for the candidate" and never presented to a hiring manager as established.
The first step does more work than it appears to. Writing the claims down as a list breaks the spell of the paragraph, because a paragraph reads as one continuous assertion while a list reads as five separate ones, four of which you have no evidence for. It also forces a distinction that saves time later: some claims are verifiable in principle, such as a credential tier or a publication count, and some are not verifiable from any public source, such as why someone left a role or how well they handled a difficult stakeholder. The second category is not a verification failure. Those are interview questions by nature, and marking them as such stops Reuben from hunting for a source that was never going to exist.
The second step turns on what "independent" means, and the word is doing real work. A source is independent when it is not the same fragment the model already consumed, and when you open it yourself rather than accepting a description of it. If the model summarized a profile page and you check the claim against the model's summary of that page, you have confirmed nothing except that the model is internally consistent. Reuben's rule is that the verification has to touch the primary artifact: the issuing body's own record for a credential, the conference's own proceedings for a paper, the repository itself for a contribution, and the candidate's own account of their own role for a role.
The third step is the one people skip, because demoting a claim feels like weakening your own pitch. It is the opposite. An unconfirmed claim that reaches a hiring manager as a fact is a liability with a delayed fuse; the same claim delivered as "the summary asserts platform leadership, which I could not confirm, so it is first on my list for the screen" is useful, honest, and makes Reuben look like someone who knows the difference. Nothing gets deleted. It gets relabeled and moved to the place where a human can resolve it, which is the conversation.
When Reuben checked his Grandmaster claim against Kaggle's public profile page, the tier read "Expert," and the whole inflated narrative collapsed from one source check. That is the general pattern worth internalizing: hallucinated profiles are usually not a scatter of small errors spread evenly across a summary. They are a coherent inflated story, and coherent stories fail at a single load-bearing point. Check the most specific, most checkable claim first, because if it breaks, you have learned something about every other sentence in the same paragraph.
Worked Example: Auditing One Synthesized Profile
Consider a real-shaped case. The AI summary read: "Priyanka Rao, senior ML engineer, 9 years experience, led recommendation systems at a Fortune 500 retailer, Kaggle Grandmaster, author of three NeurIPS papers." Read as prose, this is an excellent candidate and an easy note to send. Read as a list of claims, it is five separate assertions, every one of them specific enough to check and none of them yet checked. Reuben audited it line by line.
| Claim in the AI summary | What the source showed | Shape |
|---|---|---|
| 9 years experience | LinkedIn showed 6 years total, 2 of them in ML specifically | Inflated number |
| Led recommendation systems | Her own description said she "worked on" recommendations, not "led" them | Upgraded role |
| Kaggle Grandmaster | Kaggle showed Expert, not Grandmaster | Inflated credential |
| Author of three NeurIPS papers | A scholar search found one workshop paper, not three at the main NeurIPS conference | Invented attribute |
| Senior ML engineer at a Fortune 500 retailer | Confirmed by her own profile | Verified |
Four of five load-bearing claims were inflated or invented. Notice how the errors compound rather than cancel. The years claim inflates seniority, the verb inflates scope, the tier inflates external standing, and the publication count inflates research depth, and all four push in the same direction. A summary with random errors would have understated something somewhere. This one did not, because the model was not making mistakes about facts; it was producing the most coherent senior-ML-engineer narrative the fragments would support, and coherence in this domain means everything points up.
The candidate was still worth a conversation, and that is the part recruiters most often get wrong when an audit fails. Six years of experience with two in ML and genuine work on recommendation systems at a large retailer is a real profile with a real story, and an Expert tier on a public competition platform is a genuine signal. What the audit destroyed was not the candidate. It was the fictional version of the candidate that Reuben was about to introduce.
The profile Reuben almost sent would have set a false expectation that the interview would have embarrassingly corrected, wasting a hiring manager's afternoon and unfairly framing a real person against a fictional resume. Trace the second-order damage: the hiring manager builds a loop around leading recommendation systems, the questions assume decisions the candidate never made, the candidate underperforms against a standard nobody told her she was being held to, and the debrief records that a sourced candidate came in weaker than advertised. The candidate carries the cost of an error she had no part in and no opportunity to correct. The lesson is not that AI sourcing is useless; it is that synthesized claims are a hypothesis, and the source is the test.
Why This Is a Fairness Issue, Not Just an Accuracy One
Hallucinations do not distribute randomly. A model invents richer narratives for candidates whose public footprints are larger, which skews toward people with the time, platforms, and networks to build visible portfolios. Conference talks require travel budgets and employer support. Open-source contributions require unpaid hours. An active public profile in a technical field is easier to sustain for someone whose caregiving load is light and whose employer permits public writing. None of that is a measure of engineering ability, but all of it is a measure of how much raw material a model has to work with when it writes a summary.
The asymmetry is what makes this a fairness problem rather than a quality problem. A candidate with a large footprint gets an enriched profile: the model has real material to elaborate, and it elaborates generously. A candidate with a small footprint gets a short, cautious, accurate note, because there is nothing to elaborate from. Placed side by side in a hiring manager's inbox, the two notes look like two candidates of different calibers, when what actually differs is the volume of public text about them. The enrichment is not evidence of quality. It is evidence of visibility, and visibility is unevenly distributed for reasons that have nothing to do with the job.
If Reuben forwards inflated profiles for some candidates and thin, accurate ones for others, he is not comparing like with like, and the candidate with the quieter public record is penalized for the AI's restraint rather than rewarded for their actual work. Verifying every synthesized profile to the same evidentiary standard is therefore part of running a consistent, defensible process, the kind that holds up if a rejected candidate ever asks how the decision was made, and the kind that aligns with disparate-impact scrutiny under frameworks like the EEOC's four-fifths guideline when AI tools shape who advances.
The practical form of "same evidentiary standard" is narrow and achievable. Every profile carries the same claim list, the same three checks, and the same explicit statement of what could not be confirmed. A profile that ends up short because the candidate has little public presence is labeled as having little public presence, not as having little to offer. That single sentence of labeling is what stops a visibility gap from being read as a capability gap, and it costs one line per candidate.
Presenting What You Verified, Not What You Hoped
The last discipline is how Reuben hands a profile to a hiring manager. He labels what is confirmed and what is unconfirmed, because an honest "verified: 6 years, 2 in ML; unconfirmed: claimed leadership scope" is far more useful than a polished paragraph that implies certainty he does not have. The structure he settled on has four parts: what is confirmed and against which source, what the summary asserted that he could not confirm, what is not publicly verifiable by nature and therefore belongs in the screen, and the open questions he intends to ask. It takes a few extra lines and removes an entire category of downstream failure.
This does three things at once. It protects the hiring manager from interviewing against a fiction, because the loop gets designed around what is actually established. It protects the candidate from being measured against an inflated version of themselves they never claimed, which is the harm that matters most and the one the candidate can never see coming. And it protects Reuben, because a sourcing note that distinguishes evidence from inference is the artifact he would want to exist if anyone later asked why a candidate advanced or did not.
Over a quarter of sourcing at volume, that habit also trains the hiring managers themselves to ask for sources, which raises the standard of every conversation that follows. The change is cultural and it is durable. Once a hiring manager has been handed three notes that separate verified from unverified, the fourth note that arrives as a confident undifferentiated paragraph reads as sloppy rather than as strong. Reuben did not have to argue anyone into a verification standard. He just kept sending notes that had one until the absence of one became conspicuous.
Anti-Patterns
Forwarding the summary as the profile. This is copying the AI's paragraph into a note to a hiring manager, perhaps with the phrasing tidied, and treating the send as the end of the work. It happens because the paragraph is genuinely well written and because the sourcing queue is long; the output looks finished, so it feels finished. What goes wrong is that everything downstream inherits the invention, and the correction happens in the interview where it costs a hiring manager's afternoon and a candidate's fair hearing. The counter is mechanical rather than attitudinal: nothing reaches a hiring manager until every specific number, role, and named artifact in it has been listed and checked, or explicitly marked as unchecked.
Checking the model against the model. This is asking the AI whether it is sure, asking it to cite its sources, or re-running the summary and comparing the two versions. It happens because it is fast and feels like verification. What goes wrong is that internal consistency is not evidence; a model that inflated a role in the first pass will confidently defend the inflation in the second, and a citation it produces may point at a page that does not say what the citation claims. The counter is that verification requires opening a primary artifact yourself: the issuing body's record, the venue's own proceedings, the repository, or the candidate's own words about their own role.
Verifying only the candidates you doubt. This is running the full protocol on profiles that feel too good and skipping it on profiles that feel ordinary or on candidates who came recommended. It happens because verification takes time and doubt feels like a reasonable trigger for spending it. What goes wrong is that your intuition about which profiles deserve scrutiny is itself shaped by who resembles the people you have hired before, so the unequal application of the protocol reintroduces exactly the inconsistency the protocol exists to remove. It also means the enrichment asymmetry goes uncorrected, because the richly elaborated profiles are the ones that feel most credible and therefore get checked least. Apply the same list of checks to every synthesized profile, including the boring ones.
Practice
Each of these produces an artifact you can keep, and the first two are worth doing on live profiles rather than invented ones.
- Extract the claim list. Take one AI-synthesized candidate profile you produced this month and rewrite it as a numbered list of discrete claims. Mark each one verifiable, unverifiable, or already verified. Count how many of the verifiable claims you had actually checked before you sent it.
- Run the three checks. For that same profile, identify every specific number, every role verb, and every named artifact. Open the primary source for each one yourself. Record what the source said next to what the summary said, and note which of the three shapes each discrepancy takes.
- Test the coherence pattern. Where you found discrepancies, check whether they all push the same direction, toward a more senior and more accomplished candidate. If they do, treat the remaining unchecked claims in the same profile as suspect rather than as probably fine.
- Rewrite one note in the four-part structure. Take a sourcing note you already sent and restructure it as confirmed with source, asserted but unconfirmed, not publicly verifiable, and open questions for the screen. Compare the two versions and ask which one you would want if you were the hiring manager.
- Check your footprint asymmetry. Pull five sourcing notes from the same requisition and compare their length and richness. Where a note is short, write down whether it is short because the candidate is weaker or because the public record is thinner, and add that sentence to the note itself.
Reflection
- Which of the three shapes, inflated numbers, upgraded roles, or invented attributes, do you think has most often reached a hiring manager through your notes without being caught?
- When you last felt confident in an AI-generated profile, what specifically produced that confidence: the sources behind it, or the fluency of the writing?
- How would you know if your sourcing consistently enriches profiles for candidates with large public footprints and leaves quieter candidates with thinner notes?
- What would change in your hiring managers' behavior if every note you sent them separated what you verified from what you inferred?
- Which claims in your typical profile are not verifiable from any public source, and are you currently treating those as facts or as questions for the screen?
Glossary
- Hallucination. A confident, fluent claim produced by an AI system with no basis in fact, indistinguishable in tone from the parts of the output that are accurate.
- Synthesized profile. A candidate summary assembled by an AI from scattered public fragments rather than from a document the candidate submitted, which means no authored artifact exists to check it against.
- Inflated number. A hallucination shape in which a specific quantity, such as years of experience or team size, is stated higher than the record supports, usually as arithmetic performed on partial dates.
- Upgraded role. A hallucination shape in which participation is rounded up to leadership, typically by substituting a stronger verb while keeping the project, employer, and dates correct.
- Invented attribute. A hallucination shape in which a credential, publication, or project that does not exist is assembled from adjacent true facts about the candidate's field or colleagues.
- Load-bearing claim. An assertion the hiring decision would rest on, such as seniority, scope, or a named credential, as distinct from incidental detail.
- Independent source. A primary artifact you open yourself, such as an issuing body's record, a venue's own proceedings, or the candidate's own description of their role, as distinct from the model's paraphrase of that artifact.
- Verification protocol. A fixed sequence applied to every synthesized profile: list the claims, check each against an independent source, and demote anything unconfirmed to a question rather than a fact.
- Footprint asymmetry. The uneven distribution of public material about candidates, which causes AI summaries to be richer for more visible candidates regardless of underlying ability.
- Disparate impact. A neutral-looking practice that produces substantially different outcomes for a protected group, assessed in the United States against frameworks such as the EEOC's four-fifths guideline.
Related Lessons
- Hallucinations, Accuracy Errors, and Information Loss covers the underlying failure modes of generative systems that this lesson applies specifically to sourced candidate profiles.
- Verification Techniques: Spot-Checking Facts, Sources, and Candidates broadens the protocol here into the general verification habits that apply across every AI-assisted recruiting task.
- Research Synthesis: Building Candidate Context from Multiple Sources is the upstream skill: how to assemble candidate context well, which determines how much invention the model has to do.
- Hands-On Practice: Review AI Summaries, Identify Errors, Revise gives you a guided repetition of the audit performed in the worked example above.
- Building Confidence to Question and Override AI addresses the harder half of this discipline, which is acting on a discrepancy once you have found one.
- Documenting Decisions: Clear Records for Legal and Fairness Review covers why a sourcing note that separates evidence from inference is the artifact you want to exist later.
Closing
The habit this lesson asks for is small and it is not optional: treat a synthesized profile as a hypothesis and the source as the test. Reuben did not become a better sourcer by becoming more suspicious in general. He became a better sourcer by listing the claims, opening the primary artifacts, and refusing to let an unconfirmed sentence travel as a confirmed one. That is three steps applied uniformly, not a judgment call made under time pressure, which is exactly why it survives a busy week.
What makes it worth the minutes is that the cost of skipping it lands on the person with the least power to correct it. A hiring manager who interviews against a fiction loses an afternoon. A candidate who is measured against an inflated version of herself loses the role, never learns why, and never had the chance to disown a claim she did not make. Verification is usually presented as a quality control practice, and it is one. It is also the mechanism that keeps a sourced candidate being evaluated on what they actually did.
Key Takeaways
- Synthesized profiles hallucinate because inputs are sparse and coherence is rewarded. With no application on file, the AI fills gaps with fluent invention and nothing flags it, so every passively sourced profile is a draft to be verified, not a fact to be forwarded. The gaps in the record are where the most confident sentences appear.
- Learn the three shapes: inflated numbers, upgraded roles, invented attributes. "8 years" over a real 5, "led" over "contributed to," and certifications or papers that do not exist are the recurring failure modes, with specific quantities the most fabricated of all. Each shape has a characteristic source that settles it.
- Run a protocol, not a vibe. List the specific claims, check each against an independent public source you open yourself, and demote anything unconfirmed to a question for the candidate rather than a fact for the hiring manager. Asking the model whether it is sure verifies nothing.
- One source check can collapse a whole inflated narrative. Reuben's Grandmaster claim died the moment he read the actual Kaggle profile, which is why opening the source beats trusting the summary every time. Hallucinated profiles fail as coherent stories at a single load-bearing point, and every error points the same direction.
- Hallucinations are a fairness problem, not only an accuracy one. Richer fabrications attach to larger public footprints, so verifying every profile to the same standard keeps the comparison consistent and the process defensible under disparate-impact scrutiny. Label a thin record as a thin record, so visibility is not read as capability.
- Present what you verified, not what you hoped. Separate confirmed with source, asserted but unconfirmed, not publicly verifiable, and open questions. It protects the hiring manager from a fiction, the candidate from an inflated double, and you from having no defensible record of why anyone advanced.
Frequently Asked Questions
Can I just ask the AI to cite its sources and check those? Citations help you locate an artifact, but they are not the verification. A model can produce a plausible reference to a page that does not say what the citation claims, and asking the model to confirm its own output tests internal consistency rather than truth. Use any citation the tool offers as a starting address, then open the primary artifact yourself and read what it actually says. If the citation points at something you cannot open, treat the claim as unconfirmed rather than as sourced.
I source at volume. Is there a shortcut when I have forty profiles and an hour? Check the invented attributes first, because a named credential, publication, or repository either exists at a findable address or it does not, and that lookup is fast. If one fails, the numbers and verbs in the same paragraph came out of the same process and should be treated as suspect. If it holds, spend the remaining time on the role verb, since seniority is usually what the requisition is screening for. What you cannot shortcut is applying the same checks to every profile, because uneven verification is where inconsistency reenters.
What do I do with claims that no public source can settle? Mark them as not publicly verifiable and route them to the screen. Why someone left a role, how they handled a difficult stakeholder, and what they would do differently are interview questions by nature, and hunting for a source for them wastes time you should spend on the claims that do have sources. The failure is not that these claims are unverified; the failure is presenting them to a hiring manager as though they were established.
Does labeling claims as unconfirmed make my candidates look weak? It makes your notes look reliable, which is a different and more valuable thing. A hiring manager who receives "verified: 6 years, 2 in ML; unconfirmed: claimed leadership scope" can build a screen that resolves the open question in twenty minutes. A hiring manager who receives a confident paragraph builds a screen against a fiction and discovers the gap in the room. Over a quarter of notes, the labeling habit trains hiring managers to ask for sources, which raises the standard of every conversation.
Why is this framed as a fairness issue rather than just sloppy work? Because the errors are not evenly distributed. Models elaborate more richly on candidates with larger public footprints, and public footprints track time, platform access, and employer support rather than ability. Side by side, an enriched profile and a thin one read as two different calibers of candidate when the only difference is how much public text exists about each. Applying the same evidentiary standard to every profile is what keeps the comparison honest, and it is the kind of consistency that matters under disparate-impact scrutiny when AI tools shape who advances.
Skill.re