Hands-On Project: Design a Sourcing Workflow with AI and Guardrails
Talia leads sourcing for a 450-person logistics company, and AI made her faster in a way that started to scare her. Where she once spent an afternoon building a single Boolean string and combing LinkedIn Recruiter by hand, she could now have an AI generate a search, pull 80 profiles, enrich them, and hand her a ranked list before lunch. The speed was real. The problem was that she had stopped being able to see inside the ranking. One week she noticed the AI's top picks for a fleet-operations role skewed oddly, and when she dug in she found it had latched onto a single keyword that happened to appear on resumes from one particular former employer, and the ranking was really just sorting by that. She had been about to message the top of that list. This project builds the workflow Talia wished she had: an end-to-end sourcing pipeline that keeps AI's speed but inserts verification checkpoints and bias guardrails so a recruiter, not an unexamined algorithm, decides who gets contacted.
What You Are Building
This is a hands-on project, so it ends in artifacts rather than in agreement. By the time you finish you will hold four things for one real open requisition: a written stage map naming what AI does and what a human owns at each of the four stages, a set of ranking criteria stated explicitly enough that someone else could apply them, a verification checkpoint with a defined scope and a defined pass or fail test, and a short guardrail statement covering what the ranking may never use. Produce those four for a single role and you have a sourcing workflow you can run again, hand to a colleague, and defend to anyone who asks how a shortlist was assembled.
The reason to build the artifacts rather than hold the process in your head is that an undocumented workflow drifts. Criteria get quietly loosened when a search is not producing enough candidates. A checkpoint gets skipped on the week a requisition goes urgent. A guardrail that lives only in a recruiter's judgment leaves with that recruiter. Talia's near miss on the fleet-operations role was not a failure of care; it was a failure of a process that had no place where care was scheduled to happen.
The Pipeline, End to End
Talia maps her workflow as four stages, each with a clear AI role and a clear human role. The discipline is that AI proposes and a human disposes; nothing leaves a stage without a person able to override it. The mapping is worth writing down in exactly this form, because the failure mode of AI sourcing is not that any single stage is wrong, it is that the stages chain together and a small distortion at the front arrives at the end looking like a conclusion.
| Stage | What AI does | What the human owns | What fails if unsupervised |
|---|---|---|---|
| Search | Translates a role description into Boolean strings, surfaces adjacent titles and skills | Deciding which adjacent titles are genuinely adjacent and which are noise | A wide net pulls in volume the later stages must filter, and volume looks like success |
| Enrichment | Assembles a structured read on each profile: current role, relevant skills, evidence of the specific experience the job needs | Treating the structured read as a claim to be checked, not a fact | Gaps get filled with plausible guesses, and the workflow quietly invents qualifications |
| Fit ranking | Scores and orders candidates against the role's must-haves | Writing and auditing the criteria the score is computed from | The order looks like a judgment when it may be sorting on a single keyword |
| Verification and outreach | Nothing, until a human has passed the shortlist | Checking claims against primary sources and writing the outreach | Messages go out on the strength of an AI score alone |
Search. AI helps construct and broaden the query, translating a role description into Boolean strings and surfacing adjacent titles and skills Talia might not have thought of. It casts a wider net than she would by hand, which is genuinely useful, but a wide net pulls in noise that later stages must filter. The thing to watch here is that breadth is not quality: eighty profiles is not eighty candidates.
Enrichment. For each surfaced profile, the workflow assembles a structured read: current role, relevant skills, evidence of the specific experience the job needs. This is where hallucination risk lives, because an enrichment step that fills gaps with plausible guesses will quietly invent qualifications. A summary reads with the same confident register whether every line traces to the source or one line was inferred, and the recruiter downstream has no visual cue distinguishing the two. Talia's rule is that an enriched profile is a hypothesis about a person, never a record of one, and the workflow has to keep the primary source reachable so the hypothesis can be tested.
Fit ranking. AI scores and orders candidates against the role's must-haves. Ranking is a recommendation, never a decision, and the score is only as good as the criteria behind it. An ordered list is unusually persuasive because it arrives already sorted, which does most of the cognitive work a recruiter would otherwise do and therefore invites the recruiter to skip it. Talia's counter is to keep the criteria in written form alongside the ranking, so a surprising order has something to be checked against.
Verification and outreach. A recruiter reviews the ranked shortlist, checks claims against primary sources, and only then writes outreach. No message goes out on the strength of an AI score alone. This is the stage the whole design exists to protect, because outreach is the first irreversible act in the workflow: everything before it happens inside Talia's tooling, where a mistake costs only her attention, while a message that reaches a candidate has turned a bad shortlist into a real interaction with a real person.
Verification Checkpoints, Not Blind Trust
The workflow's backbone is a set of checkpoints where a human deliberately interrupts the automation. The most important sits between ranking and outreach. Before any candidate is contacted, Talia verifies the top of the list against the actual LinkedIn profile or portfolio rather than the AI's enriched summary, because the summary is where invented detail hides. She is checking three things: does the claimed experience actually appear in the source, does the candidate plausibly fit the role rather than just matching keywords, and did the AI drop anyone strong for a reason that does not hold up.
That third check is the one recruiters forget, and it is the one that protects the pool rather than the shortlist. The visible ranking failure is a weak candidate pushed to the top, which the first two checks catch; the invisible one is a strong candidate pushed down, which nobody sees unless they deliberately look, because the workflow never shows you what it decided not to show you. Talia reads a slice below her verification cut as well as above it, not to verify those candidates in full, but to ask whether the reason any of them ranked low is one she would defend if a person had given it.
She does not verify all 80 surfaced candidates, which would erase the time savings, and she does not trust the ranking blindly, which would erase the point of having a recruiter. She verifies the top slice the workflow is actually going to act on. The checkpoint is calibrated to the decision: the candidates about to receive outreach get human eyes, the long tail does not, and that balance is what makes the pipeline both fast and defensible. Calibrating a checkpoint to the decision rather than to the volume is the transferable idea here, and it generalizes past sourcing. You spend review effort in proportion to what an error would cost, and at the outreach boundary an error costs a real candidate's time and your credibility.
Worked Example: 80 Surfaced, 25 Verified, 8 False Positives Caught
Talia runs the workflow for a fleet-operations supervisor role. The numbers here are illustrative, but the shape is exactly what she sees in practice. The AI search surfaces 80 candidates from LinkedIn Recruiter. Enrichment builds a structured read on each, and the fit-ranking step orders them against her must-haves: multi-site fleet experience, safety-compliance background, and team-lead responsibility.
Rather than message the top of the list, Talia opens her verification checkpoint and works through the top 25 by hand, checking each ranked profile against its primary source. Eight of the 25 turn out to be false positives. When she traces why, the pattern is clear: the AI had over-weighted a single keyword, the phrase "fleet management," and several of the eight had it on their profile because they had used a software product by that name, not because they had ever managed a fleet. The ranking had confused a tool name for a qualification and pushed those eight up the list.
The mechanism is not exotic and it will recur under different keywords in any role you source for. The ranking had no concept of what "fleet management" refers to; it had a string that appeared frequently on profiles of people who did the job and applied it to profiles where the same string was doing entirely different work in the sentence. A product name in a skills list and a responsibility described in a job history look identical to a system matching text, and the distinguishing information, that one is a tool the person used and the other is something they were accountable for, lives in context the ranking did not weigh.
Talia removes the eight, and because she now understands the failure, she adjusts the ranking criteria to require corroborating evidence of actual fleet responsibility rather than the bare keyword, then re-ranks. That is the move most easily skipped under deadline. Removing eight names fixes one shortlist; rewriting the criterion so the ranking demands a second, independent signal fixes every shortlist after it, and it is cheap while the diagnosis is fresh and expensive to reconstruct a month later. Only then does she draft outreach to the verified candidates.
The eight false positives are the whole point of the exercise. Without the checkpoint, Talia would have sent personalized messages to eight people who were never a fit, wasting her time, cluttering their inboxes, and teaching the hiring manager to distrust her shortlist. The checkpoint converted a silent ranking error into a fixable, understood problem, and it did so on 25 reviews rather than 80. Note what the workflow bought at each price point. Reviewing nothing costs nothing and ships eight bad messages. Reviewing everything costs the entire time saving. Reviewing the acting slice caught the error at a fraction of the full cost, which is the trade the checkpoint exists to make.
Bias Guardrails: What the AI Must Never Do
Speed is not the only risk in AI sourcing; bias is the one that creates legal and ethical exposure. Talia builds explicit guardrails into the workflow so the AI cannot quietly narrow her pool along protected lines. The first and firmest rule: the workflow must not filter or down-rank on proxies for protected class. Graduation years that reveal age, names or affiliations that signal ethnicity or gender, a gap that suggests caregiving or military service, zip codes that correlate with race, none of these may enter the ranking. She writes this into the prompts that drive enrichment and ranking, instructing the model to ignore such signals and to flag rather than penalize non-traditional paths.
The flag-rather-than-penalize instruction is doing more work than it appears to. A non-traditional path, a career changer, a gap, an unusual sequence of titles, is genuinely harder for a ranking to score, and the default behaviour of a system asked to score something it cannot read confidently is to score it low. That is uncertainty expressed as a number, and it produces the same outcome as hostility. Surfacing those profiles for human attention instead of quietly sinking them puts the judgment where it belongs.
She is alert to the subtler failure too: a proxy that looks like a job requirement. Filtering on "graduated within the last five years" is age discrimination wearing the costume of a freshness criterion. Requiring a specific elite employer can encode a homogeneous pipeline. Talia's guardrail is to keep ranking criteria tied to demonstrable job-relevant competencies and to periodically check, in aggregate, whether her sourced pools skew in ways the role does not justify. This is the sourcing-stage expression of the same fairness discipline that EEOC adverse-impact analysis and the four-fifths rule apply downstream: do not let a tool systematically advantage one group through criteria that are not genuinely about the work.
The practical test she applies to any proposed criterion is to ask what it is a stand-in for. "Graduated within the last five years" stands in for current familiarity with recent tooling, which can be assessed directly and which plenty of people with older degrees have. "Worked at a specific well-known employer" stands in for having operated at scale, which can also be assessed directly. Whenever the direct competency is available, using the proxy instead is both less accurate and more exposed, because the proxy carries demographic correlations the competency does not. Writing criteria as competencies is the fairness control and the quality control at once.
Privacy at the Enrichment Stage
Privacy is the companion guardrail. The profiles Talia enriches contain personal data, and under regimes like GDPR and CCPA that data has to be handled with a lawful basis, limited to what the role actually requires, and not fed wholesale into tools whose data practices she has not vetted. She keeps enrichment scoped to professionally relevant, publicly available information and avoids assembling intrusive dossiers a candidate never consented to.
The unvetted-tool point is the one that most often goes wrong quietly. Pasting candidate profiles into whatever assistant is convenient moves personal data into a system whose retention, training and access practices nobody on the team has examined, and it does so at volume without anyone making a decision they would recognize as a decision. The workflow artifact should name where candidate data is allowed to go, so using a new tool becomes a change to a written scope rather than a habit that accumulates.
Where the Tools Fit
Talia's workflow is deliberately tool-agnostic in design but concrete in practice. Sourcing and profile data come from LinkedIn Recruiter. Candidates who pass verification and receive outreach are tracked in her applicant tracking system, so the pipeline has a clean record of who was contacted and why. She is careful to keep the AI assist as a layer on top of these systems rather than a black box she cannot inspect, because a workflow she cannot audit is a workflow she cannot defend. The principle outranks any specific vendor: every automated step must expose enough of its reasoning that a recruiter can check it, and every candidate-facing action must pass through a human first.
Inspectability is worth stating as a selection criterion rather than a preference, because it determines whether the rest of this workflow is possible at all. Talia's fleet-management diagnosis depended on seeing that the ranking was over-weighting one keyword. Had the ranking been a score with no visible reasoning, she would have seen only that eight of her top candidates were wrong, removed them, and had nothing to repair, so the same eight would have arrived on the next requisition. A component you cannot inspect converts every recurring error into a permanent one.
Measuring Whether the Workflow Works
A workflow you do not measure is a workflow you cannot improve. Talia tracks three signals. The false-positive rate at her verification checkpoint tells her how far she can trust the ranking and whether her criteria are drifting: a rate that climbs means the ranking is degrading. Response rates to outreach tell her whether the verified shortlist is genuinely well-matched. And periodic aggregate checks on her sourced pools tell her whether the guardrails are holding or whether bias is creeping in through some proxy she did not anticipate.
Each signal is chosen because it fails differently, which is why three are worth the effort and one would not be. The checkpoint's false-positive rate reads the ranking's internal accuracy and is available immediately, without waiting for anyone to reply. Response rates read whether the people you approved match the role as the market understands it. The aggregate pool check reads what neither of the others can see, because a workflow can have a low false-positive rate and healthy response rates while systematically never surfacing a whole category of qualified people. A shortlist that is accurate about everyone on it tells you nothing about who never made it. Treat any moved number as an instruction to change a criterion, and write the change back into the artifact so the next person inherits the correction.
Anti-Patterns
Messaging the top of the ranked list. This is treating the order as the shortlist and going straight to outreach. It happens because the list arrives already sorted, which makes checking it feel like redundant effort, and because the speed of the pipeline creates the impression that the slow step must be the unnecessary one. What goes wrong is that a silent ranking error becomes a set of real messages to real people, as Talia's eight false positives would have been, and the damage lands on candidates and on the hiring manager's trust in every later shortlist. The counter is a checkpoint sitting structurally between ranking and outreach, so that skipping it is a visible decision rather than the default path.
Verifying against the enriched summary instead of the source. This is opening the AI's structured read of a candidate, confirming it says what the ranking claimed, and calling that verification. It happens because the summary is right there and the primary profile is another click, and because the summary is genuinely accurate most of the time. What goes wrong is that the summary is exactly where invented detail lives, so checking a claim against the artifact that may have invented it confirms nothing. The counter is Talia's rule that verification means the actual profile or portfolio, and that any enriched field you are about to rely on must trace to something a person wrote about themselves.
Removing the bad candidates and leaving the criterion alone. This is deleting the eight false positives, sending the outreach, and moving on. It happens because the shortlist is now correct and the deadline is real, and because diagnosing why the ranking failed feels like optional analysis. What goes wrong is that the criterion producing the error is still in the workflow, so the same class of false positive arrives next week, and the checkpoint becomes a permanent manual tax instead of a debugging tool. The counter is to treat a pattern in verification failures as a defect report against the criteria, fix it while the diagnosis is fresh, and re-rank.
Criteria written as proxies rather than competencies. This is a ranking rule like "graduated within the last five years" or a named elite employer, standing in for something the role actually needs. It happens because proxies are easy to state, easy to compute, and genuinely correlated with the thing you want, which makes them feel like efficient shorthand rather than a substitution. What goes wrong is that the proxy carries demographic correlations the competency does not, so a freshness filter operates as age discrimination and a brand requirement encodes a homogeneous pipeline, and the workflow narrows the pool along lines the role cannot justify. The counter is Talia's test: for every criterion, name what it is a stand-in for, and if the underlying competency can be assessed directly, assess that instead.
Build Checklist
Work these in order against one real open requisition. Each step produces a piece of the finished artifact set.
- Write the stage map before you run anything. Name the four stages, and for each one write what AI does, what you own, and what fails if nobody supervises it. If you cannot state the human role at a stage in a sentence, that stage is currently running unsupervised.
- State the must-haves as competencies, not keywords. Write the two or three things the role genuinely requires evidence of, in the language of what a person did rather than what a résumé says. For every criterion, name what it is a stand-in for, and replace it with the underlying competency wherever that can be assessed directly.
- Write the guardrail statement. List what the ranking may never use: graduation years, names and affiliations that signal ethnicity or gender, employment gaps, zip codes, and any proxy for a protected characteristic dressed as a requirement. Put it into the prompts driving enrichment and ranking, including the direction to flag non-traditional paths for review rather than penalize them.
- Scope the enrichment. Decide what an enriched profile is allowed to contain, hold it to professionally relevant and publicly available information, and name the specific tools candidate data may be sent to. Anything outside that scope is a change to the artifact, not a convenience.
- Define the checkpoint: scope, test, and outcome. Decide how deep to verify based on how far down the list you actually intend to act, not on how many profiles were surfaced. Write the three checks: does the claimed experience appear in the primary source, does the person fit the role rather than the keywords, and did the ranking drop anyone strong for a reason that does not hold up. Then write what happens on a failure, including tracing patterns back into the criteria.
- Run the search and read a slice below your cut. Verify the acting slice against primary sources, never against the enriched summary, then read a thinner slice below the line for candidates demoted for a reason you would not defend if a person had given it.
- Repair the criteria before you re-rank. Where verification failures show a pattern, rewrite the criterion that produced them, add a corroborating-evidence requirement where a single signal was carrying too much weight, and re-run the ranking. Do this before drafting outreach, while the diagnosis is fresh.
- Set the measurement loop. Record the checkpoint's false-positive rate as your baseline, watch outreach response rates on the verified shortlist, and schedule a periodic aggregate look at whether your sourced pools skew in ways the role does not justify. Track who was contacted and why in your applicant tracking system so the pipeline holds a record of the decision, not just its result.
Reflection
- On your last requisition, how far down the AI-ranked list did you contact people, and how far down did you verify?
- If a hiring manager asked why a specific candidate ranked first, could you answer with a criterion, or only with the ranking?
- Which of your current screening criteria is a proxy, and what competency is it standing in for?
- When your ranking last produced an obviously wrong candidate, did you fix the shortlist, the criterion, or both?
- What does your enrichment step currently collect that the role does not actually require evidence of?
- How would you know if your sourcing was systematically missing a group of qualified people, given that the workflow never shows you who it did not surface?
Glossary
- Enrichment. The step that assembles a structured read of each surfaced profile, and the stage where invented detail is most likely to enter the workflow, because a filled gap reads identically to a sourced fact.
- Verification checkpoint. A deliberate human interruption between ranking and outreach, with a defined scope, three defined checks, and a defined outcome when a check fails. Its scope is calibrated to the decision, meaning to what an error would cost rather than to the volume of output, and it includes reading a thinner slice below the cut for qualified people the ranking pushed down, who are otherwise invisible because the workflow never displays what it chose not to surface.
- Keyword over-weighting. The failure where a ranking treats a string as a qualification regardless of the role it plays in the sentence, as when a software product named "fleet management" was scored as fleet responsibility.
- Proxy for protected class. A criterion that correlates with a protected characteristic, such as graduation year for age or zip code for race, which may not enter the ranking whether it appears openly or disguised as a job requirement.
- Aggregate pool check. A periodic look at whether sourced pools skew in ways the role does not justify, which catches interactions between individually defensible criteria that reading the criteria cannot reveal.
Related Lessons
- AI-Assisted Sourcing: Boolean Search Optimization is the search stage in depth, covering the query construction this project treats as a single box.
- Verifying Candidate Information: Spotting Hallucinations and Inaccuracies develops the verification checkpoint, including how to tell an inferred claim from a sourced one inside a fluent summary.
- Diversity and Bias in Sourcing: How AI Can Help and Harm is the fairness case behind the guardrail statement, and the material to read before defending a criteria change to a hiring manager.
- Data Minimization: Collecting Only What's Necessary covers the enrichment scope, and why the incentive to collect more at this stage needs a written limit rather than judgment.
- Metrics and Monitoring: Tracking Efficiency, Quality, and Fairness extends the measurement loop past the three signals here and into how the numbers get reviewed and acted on.
- Root Cause Analysis: Understanding Why Bias or Errors Occurred is the discipline behind repairing the criterion rather than removing the candidate, which is the step this project asks you not to skip.
Closing
Talia's original problem was not that AI was wrong. Most of the time the ranking was reasonable, and the speed it bought her was real and worth keeping. The problem was that a stack of individually plausible automated steps had chained into an output she could not see inside, and she had no scheduled moment where seeing inside it was her job. When the fleet-operations list skewed, she caught it by noticing something felt odd, which is not a control; it is luck that fails on the week she is busy.
What replaces it is not complicated and does not slow the pipeline down much. Four stages with the human role named at each. Criteria written as competencies rather than proxies, with a guardrail statement listing what the ranking may never touch. Enrichment scoped to what the role requires and to tools you have vetted. One checkpoint between ranking and outreach, sized to the slice you will act on, checking the source rather than the summary, and reaching a little below the cut for people the ranking wrongly buried. And a habit of repairing the criterion whenever verification finds a pattern. Build those artifacts on one live requisition and you own a workflow you can run again, hand over, and defend. The AI keeps the speed; the recruiter keeps the decision.
Key Takeaways
- AI proposes, a human disposes. Structure sourcing as four stages, search, enrichment, fit ranking, and verified outreach, and let nothing leave a stage without a person able to override it. Speed is only safe when a recruiter still owns the decision.
- Put a verification checkpoint before outreach. Check the ranked shortlist against primary sources, not the AI's enriched summary, because invented detail hides in the summary. Verify the slice you will act on, not the whole pool, and read a thinner slice below the cut for anyone the ranking wrongly demoted.
- Expect ranking errors and catch them. Surfacing 80 candidates and verifying the top 25 can reveal 8 false positives from the AI over-weighting a single keyword; the checkpoint turns a silent error into a fixable one and lets you correct the criteria.
- Never filter or rank on proxies for protected class. Graduation years, names, gaps, and zip codes must stay out of the ranking, and watch for proxies disguised as requirements, like a recent-graduate filter that is really age discrimination. Instruct the workflow to flag non-traditional paths rather than penalize them.
- Tie criteria to job-relevant competence and audit in aggregate. The adverse-impact discipline that EEOC analysis and the four-fifths rule apply downstream applies at sourcing: check whether your pools skew in ways the role does not justify, because individually defensible criteria can combine into a narrowed pool that reading the criteria will never reveal.
- Mind privacy at enrichment. Under GDPR and CCPA, keep enrichment scoped to professionally relevant, publicly available data, with a lawful basis, and avoid feeding personal data into unvetted tools or building dossiers candidates never consented to.
- Measure the workflow as a loop. Track the checkpoint's false-positive rate, outreach response rates, and aggregate pool composition, then adjust and measure again. The workflow is never finished, only currently tuned.
Frequently Asked Questions
How deep should the verification checkpoint go? As deep as you intend to act, and no deeper. The sizing question is not how many profiles the search surfaced but how far down the list you will actually contact someone, because the cost the checkpoint prevents lands at the moment of outreach. Talia surfaced 80 and verified 25; the ratio is not the transferable part, the rule is. Verifying everything erases the time saving that justified the pipeline and gets abandoned within a couple of requisitions. Verifying nothing ships the errors.
My ranking has never produced an obviously wrong candidate. Do I still need the guardrails? The absence of obvious errors is a statement about what you can see, not about what the workflow is doing. Keyword over-weighting of the kind Talia found was invisible until she checked profiles against sources, and the more consequential failure, a qualified candidate the ranking pushed below the cut, is invisible by construction, because nothing in the pipeline shows you who it declined to surface. That asymmetry is why the guardrails and the aggregate pool check are not contingent on having seen a problem. A workflow that is accurate about everyone on the shortlist can still be systematically wrong about who reaches it.
Can I let the AI filter out candidates automatically to save the review time? That is a different workflow with a different risk profile, and this project deliberately does not go there. The design keeps ranking as a recommendation that orders candidates for a recruiter to act on, rather than a filter that removes them before anyone looks. The difference matters because an ordering is reversible and inspectable, so a wrongly demoted candidate can still be recovered by reading below the cut, while an automated removal is silent and leaves nothing to review. It also matters for defensibility: a human deciding on documented criteria is a decision you can explain, and an unexamined automated exclusion is one you cannot.
What do I do when a criterion I need is also a proxy? Name what it is a stand-in for, then assess that thing directly. "Graduated within the last five years" is usually standing in for current familiarity with recent tooling, which plenty of people with older degrees have and which you can look for as evidence rather than infer from a date. A named elite employer is usually standing in for having operated at scale, which likewise shows up in what someone actually did. Where the underlying competency genuinely cannot be assessed directly, the proxy is not a shortcut you may take anyway; it is a signal that the requirement needs rewriting with the hiring manager.
Skill.re