←
AI for Recruiters
Capable · M17 · lesson 17 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Privacy Boundaries: Data Sharing, Tool Selection, and Compliance

15 min

David is the recruiting operations lead at a 450-person manufacturing company with a team of six recruiters. He is the person who gets the Slack message that starts, "I found this amazing AI tool, can I use it?" Last month one recruiter pasted a candidate's full resume, including name, phone number, and current employer, into a free chatbot to "clean up the formatting." Nothing visibly went wrong, which is precisely why David found it alarming: the team had no boundary for what data could go where, and no process for deciding whether a tool was safe to adopt. This lesson is the framework he built, and it has three parts: clear rules for what candidate data can enter which tools, a way to tell consumer tools from enterprise ones, and a scorecard that produces a defensible go or no-go before any tool touches real candidate information.

Why Candidate Privacy Is Not Optional

Candidate privacy is a legal requirement and an ethical obligation at the same time, and the two rarely come apart in practice. When you use AI in recruiting you are handling genuinely sensitive material: names, employment history, contact information, salary details, and sometimes interview recordings or transcripts that capture a person's voice and offhand remarks. None of that belongs to you. It was handed to you for one purpose, evaluating someone for a role, and every use beyond that purpose needs a reason.

What makes AI tools a distinct problem is that pasting data into one is not like saving a file. The data leaves your systems, lands in a vendor's, and may be retained, inspected, or used to improve a model depending on terms almost nobody reads. The recruiter who pasted the resume was not being careless in any way she would have recognized; she was doing something that felt exactly like using a text editor. The boundary has to be explicit precisely because the risk is invisible at the moment of the action.

What Candidate Data Can Go Where

David's first move was to stop treating "AI tools" as one category. The right question is never "can we use AI?" but "can this specific data go into this specific tool?" He split candidate information into two tiers. Identifying and sensitive data is the data that, if exposed, can harm a real person. Non-identifying data, stripped of anything that points back to an individual, carries far less risk.

TierExamplesBoundary
Identifying and sensitive (be careful)Candidate names; contact information such as email and phone; specific locations or home addresses; family situation or personal circumstances; salary history; medical information; demographic information; interview recordings and transcripts, which routinely contain personal detail nobody screened forApproved enterprise tools only, contractually bound to protect it
Non-identifying (generally lower risk)General background such as "5 years backend engineering"; technologies and skills with no identifiers attached; job titles and companies; years of experience; education without specific names or datesA wider set of tools, and even then only when the task actually requires it

The general rule underneath the table is short enough to remember under pressure: if sharing the information could identify the candidate or expose personal details, be careful. The practical habit that follows is to de-identify before sharing rather than after. Replace names with "Candidate A," drop the phone number and the address, and keep only what the task needs. A formatting cleanup never needs a phone number, and once you notice that, most of the risky sharing in a recruiting team simply stops happening.

Treating public profiles as automatically safe is its own trap, and it is the one most people fall into by reasoning rather than by accident. The argument sounds solid: the candidate published this on a professional network, so it is already public, so sharing it costs them nothing. What that misses is that aggregating and processing someone's profile data through a third-party tool is a different act from that person choosing to post it. They consented to publication, not to being fed into a system that retains and analyzes them. Public data goes through the same boundary as everything else.

The Consumer-Versus-Enterprise Distinction

The single most important thing David taught his team is that the same brand of AI can behave completely differently depending on the tier you are on. The difference comes down to two questions: does the provider train its models on your inputs, and how long does it retain them?

On many free, consumer tiers, the provider may retain what you type and use it to improve its models, which means a candidate's resume could end up influencing a system used by anyone. Enterprise and business tiers typically reverse this: the provider contractually commits not to train on your data, retains it only briefly or under your control, and offers the legal paperwork to prove it. The same company can offer both, under the same product name, with the difference buried in terms that a recruiter signing up with a work email will never see. That is why the logo tells you nothing. The tier and the contract decide whether candidate data is safe, and David's rule follows directly: identifying candidate data goes only into a tier backed by an enterprise agreement that forbids training on inputs.

The practical question is how a recruiter tells which side of that line they are on, and the honest answer is that they usually cannot from the interface. A consumer account and an enterprise seat can look identical while sitting under entirely different terms. So the check is administrative rather than visual: who signed up for this, under what agreement, and is that agreement on file? A tool a recruiter registered for personally is a consumer tier until proven otherwise, regardless of what the product is called or who else in the industry uses it. This is why the approved tool list in David's policy names the account, not just the product; "we use that tool" is not an answer to the question of whether candidate data may enter it.

Evaluating a Tool Before You Trust It

Before any tool handles candidate data, David runs a five-step evaluation. The steps are ordered, because the later ones are pointless if the earlier ones fail.

Step 1: check data handling practices. Ask the vendor four questions and insist on answers in writing. Where does my data go? How long is it retained? Who can access it? Is it used to train the model? Vague answers are themselves an answer. Five specific red flags should stop the conversation: data is used to improve the model, which means your candidate data trains their AI; data is shared with third parties; there is no encryption in transit; data is retained indefinitely; and there is no option to request deletion. Any one of those disqualifies the tool for identifying candidate data.

Step 2: check compliance certifications. Look for GDPR compliance where it applies to your candidate pool, SOC 2 certification, the availability of a data processing agreement, and a privacy policy clear enough that you can actually determine what the vendor does. A privacy policy that cannot be understood is not evidence of good practice.

Step 3: evaluate the tool's own data minimization. Does the tool ask for more data than it needs? A resume-parsing tool that requires a name, email address, and postal address when it only needs resume content is telling you something about how the vendor thinks. Over-collection at the tool level is a signal about everything else you cannot see.

Step 4: assess your own company's policy. What does your organization already say about sharing candidate data with third parties, and would this use comply? Teams routinely adopt tools that violate a policy nobody has read since onboarding.

Step 5: get compliance approval. For sensitive use cases, have legal or compliance review the tool before you use it at scale. The word "scale" matters: a single de-identified test is a different risk from routing every applicant through the tool, and the review should happen before the second one starts.

The ordering earns its keep because the steps get progressively more expensive. Step 1 is four questions in an email and disqualifies most weak tools within a day. Step 5 consumes a lawyer's time and a week of calendar, so you only spend it on tools that have already survived everything cheaper. Teams that run this backwards, escalating to legal first, learn to associate the process with delay and then start skipping it, which is how the vetting collapses. Run it cheap-to-expensive and most evaluations end at step one or two with a fast, honest no.

Vendor Due Diligence: The Paperwork That Matters

Three artifacts carry most of the weight in that evaluation, and they are worth understanding rather than just collecting. The first is the data processing agreement. Under the GDPR, when a vendor processes personal data on the company's behalf, Article 28 requires a data processing agreement, a contract that specifies what the vendor may do with the data, requires it to delete or return the data on request, and binds it to the company's instructions. No DPA means no identifying candidate data, full stop.

The second is independent security attestation, most commonly a SOC 2 Type II report, which shows an outside auditor has tested the vendor's security controls over a period of time rather than at a single moment. The distinction matters: a point-in-time attestation says the controls existed on a particular day, while a Type II report says they operated. Confirm the report covers the actual product you intend to use, not a different service from the same parent company.

The third is data residency: where, physically, does the data live? For EU candidates this matters because transferring personal data outside the EU requires a lawful transfer mechanism, such as the EU Standard Contractual Clauses. David asks every vendor where data is stored and processed, and whether he can choose an EU region if needed. Alongside these three, he confirms the basics: encryption in transit and at rest, a documented retention and deletion policy, and whether data is shared with any third parties at all.

Worked Example: David's Tool-Selection Scorecard

David scores every candidate tool on six factors, each rated 0, 1, or 2, where 2 is strong, 1 is partial, and 0 is a gap. The factors are training on inputs (2 means the provider never trains on your data), retention and deletion control, data processing agreement available, SOC 2 Type II, data residency options, and encryption. A tool needs a total of at least 9 of 12, with no zeros on the DPA or training factors, to get a go for identifying candidate data.

He recently scored two tools for resume parsing. Tool Alpha is an enterprise-tier AI platform. Tool Beta is a free consumer chatbot one of the recruiters loved.

FactorTool AlphaTool Beta
Training on inputs2, contractually no training on inputs0, consumer tier used inputs to improve the model
Retention and deletion control20
Data processing agreement20, none offered to individual users
SOC 2 Type II21, parent company certified but the certification did not cover the consumer product
Data residency options1, US regions only, EU on a higher plan0
Encryption21
Total11 of 12: go2 of 12: no-go

Tool Alpha cleared the threshold with no zeros on the critical factors, so the verdict was go, with the standing note that EU candidate data must wait until the EU region is enabled. That caveat is the useful part of the exercise: the scorecard does not just produce a verdict, it produces a condition the team can check later. Tool Beta's two zeros on critical factors made it an immediate no-go for any identifying candidate data.

David did not ban Tool Beta outright, and that decision is worth examining, because outright bans are the reason shadow tooling exists. He allowed it for fully de-identified, non-candidate tasks such as drafting a generic outreach template, and required that no candidate personal information ever enter it. The recruiter who liked the tool kept a version of what she wanted, the boundary held, and nobody had to pretend a useful tool was useless. What the scorecard really bought him was that a "trust me, it's fine" argument became a documented, defensible decision the team could point to later.

Five Practices for Minimizing What You Share

The scorecard governs which tools you may use. These five habits govern how much you hand them once approved, and they apply even to tools that scored well, because a strong contract is not a reason to share more than the task requires.

Practice 1: de-identify where possible. Strip the identifying detail out of the content itself before it goes anywhere, rather than relying on the vendor to ignore it. This is the habit that does the most work, because it changes the risk profile of the sharing rather than merely documenting it. It is also the one people assume will cost them output quality, which is why it is worth testing on a real profile rather than arguing about in the abstract.

Practice 2: use pseudonyms. "Candidate A, Candidate B" instead of names costs you nothing in a comparison or summarization task and removes the direct identifier entirely. The pseudonym also does something useful inside your own team: it makes it slightly harder for an evaluator to attach a name, and the demographic inferences a name can carry, to a judgment they are about to make.

Practice 3: limit to what is needed. Do not share a personal phone number if an email works, a full address if a city works, or an interview recording if a transcript works. The recording example is the sharpest of the three, because a recording carries a person's voice, accent, background environment, and every unguarded aside, none of which the summarization task needs. Each substitution reduces exposure without reducing the usefulness of the output.

Practice 4: document what you are sharing. Keep a record at the level of "for sourcing research, we share resume content only, no names or contact info." The value is that the answer exists before anyone asks the question, whether the asker is a candidate, an auditor, or a colleague deciding whether a new use is in bounds.

Practice 5: have a deletion policy. When a candidate is hired or rejected, decide in advance what happens to their data inside the AI tool. Make it an actual step owned by a named person, because deletion that depends on someone remembering to do it is not a policy.

What the de-identification habit looks like in practice is worth seeing rather than describing. A recruiter wants an AI tool to summarize a candidate's fit against a job description. The naive version pastes the resume whole: name, phone number, home address, current employer, and dates. The de-identified version passes the general background, "5 years backend engineering," the technologies and skills with no identifiers attached, the job titles and companies, the years of experience, and education without specific names or dates. The task the tool is being asked to perform is identical, because none of the removed fields were inputs to the judgment. That is the test to apply to your own sharing: for the task at hand, which of these fields is the tool actually reasoning over? The fields that fail that question are exposure without benefit, and they are usually the majority of what gets pasted.

Making the Boundaries Stick

A boundary that lives only in David's head protects nobody. He wrote the rules into a short data-sharing policy covering the two data tiers, the approved tool list with each tool's verdict, the de-identification habit, and a one-line escalation path for "can I use this new tool?" requests. New tools route through the scorecard before approval, and the approved list names which tool is cleared for which kind of data, because "approved" without a scope is how a tool cleared for de-identified drafting ends up processing resumes.

He also added a deletion step to the candidate lifecycle. When a candidate is hired or rejected, the team confirms their data is removed from any AI tool that touched it, relying on the deletion clause in each vendor's DPA. Once a quarter, David audits actual usage against the policy: which tools are people really using, and is any identifying data flowing somewhere it should not? That audit is the part most teams skip and the part that makes the rest true. The boundaries hold because they are written down, tied to a clear approval process, and checked, not because everyone remembers them.

Anti-Patterns

Sharing unnecessary data. The recruiter pastes a candidate's name, email, and phone number into an AI tool for a task that needed none of them. It happens because copying the whole document is faster than editing it down, and because nothing visibly bad occurs when you do. What goes wrong is unnecessary exposure of personal information that the candidate never agreed to, sitting in a system you do not control. The fix is the de-identification habit: strip identifiers first and share only what the task requires, which usually takes seconds.

Not vetting tools for compliance. The team adopts a popular AI tool on its reputation without checking how it handles data. It happens because popularity reads as safety, and because the vetting process feels like bureaucracy standing between a recruiter and a useful tool. What goes wrong is that you may be exposing candidate data or violating regulations without any awareness of it, and you will find out from an auditor or an incident rather than from the vendor. The fix is to check data handling before using the tool at scale, which is exactly what the five-step evaluation and the scorecard exist to make quick.

Assuming public information is safe to share. The team reasons that a professional profile or a public code repository is already public, so routing it through a third-party AI tool is harmless. It happens because the logic is genuinely appealing and the distinction is subtle. What goes wrong is that sharing data via a third-party tool is a different act from a candidate choosing to publish it: you are aggregating and processing their information, often combining sources they never connected themselves. The fix is to run even public information through the same privacy framework you apply to everything else.

Practice

  • Audit your current tool usage. For each AI tool you use, write down what candidate data you are actually sharing with it, whether that data is necessary for the task, and whether you could de-identify it instead.
  • Evaluate a new tool end to end. Take a tool you are considering and run the five steps: data handling practices, compliance certifications, the tool's own data minimization, your company policy, and whether it needs compliance review. Would you recommend it, and on what evidence?
  • Write your data-sharing policy. For your team, define what candidate data can be shared with AI tools, when, and with which tools specifically. One page is enough if it names the tiers and the approved list.
  • Test de-identification. Take a real candidate profile and rewrite it sharing only job-relevant information with no identifying details. How much genuinely useful information do you lose? For most tasks the honest answer is almost none, which is the point of the exercise.
  • Build your tool evaluation checklist. Turn the factors in this lesson into a checklist covering data handling, compliance, retention, deletion, and encryption, then use it before approving anything.

Reflection

  • What is the most sensitive candidate data you currently share with an AI tool, and does the task genuinely require it?
  • If a candidate asked you exactly what data about them you had shared, and with which tools, what would you be able to say?
  • What would change about how you use AI if you treated privacy as a design constraint rather than a review step at the end?

Glossary

  • Data minimization. Sharing only the data necessary for the task, and nothing more.
  • De-identification. Removing information that could identify a specific individual.
  • GDPR. General Data Protection Regulation, the EU privacy law, which applies broadly.
  • SOC 2. A security and organizational controls certification. A Type II report covers how controls operated over a period rather than at a single moment.
  • Data processing agreement (DPA). A contract specifying how a vendor handles your data. Required under GDPR Article 28 when a vendor processes personal data on your behalf.
  • Standard Contractual Clauses. A lawful mechanism for transferring EU personal data outside the EU.
  • Data residency. Where data is physically stored and processed.

Closing

Candidate trust depends on protecting candidate privacy, and responsible privacy practice is a competitive advantage rather than a tax on speed. Privacy is not a checkbox you clear once; it is an ongoing practice, which means regularly asking what data you are sharing, whether it is necessary, and whether it is protected. The concrete version of that is simple enough to start this week: audit the AI tools you use, document what data each one receives, find the sharing that is not necessary, and change it. David's team did not become slower once the boundaries existed. They became able to answer a question they previously could not, which is where the confidence to adopt new tools actually comes from.

Key Takeaways

  • Ask which data goes into which tool, not whether to use AI. Split candidate information into identifying or sensitive versus non-identifying. Be careful with names, contact information, locations, personal circumstances, salary history, medical and demographic data, and interview recordings. De-identify before sharing, and run even public profile data through the same framework.
  • The tier and contract decide safety, not the brand. Consumer tiers may train on your inputs and retain them; enterprise tiers typically contractually forbid training and limit retention. Identifying candidate data goes only into a tier backed by an enterprise agreement.
  • Evaluate tools in five steps. Data handling practices, compliance certifications, the tool's own data minimization, your company policy, then compliance approval before use at scale. Five red flags disqualify a tool outright: training on your data, third-party sharing, no encryption in transit, indefinite retention, and no deletion option.
  • Require a data processing agreement before sharing personal data. Under GDPR Article 28 a processor must operate under a DPA that controls use and guarantees deletion on request. No DPA means no identifying candidate data, with no exceptions.
  • Verify security and data residency independently. Look for a SOC 2 Type II report covering the actual product, confirm encryption in transit and at rest, and ask where data physically lives, because moving EU personal data abroad needs a lawful transfer mechanism such as the Standard Contractual Clauses.
  • Score tools before you trust them. A rubric covering training, retention, DPA, SOC 2, residency, and encryption, with a minimum threshold and no zeros on the critical factors, turns "trust me" into a documented go or no-go with a recorded condition to recheck later.
  • Allow weak tools only for de-identified, non-candidate work. A tool that fails the scorecard for candidate data can still help with generic drafting, as long as a hard rule keeps all candidate personal information out of it. Outright bans push usage into the shadows.
  • Minimize what you share even with approved tools. Pseudonymize, substitute the lighter-touch artifact where it works, document what you share, and make deletion an actual step in the candidate lifecycle.
  • Write the boundaries down and audit them. A short data-sharing policy, an approved tool list scoped by data type, a deletion step at hire or rejection, and a quarterly usage audit are what keep the rules real instead of relying on memory.

Frequently Asked Questions

Is it really a problem to paste a resume into a chatbot? Nothing happened. Nothing visible happened, which is the difficulty. On many consumer tiers the provider may retain the input and use it to improve models, so the exposure is real even though there is no incident to point at. The candidate handed you that resume for one purpose and did not agree to it entering a third-party system. Treat the absence of a visible consequence as a poor evidence base rather than an all-clear, and apply the boundary regardless.

The candidate's profile is public. Why does sharing it need approval? Because publication and processing are different acts. The candidate chose to make information visible on a professional network; they did not choose to have it aggregated, analyzed, and retained by a vendor they have never heard of, potentially combined with other sources they kept separate. The convenience of the public-data argument is exactly what makes it the third anti-pattern in this lesson. Public data goes through the same framework.

Our vendor says they are GDPR compliant. Is that enough? It is a claim, not an artifact. Ask for the specific things that make the claim checkable: a data processing agreement under Article 28, a SOC 2 Type II report covering the product you will actually use rather than a sibling service, a written answer on whether inputs train the model, a documented retention and deletion policy, and confirmation of where data is stored. If EU candidate data is involved, ask what transfer mechanism applies. A vendor who cannot produce these quickly has answered your question.

How do we handle a recruiter who really wants a tool that fails the scorecard? Separate the tool from the data. David's answer with Tool Beta was to permit it for fully de-identified, non-candidate work such as drafting a generic outreach template, while hard-blocking any candidate personal information. That preserves most of what the recruiter wanted, keeps the boundary intact, and avoids the outright ban that pushes usage into tools nobody has evaluated at all. Record the decision and its scope on the approved list so it is not relitigated monthly.