Privacy, Data Security, and Candidate Trust Risks
Daniel Okafor leads talent acquisition for a 600-person healthcare technology company, managing a team of nine recruiters who together touch roughly 4,000 applicants a quarter. When his team started pasting resumes into a consumer chatbot to draft screening summaries and feeding interview recordings to a transcription tool, Daniel realized something uncomfortable: every one of those candidates had handed his company sensitive personal data to be evaluated for a job, and nobody had asked whether that data was safe once it left the applicant tracking system. He was not worried about a single villain. He was worried about a hundred small, ordinary decisions that, added together, could leak a candidate's medical accommodation request into a model's training set or expose 4,000 records in a breach. This lesson is the risk framework Daniel built to get ahead of that before it became a headline.
Candidate Data Is a Trust Asset You Are Holding
When a candidate applies for a role, they share work history, contact information, education, availability, compensation expectations, medical accommodations, visa status, and sometimes deeply personal details about their situation. They are making a vulnerable choice, handing over information that could affect their career, their reputation, and in some cases their safety. That trust is a liability on your books whether or not you have written it down. Introducing AI does not create the obligation; it multiplies the ways the obligation can be broken, because every new tool is another copy of the data somewhere you do not directly control.
Daniel started by naming the data, because you cannot manage a risk you have not described. There is direct data the applicant types into the form: name, email, phone, work history, education. There is inferred data the AI manufactures, such as a screening tool guessing seniority from job titles or an interview-analysis tool guessing communication style from a recording. There is behavioral data about how candidates engage, like whether they opened an email. And there is derived data assembled by stitching sources together: a LinkedIn profile cross-referenced with a GitHub history and a background check. Each category has a different failure mode, and lumping them together hides the ones that matter.
The category that kept Daniel up at night was sensitive data. A candidate who requests a sign-language interpreter for an interview has just disclosed a disability, which under the Americans with Disabilities Act is protected and must be kept separate from hiring decisions. Salary history, religious observance that surfaces in scheduling, ethnicity inferred from a name: none of it belongs in a prompt window. The moment any of this becomes input to an AI system it can be logged, stored, and sometimes reused. Candidates increasingly ask the three questions that follow directly from this: will my resume be used to train AI models, who will see my application, and how long will you keep my data.
You Cannot Outsource Accountability
The most expensive misunderstanding in this area is the belief that using a vendor moves the risk onto the vendor. It does not. You remain responsible for how candidate data is used even when a third party processes it. You own the compliance burden, the notification duty, and the trust relationship with a candidate who applied to your company and has never heard of your transcription vendor. When a vendor is breached, it is your candidates who must be told and your regulator who must be notified. Contracts allocate cost and recourse between companies; they do not hand off the obligation.
This reframes vendor selection as a risk decision rather than a procurement decision. Certifications help, but they are indicators, not assurances. A vendor holding a security attestation and a vague retention policy is still a risk you have taken on. Daniel's rule is that certifications set a floor for the conversation and never end it, because the questions that actually reduce exposure are specific and answerable in writing: what is encrypted, who can access it, how long is it kept, and what happens the day something goes wrong.
The Four Categories of Exposure
Daniel grouped his exposure into four buckets so his team could reason about them clearly. The first is data leakage into model training. When a recruiter pastes a resume into a free consumer chatbot, the default terms on many free tiers allow the provider to use that input to improve future models. A candidate's history could, in principle, surface as a fragment in another company's output. The fix is structural rather than behavioral: use enterprise tooling that contractually guarantees inputs are never used for training. A written no-training guarantee is the single highest-leverage control on this list, because it removes the risk from the recruiter's memory and puts it in a contract.
The second is a breach. The more candidate data you hold, and the more places it sits, the larger the target. A consumer account with a weak password, a transcription vendor with no encryption at rest, a spreadsheet of 4,000 applicants on someone's laptop: each is an unlocked door. The third is over-collection. Every field you do not need is a liability you chose to take on. If Daniel's team never uses a candidate's home address before an offer, collecting it early only creates risk with no benefit, and it enlarges the population of records that must be reported on if something leaks.
The fourth is candidate-trust erosion, which is slower and harder to see because it never generates a ticket. Candidates who discover after the fact that a recording was scored for personality, or that their data was kept for years, do not usually file complaints. They quietly tell their network, and the best applicants stop coming. This is the risk with no incident report and no regulatory letter, which is exactly why it goes unmanaged. Daniel tracks it deliberately, because a slow decline in his inbound pipeline is more expensive over three years than any fine he is realistically likely to receive.
Five Ways AI Specifically Widens the Attack Surface
Beyond the four categories, Daniel keeps a shorter list of risks that exist only because AI tooling is in the workflow, and that a security review built for an era when candidate data sat in one applicant tracking system tends to miss. Each is phrased as a question he can put to a vendor in writing.
Risk 1: Cloud Storage and Geographic Data Residency
When you use an AI tool, candidate data usually moves to the vendor's cloud, which raises questions most recruiters never think to ask: where exactly is the data stored, in which country, and is it encrypted both in transit and at rest. GDPR restricts transferring EU personal data outside the EU without strong safeguards, so if an AI vendor stores data on servers outside the EU you need a data processing agreement with standard contractual clauses. Failing to put that in place is a compliance violation on its own, independent of whether anything is ever leaked.
The questions Daniel puts to every vendor in writing are: where is our data stored, can we restrict it to specific regions, how is it encrypted in transit and at rest, and who controls the encryption keys. If a vendor cannot say who holds the keys, they cannot credibly say who can read your candidates' records.
Risk 2: Data Retention and Deletion
How long does the vendor keep candidate data after you are finished with it? Vague retention language is common and each variant hides a different problem. "Until you delete it" puts the entire burden on you and guarantees drift. "Indefinitely" is a standing compliance risk. "As long as your account is active" leaves what happens after you churn unanswered. The remedy is to write your own retention schedule first, then require vendors to agree to it contractually, as concretely as a commitment to delete all candidate data within a fixed number of days of your instruction.
Risk 3: Unauthorized Use of Data for Model Training
Some AI vendors use customer data to train or improve their models. In a recruiting context that means your candidates' resumes, interview feedback, and hiring decisions are feeding the training data for a commercial product sold to other companies, without those candidates ever being asked. This is a GDPR violation if you have not informed candidates and established a lawful basis, and it is an ethical problem well outside the EU too. Candidates applied to your company for a job; they did not apply to become training data. Check the vendor's terms for whether data is used for model improvement and whether you can opt out, and get the answer in writing.
Risk 4: Third-Party Access and Subprocessors
Does the vendor share data onward with other companies for analytics, compliance, or benchmarking? Every additional company that touches candidate data multiplies the exposure, and subprocessors are the part of the chain buyers most reliably forget to look at. GDPR requires you to approve subprocessors, so a vendor that quietly adds one has put you in breach through no action of your own. The protection is contractual: an agreement that limits subprocessor use, requires notice of changes, and gives you the current list rather than a promise that one exists.
Risk 5: No Transparency and No Audit Capability
Can you audit how a vendor uses your candidate data? Can you see logs of who accessed what and when? Without that, you cannot answer a candidate's access request, scope a breach, or demonstrate compliance to a regulator. A vendor who cannot tell you how many candidate records were processed in a month and whether any access was unauthorized is a transparency risk regardless of how good their security actually is, because an unverifiable control is indistinguishable from an absent one when you are the party being asked to prove it.
What the Exposure Actually Costs
Leadership engaged with this topic properly only once Daniel stopped describing it as a principle and started describing it as a number. Under GDPR, fines reach up to 20 million euros or 4 percent of annual global revenue, whichever is higher; for a company with 100 million euros of revenue, the 4 percent branch is 4 million euros. Regulators have levied privacy fines in the hundreds of millions against large technology companies, and hiring practices have drawn litigation independently of privacy law. The point is not that a mid-market recruiting team draws the largest fine on record. It is that the ceiling is set by revenue, not by the size of the recruiting function.
The cost also arrives in forms that never appear on a regulator's website. Class-action litigation is common after breaches involving large numbers of individuals, remediation consumes engineering and legal time budgeted for something else, and the reputational cost lands precisely where a recruiting function is least able to absorb it: in the willingness of strong candidates to hand you their information at all. Daniel's framing to his executive team was deliberately blunt. The recruiting funnel is one of the largest collections of personal data the company holds, and it is usually the one with the least security attention per record.
How Trust Is Lost and Rebuilt
Trust is lost in silence and rebuilt in transparency. The fastest way to erode it is to use AI invisibly, screening, scoring, or analyzing candidates without telling them. When a candidate finds out later, the harm is not only the analysis itself but the discovery that they were never given a choice. New York City Local Law 144 codifies that instinct into law: an employer using an automated employment decision tool on candidates in the city must notify them at least ten business days in advance, disclose the categories of data used, and publish the results of an independent annual bias audit. Daniel treats it not as a regional edge case but as a preview of where disclosure norms are heading.
Rebuilding trust is concrete work, not a statement of values. Tell candidates plainly, in the job posting and in the application flow, that AI assists with parts of the process, what it does, and that a human makes every advancement and rejection decision. Offer a path to request human-only review, and honor deletion requests promptly. Daniel found the disclosure did not scare candidates away; it did the opposite. A clear, confident explanation of responsible AI use signaled that his company took candidate data seriously, which is exactly the signal strong candidates use when choosing between offers.
The Practices That Actually Move Trust
Beyond disclosure, a handful of ordinary behaviors carry disproportionate weight. Keep humans in the high-stakes moments: use AI for screening and scheduling, keep people for interviews, offers, and rejections. Explain decisions rather than issuing a generic line about a stronger candidate, because a specific, honest reason shows respect and helps the person improve. Respect candidate time by not making them retype a resume and not leaving them without an update for weeks. Treat candidates consistently, since responding faster to referrals than to cold applications is read as bias whether or not it was. The return is not sentimental: candidates rejected respectfully become future applicants, customers, and referrers.
The Retention Clock Is an Exposure Clock
Every record you hold is a record you can lose, so a retention schedule is a security control as much as a compliance artifact. Deletion windows of 30 to 90 days after a decision are common for screened-out applicants, and the defensible pattern is a different clock per category rather than one blanket rule, then automated deletion, since manual deletion is easy to forget and impossible to evidence. The schedule below is what Daniel documented and wired into his systems so the exposure window closes on its own.
| Candidate status | Retention period | Reason |
|---|---|---|
| Rejected at initial screening | 90 days | Brief follow-up period, then delete |
| Rejected after interviews | 1 year | Possible rehire if the role reopens; statute of limitations for hiring disputes |
| Rejected from the final candidate pool | 1 year | Ability to defend the hiring decision if it is challenged |
| Hired and onboarded | 3 to 7 years | Employment records, tax compliance, potential disputes |
| Video interviews, candidate not hired | 90 days | Minimization; no ongoing need to retain a recording |
| Reference checks, candidate not hired | 1 year | Supporting documentation for the hiring decision |
Access control belongs in the same conversation, because retention determines how much data exists and access determines how many people can reach it. Role-based access means recruiters, hiring managers, and HR staff can see candidate data and nobody else can; an applicant tracking system open to the whole company is a finding, not a convenience. Audit logs should record who accessed which records and when, specifically enough to reconstruct an incident. Review access quarterly, and if your system cannot produce logs at all, raise that with the vendor as a requirement.
Building a Risk Register: A Worked Example
To turn anxiety into action, Daniel built a risk register. Each row names a risk, rates its likelihood and its impact on a one-to-five scale, multiplies them into a score, and assigns an owned mitigation. The numbers below are his own illustrative ratings rather than any benchmark, but the method is the transferable part: it converts a diffuse sense of unease into a ranked, ownable list on a single page.
Risk 1: resume data leaking into model training. Likelihood 4, impact 4, score 16. The mitigation is to ban free consumer chatbots for any candidate data and standardize on an enterprise tool with a written no-training guarantee. This was the highest-scoring risk on the register and the first one Daniel closed, precisely because the control is structural and does not depend on nine recruiters remembering a rule under deadline pressure.
Risk 2: a breach of stored candidate records. Likelihood 2, impact 5, score 10. The mitigation is to keep candidate data inside the applicant tracking system rather than scattered across spreadsheets, encrypt at rest, enforce single sign-on, and limit access to the recruiters who need it. Critically, Daniel documented a breach-response plan, because under GDPR a personal-data breach affecting EU candidates must be reported to the supervisory authority within 72 hours of becoming aware of it. A plan written after the clock has started is already late.
Risk 3: over-collection of sensitive fields. Likelihood 4, impact 3, score 12. The mitigation is to strip the application form down to what each stage needs and never store accommodation disclosures alongside evaluation data. Risk 4: a candidate discovering undisclosed AI scoring. Likelihood 3, impact 3, score 9. The mitigation is to publish an AI-use disclosure and route adverse action through human review. Sorting by score, 16, 12, 10 and 9, gave Daniel an unambiguous order of work for the quarter.
The 72-Hour Clock: Preparing for the Incident
GDPR gives you 72 hours from awareness to notify the supervisory authority, and 72 hours is not enough time to also decide who is in charge, find the regulator's contact details, and draft a candidate notification from a blank page. Everything that can be written in advance should be. Daniel's pre-breach package is short and lives where his team can find it at two in the morning, which is the only test of an incident plan that has ever mattered.
- Incident response team. Named roles: data protection officer if you have one, legal, security, HR, and a single decision-maker.
- Contact information. The relevant data protection authorities in your jurisdiction and in your candidates' jurisdictions, gathered before you need them.
- Notification templates. Draft candidate letters and regulator notifications in advance so the clock is spent on facts rather than on wording.
- Documentation process. A defined way to record what happened, what was decided, and what was communicated; regulators will ask.
- Recovery procedures. An approach to closing the underlying vulnerability, not only containing the incident.
With that in place, the response has a realistic shape. In the first four hours the team convenes, scopes how many candidates and which data are affected, secures the affected systems, and starts documenting. From hour four to hour twenty-four the work is investigation with IT and security to find root cause, contacting vendors if their systems were involved, and determining whether the breach meets the threshold for notification. Hours twenty-four to forty-eight are for drafting the regulator notification and consulting legal, and the notification itself goes out inside the 72-hour window rather than at the edge of it. Candidate notification follows in the first week, written as a clear and honest account of what happened, what you are doing about it, and what the candidate should do, including guidance on credit or identity monitoring where the exposed data warrants it. Remediation runs across the following weeks: fix the vulnerability, add controls, and review the remediation plans of any vendor involved.
The Mitigations That Carry the Most Weight
Across the whole register, three mitigations did the heaviest lifting. The first is enterprise tooling with a no-training guarantee and a signed data processing agreement specifying what the vendor may access, how long they may retain it, where it is stored, who their subprocessors are, whether you may audit them, and how breaches are reported. A vendor's public privacy policy protects the vendor; the agreement protects your candidates. If a vendor cannot produce one at all, that absence is the finding.
The second is data minimization, which shrinks the breach surface and the compliance burden simultaneously. Collect late, retain briefly, delete on a schedule. Fields that reliably do not belong in an early-stage application include Social Security number before an offer, date of birth, marital status or number of children, religion, ethnicity, political affiliation, and criminal or credit history unless the role genuinely requires it. Each expands your legal and bias exposure at once, for information you will not use in the decision in front of you.
The third is transparency paired with human oversight. Disclosure satisfies the direction regulation is moving in, and human review of any adverse action keeps a person accountable for a decision affecting someone's livelihood. None of these three are exotic. They separate a team that gets to use AI confidently from one that gets to explain a data incident to four thousand people who trusted it with their resumes.
Anti-Patterns
Treating the vendor contract as a transfer of risk. Teams sign with a certified vendor and conclude the exposure now belongs to somebody else. When the vendor is breached, you notify your candidates, you notify your regulator, and you absorb the reputational damage. The fix is to treat vendor selection as risk retention: assume you will own every consequence, and price the diligence accordingly.
Auditing the tools and ignoring the shadow workflow. The formally procured screening platform gets a security review while a recruiter quietly pastes resumes into a free chatbot to save twenty minutes. The reviewed surface is not the real surface. The fix is a sanctioned enterprise tool genuinely good enough for the task, because a ban with no approved alternative just moves the activity out of view.
Writing the incident plan during the incident. The 72-hour clock starts at awareness, not when you feel ready, and teams that never named an incident owner spend the first day deciding who is in charge. The fix is a short pre-breach package drafted in calm conditions.
Storing an accommodation disclosure with the evaluation record. A candidate asks for an interpreter or extra time, and the request lands in the same notes field as the interview scores. That co-location is both a privacy failure and a discrimination risk, because it puts protected information in front of the people making the decision. The fix is a separate, access-restricted channel for accommodation requests, and the same logic applies to every sensitive field that arrived in your application form because it was in the default template rather than because anyone uses it before an offer.
Practice Prompts
- Map the data. Build a one-page inventory of every candidate data element your team collects and every place it lives, including the informal places: resumes in the applicant tracking system, phone numbers duplicated into a sourcing CRM, interview notes in a shared document, recordings on a vendor's servers. The duplicates are the finding.
- Score your register. Write your own four to six top risks, rate likelihood and impact from one to five, multiply, and sort. Then check whether your current effort is going to the highest scores or to the most familiar ones.
- Interrogate one vendor. Pick the AI tool your team uses most and get written answers to: where is data stored, is it encrypted in transit and at rest, who holds the keys, is our data used for model improvement, can we opt out, who are the subprocessors, can you produce access logs. Note which were hard to answer.
- Draft the disclosure. Write the paragraph you would put in a job posting explaining that AI assists with screening, what it does, that a human makes every advancement and rejection decision, how long you keep data, and how to request human-only review. Remove anything written to protect you rather than to inform the candidate.
- Rehearse the clock. Table-top a breach of your applicant tracking system. Who convenes in hour one? Who scopes it? Who contacts the regulator, and do you have the contact details today?
Reflection
- If a candidate asked today exactly what data you hold on them and where it lives, how long would it take to answer accurately, and what would you find while looking?
- Which of your current AI tools would you be uncomfortable describing, in plain language, to the candidates whose data flows through it?
- Where in your process does a protected disclosure, such as an accommodation request, currently sit relative to the evaluation record?
- Your most senior leader asks what a privacy failure in recruiting would cost this company. What do you say, and on what evidence?
- Which risk on your list gets the least effort relative to its score, and what is the honest reason?
Glossary
- Data processing agreement. A contract governing what a vendor may do with personal data on your behalf: permitted uses, retention, storage location, security standards, subprocessors, audit rights, and breach notification.
- Subprocessor. A third party a vendor engages to help process your data. Under GDPR their use must be approved, so an undisclosed addition breaches your own obligations.
- No-training guarantee. A contractual commitment that inputs submitted to an AI tool will not be used to train or improve the provider's models.
- Data residency. The geography in which data is physically stored and processed. Moving EU personal data outside the EU requires a lawful transfer mechanism such as standard contractual clauses.
- Encryption in transit and at rest. Protection of data as it moves between systems and as it sits on storage. Both are required.
- Audit log. A record of who accessed which data and when, detailed enough to scope an incident and answer a candidate's access request. Paired with role-based access control, which restricts access by job function so only recruiters, hiring managers, and HR staff reach candidate records.
- Risk register. A ranked list of risks with likelihood, impact, a computed score, and a named owner per mitigation.
- Automated employment decision tool. The category covered by New York City Local Law 144, carrying advance-notice, data-disclosure, and annual independent bias audit obligations.
Related Lessons
- Data Privacy Fundamentals: GDPR, CCPA, FCRA, and Regional Requirements for the full legal landscape this lesson touches only where it creates exposure.
- Privacy Boundaries: Data Sharing, Tool Selection, and Compliance for the step-by-step tool evaluation and the paperwork that supports it.
- Data Minimization: Collecting Only What's Necessary for the field-level audit that shrinks the surface described here.
- Consent and Transparency: What Candidates Need to Know for what disclosure must contain and how consent is obtained.
- Privacy as a Candidate Right and Organizational Responsibility for the ethical frame underneath the risk frame.
- Compliance Risks and Legal Exposure for the wider liability picture.
Closing
Daniel did not solve privacy risk. He made it visible, ranked, owned, and small enough to work through in a quarter, which is the achievable outcome. Candidate data will keep flowing through more tools than you can personally inspect, and the honest position is not that nothing will ever go wrong but that you know where the data is, you closed the largest gaps first, and you can respond correctly when something does. The recruiters who get to keep using AI confidently are the ones who did this unglamorous work before an incident forced them to.
Key Takeaways
- Name the data before you protect it. Recruiting holds direct, inferred, behavioral, derived, and sensitive data. Sensitive categories such as disability disclosures under the ADA and salary history must never enter a prompt window, because anything sent to an AI system can be logged, stored, and sometimes reused.
- You cannot outsource accountability. You are responsible for candidate data even when a vendor processes it. If the vendor is breached, you notify candidates and regulators, so vendor selection is a risk decision, not a procurement one.
- Treat the no-training guarantee as your top control. Free consumer chatbot tiers may use inputs to improve models. Enterprise tooling that contractually excludes business inputs from training closes the highest-scoring leakage risk structurally, not by asking recruiters to remember a rule.
- Five AI-specific risks widen the surface. Cloud storage and data residency, vague vendor retention, use of your data for model training, undisclosed subprocessors, and absent audit capability are all invisible to a security review designed for one applicant tracking system.
- Know your breach clock. Under GDPR, a personal-data breach affecting EU candidates must be reported to the supervisory authority within 72 hours of awareness. Name the incident team, gather regulator contacts, and draft the notification templates before an incident.
- The exposure ceiling is set by revenue, not by team size. GDPR fines reach up to 20 million euros or 4 percent of annual global revenue, whichever is higher, and class actions and remediation cost follow independently of any fine.
- A risk register turns worry into a work plan. Rate each risk by likelihood and impact, multiply to a score, assign an owned mitigation, and work the list in score order.
- Retention is a security control. Every record kept past its purpose is a record that can be lost. Set a different clock per category, automate the deletion, and pair it with role-based access and audit logs.
- Transparency is a trust-builder, not a liability. Disclosing AI use, as New York City Local Law 144 requires with at least ten business days of advance notice, data-category disclosure, and an annual independent bias audit, signals seriousness that strong candidates weigh.
Frequently Asked Questions
Must we comply with GDPR if we do not recruit in Europe?
Yes, if you have any EU candidates. GDPR applies to the processing of personal data belonging to EU residents regardless of where your company or your servers are based, so if you post roles on a global job board you have EU candidates whether or not you targeted them. Most organizations adopt GDPR as the baseline standard because it is generally the strictest; comply with it and you are usually close to compliant elsewhere, with the specific state and national laws that apply to you still worth checking individually.
Do we need explicit consent to use AI to screen resumes?
It depends on what the screening actually does. If AI extracts qualifications from a resume for a recruiter to review, which is a normal recruiting task, explicit consent may not be required. If AI makes predictions about personality, likelihood to stay, or cultural fit, you likely do need it. If the vendor uses the data to train their model, you definitely do. When in doubt, ask and disclose; it costs very little and protects both the compliance position and the trust relationship.
How do we respond to candidate requests for data access or deletion?
Build a repeatable process rather than handling each request improvisationally: document the request, gather the data from your applicant tracking system and every vendor that holds a copy, provide access or delete as requested, and confirm completion back to the candidate. Respond within the timeline your applicable regime sets. This is legally required under regimes including GDPR and CCPA, and slow or absent responses can themselves draw regulatory penalties, entirely separately from whatever the underlying data issue was.
Are vendor security certifications enough assurance on their own?
No. Certifications are useful indicators and a reasonable entry requirement, but they do not transfer the risk assessment to the auditor. Ask the detailed questions anyway: how is data encrypted, who has access, what is the incident response process, how long is data retained, and who are the subprocessors. A vendor with a strong certification and a vague retention policy is still a live risk, so use certifications as a floor and supplement them with a specific assessment and a contract.
What is our liability if a vendor's system is breached?
You are liable. You are responsible for candidate data even when a vendor handles it, so a breach in their system still obliges you to notify affected candidates and the relevant regulators. The exposure includes regulatory fines, which under GDPR reach up to 4 percent of annual global revenue, candidate litigation, and reputational damage in the talent market. This is why due diligence, a real data processing agreement, and ongoing monitoring matter: they are the only levers you have over a risk you cannot hand away.
Skill.re