←
AI for Recruiters
Proficient · M26 · lesson 26 of 32 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Retention and Deletion: Reasonable Timelines and Clean Data Practices

15 min

Dana leads talent acquisition at a 600-person fintech that hires across the US and the EU, and she inherited a problem she did not create: an applicant tracking system holding eleven years of candidate records, roughly 240,000 profiles, almost none of which anyone could explain a reason to keep. Resumes, AI screening notes, interview scorecards, rejected applicants from 2016, demographic self-identification forms, and free-text recruiter comments all sat in one undifferentiated pool. When a German candidate filed a right-to-erasure request and, three weeks later, a discrimination charge landed that required Dana to produce records for a specific role, she realized she had the worst of both worlds: she was keeping data she had no right to keep, and she could not reliably find the data she was legally obligated to retain. This lesson is the retention and deletion discipline Dana built to fix that, and the schedule she now runs the whole function on.

Why Retention Is a Two-Sided Obligation

Most recruiters think about data retention as a single question: how long do we keep things? It is actually two opposing obligations pulling against each other, and a defensible practice lives in the space between them. On one side, anti-discrimination law in the United States requires you to keep certain records long enough to be auditable. The EEOC, under Title VII regulations, requires employers to preserve all personnel and employment records relevant to a hiring decision for at least one year from the date the record was made or the action was taken, whichever is later. For federal contractors, OFCCP rules generally extend this to two years. If a charge is filed, you must preserve all related records until final disposition of the matter, no matter how long that takes. Throwing data away too early is not a privacy virtue; it can be spoliation of evidence.

On the other side, privacy law pushes the opposite way. The GDPR's storage limitation principle, in Article 5(1)(e), says personal data must be kept in a form that permits identification of individuals for no longer than is necessary for the purposes for which it was processed. Data minimization, in Article 5(1)(c), says you should not collect or hold more than you need in the first place. And the right to erasure, Article 17, gives individuals the ability to request deletion of their data when it is no longer necessary, when consent is withdrawn, or when there is no overriding legitimate basis to keep it. Dana's eleven-year pool violated all three.

The discipline is not picking a side. It is defining, per data category, the shortest retention period that still satisfies your legal recordkeeping floor, and then actually deleting on schedule. Notice what that formulation rules out. A single organization-wide "keep everything for seven years" rule fails the privacy half even though it clears the recordkeeping half comfortably, and a blanket "delete after twelve months" rule fails the recordkeeping half for anything touched by a charge. Neither is defensible, and both feel decisive, which is why teams reach for them. The category-by-category version is more work exactly once, at design time, and less work every month afterwards.

The Clean Data Principle: Less Is Safer

Before Dana could build a schedule, she had to confront how much of her data should never have been collected or kept at all. Clean data practice starts with minimization at the point of capture. Every field a recruiting system stores is a field that can be breached, subpoenaed, mishandled, or used to infer a protected characteristic. The free-text recruiter comments in Dana's system were the clearest hazard: notes like "seemed like a culture fit" or "might have childcare constraints" are both useless for a real hiring defense and radioactive in a discrimination case, because they document subjective impressions that correlate with protected status.

Look closely at why those two examples are worse than merely unhelpful. Neither states a job-related reason, so neither can support a decision if it is ever challenged, which means they carry no defensive value whatsoever. But both are readable by an investigator as evidence about what the decision-maker was actually weighing, and the second reads directly onto a protected characteristic without ever naming it. A comment that helps nobody and can only hurt you is not a neutral artifact of an informal culture; it is an unforced liability the system invited by offering an empty box. Clean data means structured, role-relevant fields wherever possible, and ruthless discipline about what free text is allowed to hold.

Minimization also means separating data by sensitivity. Demographic self-identification data, collected for EEO and OFCCP reporting, must be kept apart from the records used to make hiring decisions, precisely so that a decision-maker cannot see it and so an auditor can confirm it was not a factor. AI screening artifacts deserve their own category too: the prompt used, the model's output, and the recruiter's override decision are exactly what you will need if you ever have to show that an automated tool did not produce a disparate impact. Lumping everything into one pool, the way Dana inherited it, makes both deletion and defense impossible, because you cannot delete one category without risking another.

That last sentence is the operational trap worth naming plainly, because it explains why Dana's predecessor kept everything. Faced with one undifferentiated pool, any deletion carries an unknown risk of destroying something a charge might need, so the safe-feeling choice is always to defer. The pool then grows, which makes the next deletion decision harder still, and the organization arrives at eleven years of records not through a decision but through eleven years of postponement. Segregation by category is what breaks that loop, because it converts one impossible judgment into several routine ones.

A Worked Retention Schedule

Here is the schedule Dana built, category by category. The periods below are an illustrative design, not legal advice for any specific jurisdiction; the point is the reasoning, and the legal floors named are stated accurately. Every category gets a defined retention period, a basis, and a deletion trigger. Those three fields are not bureaucratic decoration. The period tells the system when to act, the basis is what you say to a regulator or an investigator who asks why that number, and the trigger is what makes the deletion actually happen rather than remain an intention.

Data category Retention period Basis Deletion trigger
Applications and resumes, non-hired candidates Two years from the hiring decision The OFCCP two-year recordkeeping rule, applied company-wide as a uniform floor Monthly automated purge of decisions older than 24 months, unless flagged for litigation hold
AI screening outputs and override records Same two-year clock as the application they relate to The same recordkeeping obligation: a screening output that influenced a decision is a record relevant to that decision Stored linked to the application so they are preserved or deleted together
Interview scorecards and structured evaluations Two years from the decision for non-hired candidates Recordkeeping for the hiring decision they document For hired candidates, converts to an employee record under the separate, longer employee-records policy
Demographic and EEO self-identification data The period required for EEO-1 and affirmative action reporting EEO and OFCCP reporting obligations Its own schedule, independent of the application record, in a segregated system
Talent-pool and silver-medalist data 12 months from last meaningful contact Legitimate business interest in future contact, bounded by GDPR storage limitation Re-consent prompt; deleted if the candidate does not re-engage or affirmatively consent
Anything under litigation hold Until final disposition of the matter The EEOC preservation requirement Frozen and exempted from automated deletion, then released back to the normal schedule

Several of those rows repay a closer look. Dana applies the OFCCP two-year rule company-wide as her floor, even for business units that are not federal contractors, because it comfortably exceeds the EEOC one-year minimum and keeps her practice uniform. That choice costs her twelve extra months of storage on some records and buys her something more valuable: one number to explain, one rule to automate, and no scenario where a record is deleted because someone misclassified which business unit a requisition belonged to.

The linkage between AI screening outputs and the application they relate to is the row most often missing in practice. If the screening output influenced a decision, it is a record relevant to that decision and must be auditable for the same period, so storing it on a separate clock creates two failure modes at once: the artifact can expire while the application it explains survives, leaving you with a decision you cannot reconstruct, or it can outlive the application and sit in the system with no retention basis at all. Preserving and deleting them together is what keeps the pair coherent.

Talent-pool data is the trickiest category, because keeping a strong candidate's profile to contact them later is a legitimate business interest, but it cannot be indefinite. Dana sets a 12-month retention from last meaningful contact, with a re-consent prompt. If the candidate does not re-engage or affirmatively consent to remain in the pool, the record is deleted. Under GDPR, keeping someone in a talent pool indefinitely without a fresh basis is exactly what storage limitation prohibits, and the re-consent prompt is what refreshes the basis rather than merely restarting a timer.

Handling Right-to-Erasure Requests

When the German candidate asked Dana to delete his data, she could not simply hit delete, and that nuance matters. The right to erasure under Article 17 is not absolute. It yields where the data is still necessary for compliance with a legal obligation or for the establishment, exercise, or defense of legal claims. So Dana's response was not "yes" or "no" but a reasoned split: she deleted the marketing and talent-pool portions of his data immediately because there was no overriding basis to keep them, and she retained the specific application and decision records for the duration of her recordkeeping obligation, informing him in writing that those records were held under a legal-obligation exception and would be deleted when that period expired.

That written, reasoned response is itself the defensible artifact. It does three things at once that a bare yes or no cannot. It gives the candidate a real outcome on the part of his request that was straightforwardly grantable rather than making him wait on the part that was not. It states the basis for the retained portion, which is the thing a supervisory authority would ask for if he escalated. And it commits to a date, which converts an indefinite refusal into a bounded one. An erasure process that has no way to distinguish "delete now" data from "retain under legal obligation" data will either over-delete and destroy evidence or over-retain and breach privacy law. There is no third option available to a team that cannot tell the two apart.

The operational requirement underneath this is the ability to find one person's data across every system. Dana's eleven-year pool failed this test completely; she could not confidently locate every copy of one candidate's records across the ATS, the email system, exported spreadsheets, and AI tool logs. Note that only the first of those four is a system anyone would think to search. Exported spreadsheets in particular are where retention policies go to die, because a copy taken years ago for a hiring committee sits outside every rule the ATS enforces and is invisible to every job that runs against it. A retention policy you cannot execute is not a policy. Knowing where personal data lives, and being able to retrieve or delete it per individual, is the precondition for honoring both erasure rights and litigation holds.

Automating Deletion So It Actually Happens

The single most common failure Dana found was that deletion was always somebody's good intention and never anybody's job. Manual deletion does not happen, because it is never urgent and always feels risky to the person clicking the button. Both halves of that sentence matter. Nothing breaks on the day a record is not deleted, so the task loses every scheduling contest it enters, and the individual asked to perform it bears all the downside of deleting something that turns out to be needed while the upside accrues to the organization in the abstract. Under those incentives, the rational individual defers, and the deferral is invisible.

The fix is automation governed by the schedule. Each data category gets a retention period encoded as a rule in the system, a recurring job that identifies records past their period, and a litigation-hold flag that exempts frozen records. The job runs on a fixed cadence, logs what it deleted and why, and produces an audit trail showing the policy was applied consistently. That audit trail is doing double duty: it proves to a privacy regulator that you delete on schedule, and it proves to a discrimination investigator that any missing record was deleted under a documented, uniformly applied policy rather than to hide something specific.

Consistency is the legal shield here. A deletion that happens to one candidate's records but not another's looks like spoliation. A deletion that happens to every record of a given category at a given age, logged and uniform, looks like good governance. The difference between those two interpretations is entirely whether you can show the rule and show it was applied to everyone. This is also why the litigation-hold flag has to live inside the automated job rather than beside it as a manual step. A hold that depends on someone remembering to pause the purge will eventually fail on the one matter where it counted, and the resulting gap will be a hole in exactly the population under scrutiny, which is the worst possible shape for a deletion record to have.

Building the Policy and Keeping It Alive

A retention schedule is not a one-time project; it is a living document with an owner, a review cadence, and a clear home in the privacy program. Dana's schedule names a single accountable owner, lists every data category with its period, basis, and trigger, and gets reviewed at least annually or whenever a relevant law changes. The single owner is doing more work than it appears. A schedule owned by a committee is a schedule where every category is somebody's job in principle and nobody's in practice, which is the same failure mode as manual deletion moved up a level.

New AI tools are the most common reason a schedule drifts out of date: every time the function adopts a tool that generates a new kind of record, that record needs a category and a retention period before it goes into production, not after it has quietly accumulated for two years. The sequencing is the whole point. A record type that enters production without a category does not sit still waiting to be classified; it starts accumulating immediately, and by the time anyone notices, the decision is no longer "what period should this have" but "what do we do with two years of records we have no basis for." The same discipline that governs why you keep data governs whether you should have collected it at all, which closes the loop back to minimization.

Done well, this turns retention from a source of risk into a source of confidence. Dana can now answer an erasure request in days, produce a complete set of records for any charge within her retention window, and tell an auditor exactly what she keeps, why, and when it is destroyed. The eleven-year pool is gone, deleted in stages under a documented policy, and what remains is the minimum she is required to hold, held only as long as she is required to hold it.

Anti-Patterns

Keeping everything, just in case. This is the eleven-year pool: no deletion rule, no categories, and a general sense that data might be needed someday. It happens because retaining is passive and deleting is an action someone has to take and defend, and because the recordkeeping obligations are real enough to make caution feel principled. What goes wrong is that the caution only addresses one of two obligations, so the organization is continuously in breach of storage limitation and data minimization while accumulating records that expand the blast radius of any breach or subpoena. The counter is a category-by-category schedule where every category has the shortest period that clears its legal floor, so keeping data becomes a decision with a stated basis rather than the default that happens when nobody decides.

Deleting on a single blanket clock. This is the opposite overcorrection, a policy that purges everything at twelve or twenty-four months regardless of what it is. It happens after a privacy review, because one number is easy to explain and easy to automate, and because it genuinely does fix the over-retention problem. What goes wrong is that a blanket clock cannot honor a litigation hold it does not know about and cannot distinguish demographic data on its own reporting schedule from a talent-pool record on a 12-month re-consent basis, so it deletes into an active matter and calls it compliance. The counter is period, basis, and trigger defined per category, with holds implemented as an exemption inside the job rather than a note beside it.

Treating erasure as a yes-or-no question. This is answering a right-to-erasure request by either deleting the candidate's entire footprint or refusing on the grounds that records are legally required. It happens because Article 17 is described as a right to deletion, which sounds binary, and because splitting the request requires knowing which records sit under which basis. What goes wrong is that both answers are wrong in the same case: full deletion can destroy records held under a recordkeeping obligation or needed to defend a claim, and blanket refusal keeps marketing and talent-pool data that has no overriding basis at all. The counter is the reasoned split, delivered in writing, naming what was deleted, what was retained, under which exception, and when it will be deleted.

Leaving deletion to human diligence. This is a written schedule with correct periods that nobody has encoded into a job. It happens because writing the policy feels like completing the work, and because the first few months of manual compliance actually succeed while attention is high. What goes wrong is that deletion is never urgent and always personally risky, so it loses every scheduling contest and quietly stops, and the resulting pattern of some records deleted and others not is the shape that reads as spoliation rather than governance. The counter is an encoded rule per category, a recurring job that logs what it deleted and why, and a hold flag the job respects, so uniformity is a property of the system rather than of anyone's memory.

Letting a new tool into production without a category. This is adopting an AI screening or interview tool and discovering months later that it has been generating a record type the schedule does not cover. It happens because the retention question belongs to nobody during a tool evaluation, and because the new artifacts look like a byproduct rather than a record. What goes wrong is that the records accumulate from day one with no period and no basis, so by the time anyone asks, the question has changed from what period to assign into what to do with a backlog you cannot justify holding. The counter is to make a retention category a precondition of going live, assigned by the schedule's named owner before the first record exists.

Practice

These exercises rebuild Dana's schedule in the order she built it, and the first one is what makes the rest possible.

  • Inventory where candidate data actually lives. List every place one candidate's data could exist: the ATS, email, exported spreadsheets, AI tool logs, shared drives, anything else. For each, note whether you could locate and delete that person's records on request. The systems you cannot answer for are the ones that will break both an erasure request and a litigation hold.
  • Categorize what you hold and mark what should never have been collected. Sort your recruiting data into categories by sensitivity and purpose. Flag free-text fields that document subjective impressions rather than job-related reasons, and check whether demographic self-identification data is stored where a hiring decision-maker can see it.
  • Write period, basis, and trigger for every category. For each category, set the shortest retention period that still clears its legal recordkeeping floor, state the basis in one sentence you would be willing to give a regulator, and define the specific event or job that causes deletion. Any category missing one of the three is not finished.
  • Trace a right-to-erasure request end to end. Take a real past candidate and work out which parts of their footprint you would delete immediately and which you would retain under a legal-obligation exception. Then draft the written response, naming what was deleted, what was retained, on what basis, and when it will go.
  • Test the litigation hold before you need it. Pick a role family and establish exactly what would freeze if a charge were filed tomorrow, who sets the flag, and whether your deletion job would respect it. If the hold depends on a person remembering to pause a purge, you have found the failure.

Reflection

  • How many years of candidate records are you currently holding, and can you state a basis for the oldest one?
  • If a candidate asked you to delete their data tomorrow, how many systems would you have to search, and which one would you forget?
  • What do your recruiters write in free-text fields, and how would those notes read to an investigator?
  • Which of your data categories has a retention period that nobody can explain the reasoning behind?
  • Is deletion in your organization a job that runs, or an intention someone holds?
  • What record types has your newest AI tool started generating, and where do they sit in the schedule?

Glossary

  • Storage limitation. The GDPR principle in Article 5(1)(e) that personal data must be kept in a form permitting identification of individuals for no longer than is necessary for the purposes for which it was processed.
  • Data minimization. The GDPR principle in Article 5(1)(c) that you should not collect or hold more personal data than you need in the first place.
  • Right to erasure. The Article 17 right allowing individuals to request deletion of their data when it is no longer necessary, when consent is withdrawn, or when there is no overriding legitimate basis to keep it. It is not absolute.
  • Legal-obligation exception. The circumstance in which the right to erasure yields, where data is still necessary for compliance with a legal obligation or for the establishment, exercise, or defense of legal claims.
  • EEOC recordkeeping floor. The Title VII requirement to preserve all personnel and employment records relevant to a hiring decision for at least one year from the date the record was made or the action was taken, whichever is later.
  • OFCCP two-year rule. The federal contractor recordkeeping period that generally extends the retention floor to two years, and which Dana applies company-wide for uniformity.
  • Litigation hold. The freezing of every record relevant to a filed charge or claim, exempt from automated deletion until final disposition of the matter, then released back to the normal schedule.
  • Spoliation. Destruction of evidence. Deleting too early, or deleting selectively rather than uniformly, is what turns a purge into this.
  • Period, basis, and trigger. The three fields every category in a retention schedule needs: how long it is kept, the stated reason for that length, and the specific event or job that causes deletion.
  • Segregation by sensitivity. Storing demographic and EEO self-identification data apart from records used to make hiring decisions, so a decision-maker cannot see it and an auditor can confirm it was not a factor.
  • AI screening artifacts. The prompt used, the model's output, and the recruiter's override decision, kept on the same clock as and linked to the application they relate to.
  • Re-consent prompt. The step that refreshes the basis for holding talent-pool data, rather than merely restarting a timer, when a candidate is asked whether they wish to remain in the pool.

Closing

Dana's two bad weeks were not caused by ignorance of the law. She knew that candidates had erasure rights and she knew that charges required records. What she did not have was a structure that let her act on both facts at once, and without that structure the two obligations simply cancelled into paralysis: she could not delete safely and she could not produce reliably, so she did neither well. Eleven years of records were the visible symptom. The absence of categories was the cause.

The fix is available to any recruiting team willing to do the design work once. Name your categories. Give each one the shortest period that clears its legal floor, a basis you would say out loud to a regulator, and a trigger that fires without anyone deciding. Segregate demographic data from decision records, link AI artifacts to the applications they explain, and put holds inside the job rather than beside it. Then give the schedule an owner and a review date, and refuse to let a new tool reach production without a category. What you get back is not just compliance. It is the ability to answer a candidate in days, an investigator in full, and an auditor without hesitating.

Key Takeaways

  • Retention is two opposing obligations, not one. US anti-discrimination law sets a floor you must keep above; the EEOC requires relevant hiring records for at least one year and OFCCP requires two years for federal contractors. GDPR storage limitation and the right to erasure set a ceiling you must keep below. A defensible schedule is the shortest period that still clears the legal floor for each category.
  • Minimize at capture, then segregate by sensitivity. The safest data is the data you never collected. Keep structured role-relevant fields, discipline free-text comments that document subjective impressions, and store demographic and EEO data separately from decision records so it cannot influence or appear to influence a hiring decision.
  • Build a category-by-category schedule with period, basis, and trigger. Non-hired applications and their linked AI screening outputs on a two-year clock, talent-pool data on a 12-month re-consent clock, EEO data on its own segregated schedule, and a deletion trigger defined for each so nothing is kept by inertia.
  • The right to erasure is conditional, not absolute. Article 17 yields where data is still needed for a legal obligation or to defend legal claims. Answer requests with a reasoned split: delete what has no overriding basis, retain decision records under the legal-obligation exception, and document the reasoning in writing.
  • Litigation holds always override routine deletion. When a charge is filed, freeze every relevant record until final disposition. Deleting on a normal schedule into an active matter is spoliation, and the preservation duty outranks any purge.
  • Automate deletion or it will not happen. Encode each period as a rule, run a recurring job that logs what it deleted and why, and exempt held records. Uniform, logged deletion is a legal shield; ad hoc deletion looks like evidence destruction.
  • Keep the schedule alive as tools change. Assign an owner, review at least annually, and give every new AI tool a retention category before it reaches production so new record types do not silently accumulate outside the policy.

Frequently Asked Questions

Can we just keep everything and avoid the deletion problem entirely? That is the choice Dana inherited, and it fails the privacy half of the obligation continuously rather than occasionally. Storage limitation and data minimization are not triggered by a request; they describe the state your data is supposed to be in at all times, so an eleven-year pool is in breach on every day it exists. It also makes the recordkeeping half harder rather than easier, because a pool nobody has categorized is a pool nobody can search reliably when a charge requires specific records. Keeping everything buys the appearance of safety and delivers the opposite on both sides.

A candidate asked us to delete everything. Do we have to? Not necessarily everything, and the correct response is a split rather than a verdict. The right to erasure yields where data is still necessary for compliance with a legal obligation or for the establishment, exercise, or defense of legal claims, so application and decision records within your recordkeeping period can generally be retained on that basis while marketing and talent-pool data, which has no overriding basis, goes immediately. Put the split in writing, name the exception you are relying on for the retained portion, and state when it will be deleted. That written reasoning is the artifact that makes the decision defensible.

What happens if our automated purge deletes records that a charge later needs? This is precisely what the litigation-hold flag exists to prevent, and why it has to be implemented inside the deletion job rather than as a separate manual step. Once a charge or claim is filed, the preservation duty attaches to every relevant record until final disposition and outranks any routine schedule. If a purge has already run before the hold was set, the defensibility of what remains rests on being able to show the rule and show it was applied uniformly, which is the entire reason the job logs what it deleted and why. Uniform deletion under a documented policy is a different thing from selective deletion, and the log is what distinguishes them.

Our AI vendor stores candidate data too. Does our schedule cover that? Only if the contract makes it. A retention schedule governs the systems you control, and screening artifacts sitting in a vendor's environment are outside it unless deletion terms were negotiated. This is also why AI screening outputs need a category of their own on your side: the prompt, the output, and the recruiter's override are what you would rely on to show that an automated tool did not produce a disparate impact, so you need them preserved on the same clock as the application and deleted with it, wherever they physically live.