←
AI for Recruiters
Strategic · M11 · lesson 11 of 33 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Documentation and Evidence: Building a Trail for Compliance

15 min

Omar runs talent acquisition for a 1,200-person logistics company that hires about 600 warehouse and operations roles a year. Last spring his team turned on an AI resume-screening tool to triage a backlog of roughly 4,000 applications. It worked beautifully until a rejected candidate filed a charge with the Equal Employment Opportunity Commission alleging that the screening disadvantaged applicants over 40. The EEOC sent a request for information. Omar sat down to assemble the company's response and discovered the uncomfortable truth at the center of this lesson: his team had made thousands of fast, mostly reasonable decisions, but almost none of them had been written down. The AI had produced scores. Nobody had recorded what the scores were based on, who reviewed them, or why a given candidate had been cut. Omar could not reconstruct the trail because the trail did not exist.

Why the Record Matters More When AI Is Involved

Documentation feels like busywork when you are moving fast, but it serves two purposes at once. The first is defense: if a hiring decision is ever questioned by a candidate, a regulator, or a court, your records are the evidence that decisions were made on job-related criteria and applied consistently. The second is improvement: without a record you cannot analyze whether your process treated similar candidates similarly, whether selection rates differed across groups, whether bias crept in at a particular gate, or whether the AI tool drifted over time. Teams that skip documentation believe they are trading paperwork for speed; what they are actually trading is their ability to answer a question they will eventually be asked.

When an automated tool screens candidates, the stakes rise on both counts. A human screener can, if pressed, explain their reasoning after the fact. An algorithm cannot be deposed. The only account of why the tool ranked a candidate the way it did is the account you captured at the time, which means you need to document which candidates were screened by AI, how the tool made its decision, what criteria it used, and what human review happened afterward. Regulators increasingly expect that account to exist. The EEOC has been explicit that using AI in hiring does not relieve an employer of its obligations under Title VII, the Age Discrimination in Employment Act, and the Americans with Disabilities Act, and that employers remain responsible for the outcomes of tools they deploy, including tools built by vendors. If the tool produces an adverse impact on a protected group, "the vendor told us it was validated" is not, by itself, a defense. Your own evidence trail is.

What to Document: The Essentials

Hiring decisions involve a lot of information, and you do not need to keep all of it. You need to keep enough to recreate why a decision was made. Seven categories carry that weight. Candidate identity and source is the baseline: who the candidate is and where they came from, whether direct application, referral, recruiter outreach, or a job board. Source matters more than it looks, because it is one of the dimensions along which inconsistent treatment tends to show up. Resume and application materials should be retained exactly as the candidate submitted them, since that is the source data behind every downstream decision, and a parsed or summarized version is not a substitute. Screening assessments covers any skills test, coding challenge, or structured assessment score, and, when AI screened the candidate, the score or ranking the tool produced.

Decision criteria must be written down before screening begins, not reverse-engineered afterward. "To advance to phone screen, candidates must have Python experience, three years of relevant experience, and a bachelor's degree or equivalent" is explicit enough that you can later verify decisions actually followed it. Screener notes record the reviewer's assessment in specific terms: "Candidate has the required skills. No mention of relevant project experience in the resume. Will ask about this in the phone screen." Interview feedback should come from each interviewer on the dimensions they were assigned, with enough detail to be checked: "Technical skill, 4 of 5: strong understanding of distributed systems, answered the question on X correctly, had gaps on Y but showed a learning mindset. Communication, 4 of 5: explained concepts clearly and asked clarifying questions." Finally, the decision rationale pulls the inputs together into a conclusion: "Advancing to second interview. Candidate meets requirements. Technical skills strong. Communication solid. Fit concerns need more exploration."

When a tool is in the loop, six fields carry the evidentiary weight, and a record missing any of them cannot connect an outcome to an accountable process. The job-related criteria, tied to the actual demands of the job: "forklift certification, two years warehouse experience, ability to lift 50 pounds" is defensible, while "culture fit" and "polish" are not, and they are exactly the kind of vague criteria that invite an adverse-impact claim. The AI tool and its version, plus the date range it was in use, because vendors update models and a tool that screened cleanly in January may behave differently in June. The output the tool produced, meaning the score or recommendation for that candidate rather than only the final human decision, because the raw output is what a bias audit examines for disparate selection rates. The named human reviewer who reviewed that output and made or confirmed the decision, since meaningful human oversight is both good practice and, in many frameworks, a legal expectation. The rationale in terms of the criteria: "Rejected: no forklift certification, a required qualification for this role." And the date, because a regulator can tell the difference between notes written the day of the decision and notes written the week of the charge.

How Much Detail, and When

The question recruiters ask most is how much detail is enough. Enough to recreate your thinking, and not so much that documentation becomes burdensome and stops getting done. The right standard is proportionality: the closer a decision sits to a boundary, or the more likely it is to be questioned, the more detail it warrants. For routine, clear-cut decisions, a brief note tied to a criterion is sufficient: "Does not meet minimum years of experience requirement." For decisions near the threshold, where a candidate is neither clearly qualified nor clearly unqualified, write more: "Two years of directly relevant experience against a three-year requirement, but a strong recent trajectory in a related role; advancing to phone screen to assess depth." For offer-level decisions, document the full picture, including the interview feedback, the team's reasoning, any concerns raised, and how those concerns were resolved.

Two situations deserve extra care. Any decision that could plausibly be challenged, especially rejecting a candidate who shares a protected characteristic where the call was at all arguable, should be documented thoroughly and honestly. So should the reverse case: if you extend an offer to an employee's friend or family member, document the genuine, job-related qualifications behind the decision and be prepared to explain them. For AI-assisted decisions specifically, document the tool's methodology as well as its output: how the tool scored this candidate, what criteria it used, what score it produced, who reviewed that score, and what the reviewer concluded. The governing standard across all of these is a single test. Documentation should be specific enough that someone unfamiliar with the candidate could understand why the decision was made and could verify that the decision followed consistent criteria.

The word honestly is load-bearing there. The most dangerous record is one that tells the wrong story: if a screener's note says "meets requirements" but the real reason for rejection was an unmeasured, subjective impression, the documentation contradicts the actual decision and looks like a cover-up the moment it is examined. Document what truly drove the decision, write down the observation that mattered, and record the concern a reference surfaced. Honest records are more defensible than tidy ones.

Structured Formats and Organization

Documentation can be unstructured, as free-form notes, or structured, as scorecards and forms. Structured documentation is preferable because it forces the same fields for every candidate, which is what makes consistency possible and analysis feasible later. A simple structured screening form captures the candidate name, the source (direct application, referral, LinkedIn, or other), the screener's name, and the date; then each required criterion with a yes, no, or not-assessed answer, such as Python required, years of relevant experience, and degree or equivalent; then the optional criteria on the same scale, such as project leadership experience or public speaking experience; and finally an overall assessment recording whether the candidate meets minimum requirements and the recommendation, which is advance to phone screen, hold for future, or reject. The structure is doing real work: it ensures every screener decides against the same criteria, and it makes it straightforward to check afterward whether those criteria were applied fairly across candidates.

Organization matters as much as capture. Each candidate should have a single file, physical or digital, containing all of their materials: resume, application, assessment scores, interview feedback, and decision rationales. Organize by hiring cycle, role, and candidate name so records can actually be retrieved when someone asks. If you use an applicant tracking system, confirm it is capturing and retaining everything necessary rather than only what it displays by default; many systems retain the final disposition but quietly discard the intermediate algorithmic output, which is precisely the evidence a bias audit needs.

A Worked Example: Reconstructing the Trail

Return to Omar. The EEOC charge concerns a 47-year-old applicant, call her Diane, who applied for a warehouse operations role and was rejected at the AI-screening stage. The agency's request for information asks the company to explain the basis for the decision and to produce records. Omar had two things: Diane's original resume in the applicant tracking system, and a numeric score from the screening tool, 62 out of 100. What he did not have was almost everything that mattered. No written record of the job-related criteria the score was meant to measure. No note of which version of the tool produced the 62. No named human reviewer attached to the rejection, because the team had set the tool to auto-reject everything below 70 and nobody had looked at her file. No rationale connecting the 62 to anything specific about the role, and because the auto-reject ran in a batch, no contemporaneous decision note at all, just a status flag.

Compare that to the record Omar needed. With a complete trail, the response writes itself: the role required forklift certification and two years of warehouse experience, both documented before screening began. The tool, version 3.2, in use from March 1 to April 15, scored candidates against those criteria. Diane scored 62 because, per the captured output, the tool found no forklift certification in her materials, which the documented criteria listed as required. A named recruiter reviewed the auto-flagged rejections on April 3, confirmed the missing certification against the resume, and recorded the rationale. The selection-rate analysis run that quarter showed no statistically meaningful difference in pass rates between applicants over 40 and under 40. The first version invites an inference of discrimination, because the absence of records looks like the absence of legitimate reasoning. The second is a defense. The decision may have been identical in both cases; only the documentation differs, and the documentation is what the EEOC evaluates.

Retention: How Long to Keep Records

Records only help if they still exist when you need them. Federal rules set a floor. Under the EEOC's recordkeeping regulations, employers covered by Title VII and the ADA must generally preserve personnel and employment records, including application materials and records relating to hiring decisions, for a defined period after the record is made or the action is taken. The widely applied baseline for most private employers is one year, and the ADEA's recordkeeping requirements call for retaining certain hiring records for a similar minimum period. Critically, if a charge of discrimination is filed, the obligation changes: you must preserve all records relevant to that charge until the matter is finally resolved, regardless of the normal retention clock. Federal contractors and some state laws impose longer minimums, which is why many recruiting teams adopt an internal standard of three years as a practical policy, comfortably above the federal floor and closer to the window in which a legal challenge realistically arrives.

The practical takeaway is to pick a period that meets or exceeds every legal floor you are subject to, apply it consistently, and never delete records connected to a pending or anticipated complaint. Verify the exact periods for your jurisdiction and contractor status rather than relying on a remembered number, because the floors vary and a longer obligation always governs. Just as important, confirm your applicant tracking system actually retains the AI-specific fields, meaning tool version, raw score, and reviewer, for the full period.

NYC Local Law 144 and Data Minimization

If you hire for positions in New York City, a separate regime applies on top of federal recordkeeping. NYC Local Law 144 governs automated employment decision tools. It requires that such a tool undergo an independent bias audit within the year before it is used, that a summary of the audit results be published publicly, and that candidates and employees receive notice that an automated tool will be used, typically at least ten business days before it is used, along with information about the job qualifications and characteristics the tool assesses. Your documentation trail is what makes compliance with these obligations provable: the audit on file, the published summary, the dated notices sent to candidates, and the per-candidate records showing the tool was applied as described. A bias audit cannot be conducted, and a Local Law 144 inquiry cannot be answered, without exactly the captured outputs and criteria described earlier in this lesson.

Data protection law pulls in a complementary direction worth holding in tension with retention. If you process candidate data subject to the European Union's General Data Protection Regulation, the data-minimization and storage-limitation principles say you should collect only the personal data you need and keep it only as long as you have a lawful basis. This is not a contradiction with recordkeeping; defending against discrimination claims and meeting legal retention duties is a legitimate basis for keeping the relevant records for the required period. The discipline GDPR adds is intentionality: keep the fields that serve compliance and improvement, set a defined expiry aligned to your legal obligations, and do not hoard unrelated personal data just in case. A trail complete on the dimensions that matter and minimal everywhere else is both more defensible and more respectful of candidates.

Building the Trail Into the Workflow

The reason Omar had no records was not laziness. It was that documentation lived outside the workflow, as an afterthought to be done later, which in practice means never. Make capture automatic instead. Configure the applicant tracking system so the tool's version and score write to the candidate record without anyone remembering to do it. Require a named reviewer and a rationale before any AI-flagged rejection can be finalized, so the oversight step cannot be skipped under deadline pressure. Run selection-rate analysis as a standing process rather than a fire drill triggered by a complaint. And assign ownership, because a standard that is everyone's job is no one's job: one person owns the policy, audits a sample of records each cycle, and confirms the AI-specific fields are actually being retained.

Three Anti-Patterns

Retrospective documentation. Decisions get made verbally in the hallway. Candidates are rejected by email with no documented feedback. Offers go out after a phone call. Months later someone asks why a particular candidate was not hired and nobody remembers. It happens because documentation feels like extra work when you are busy and skipping it carries no immediate cost. What goes wrong is that you have no record when a decision is questioned, no defense if a candidate claims discrimination, and no way to analyze whether decisions were consistent. The fix is to build documentation into the workflow so it happens as the work happens, and to make it someone's explicit responsibility.

Over-documentation that nobody reads. A team implements a comprehensive system: screeners fill out detailed forms, panels write long narrative feedback, offers carry exhaustive rationales. It is thorough, it takes so much time that screeners feel slowed down, and nobody ever reads any of it unless there is a legal issue. It happens through an abundance of caution, the instinct to document everything in case anything is needed. What goes wrong is that hiring slows, people resent the process, and records nobody reads never improve anything. The fix is proportionality: document what compliance and improvement require and no more, efficiently enough that it survives a busy quarter.

Documentation that tells the wrong story. Screening feedback reads "candidate has required experience," but the actual decision turned on something unmeasured, such as a secondhand report that the person would not fit the culture, unverified and untested. This is the legally dangerous case, because the documentation asserts that objective criteria were applied while the underlying decision was subjective. It happens because screeners want their decisions to look defensible, so they document formal criteria even when those are not what drove the call. What goes wrong is that the documentation contradicts the actual decision-making when challenged, which reads as dishonest. The fix is to document what actually drove the decision: if culture fit played a role, write down the specific observation behind that judgment, and if a reference surfaced a concern, write down the concern.

Practice

  • Design a structured screening form for one specific role. Decide what every screened candidate must be assessed against, and build the form so that any screener using it evaluates candidates the same way.
  • Audit your current practice. For your last ten hires and ten rejections, review the documentation that exists. Is there enough to understand why each decision was made? Where are the gaps, and do they cluster anywhere?
  • Create an organizing system for candidate records. Decide how materials are stored by hiring cycle, role, and candidate so they can be retrieved quickly for analysis or defense.
  • Write documentation for a decision you made recently, a hire or a rejection. Then hand it to someone unfamiliar with the candidate and ask whether they can follow your reasoning and identify what is missing.
  • Run the discrimination-claim scenario. If a candidate claimed they were rejected on a protected basis, what could you actually produce today to defend the decision, and what would it take to close the gaps?

Reflection

  • What is the most common reason you reject candidates, and is that reason explicit in your documentation or only in your head?
  • If a candidate from a protected group asked why they were rejected, what would you be able to tell them based on the record you have?
  • What do you wish you had documented about a past hire who did not work out, and what would that information have told you?
  • How would you explain to your team why documentation matters without making it sound like you do not trust their decisions?
  • When you are rushed, which documentation do you skip first? How could you make that specific piece efficient enough that you stop skipping it?

Glossary

  • Decision rationale. The documented reasons why a specific hiring decision, whether advance, reject, or offer, was made.
  • Adverse impact. A hiring practice or decision that disproportionately affects members of a protected group. Good documentation is how you demonstrate that decisions were made on legitimate, job-related criteria.
  • ATS, applicant tracking system. Software that manages recruiting from job posting through hire, and which should capture and retain hiring documentation including AI-specific fields.
  • Defensible decision. A hiring decision supported by documented evidence showing that criteria were applied fairly and consistently.

Closing

Documentation is the infrastructure of fair, defensible hiring. It looks like overhead when you are busy and it is actually the foundation that lets you be confident in your decisions and improve on them. The investment is front-loaded and the return arrives later, which is exactly why it gets skipped, and exactly why the teams that build it into the workflow are the ones that still have it when they need it. Omar's team now captures the criteria before screening starts, the tool version with the score, a named reviewer on every AI-flagged rejection, and a rationale tied to a posted requirement. None of it takes long, and all of it would have answered the EEOC's request in a week rather than a quarter.

Key Takeaways

  • Documentation serves two purposes. Compliance, as evidence of fair practices if you are questioned, and improvement, as the only way to learn whether your decisions were consistent.
  • The record is the defense when AI is involved. An algorithm cannot be deposed, so the only account of why a tool ranked a candidate is the one you captured at the time. The EEOC holds employers responsible for the outcomes of tools they deploy, including vendors' tools.
  • Capture the essentials for every candidate, and six extra fields when AI is involved. Identity and source, the resume as submitted, assessment and AI scores, the criteria written before screening, screener notes, interview feedback by dimension, and the rationale. For AI-assisted decisions: the job-related criteria, the tool and its version, the raw output, the named reviewer, the rationale, and the date.
  • Match detail to risk, and prefer structure. Brief notes for clear-cut calls, thorough documentation for near-boundary, offer-level, and challengeable decisions. Structured forms and scorecards beat free-form notes because they force consistency and enable analysis.
  • Document what actually drove the decision. A record that asserts objective criteria while the real basis was subjective is the most dangerous kind, because it contradicts the decision-making the moment it is examined. Honest documentation is more defensible than tidy documentation.
  • Retain to the longest obligation, and freeze on a charge. Meet or exceed every federal, contractor, and state minimum, keep AI-specific fields for the full period, adopt a practical internal standard such as three years, and preserve all records relevant to a filed charge until it is finally resolved.
  • Build the trail into the workflow and name an owner. Automate capture, require a reviewer and rationale before an AI-flagged rejection finalizes, run selection-rate analysis on a schedule, and give one person accountability.

Frequently Asked Questions

How much documentation is enough for a routine rejection? One sentence tied to a stated criterion. The test is not length, it is whether someone unfamiliar with the candidate could understand the basis and verify it against the criteria you wrote down before screening. What is not acceptable is varying that depth candidate by candidate rather than by decision type, since uneven documentation across comparable candidates is what a discrimination claim is built from.

Our vendor says their tool is validated and bias-tested. Is that our documentation? No. It is a useful input to your due diligence and worth keeping on file, but the EEOC has been clear that employers remain responsible for the outcomes of tools they deploy, including tools built by vendors. If the tool produces an adverse impact, the vendor's assurance is not a defense. You still need your own per-candidate records: criteria, version, output, reviewer, rationale, and date.

How long should we keep records? Long enough to satisfy every obligation that applies to you, which means checking your jurisdiction and your federal contractor status rather than relying on one remembered number. The federal baseline for most private employers is one year under EEOC recordkeeping regulations, with a similar minimum for certain records under the ADEA, and many teams set an internal standard of three years to stay comfortably clear of longer state and contractor requirements. If a charge is filed, preserve everything relevant to it until the matter is finally resolved, regardless of the normal clock.