Documentation Standards: What to Document and How Much Detail
Bianca runs talent acquisition for a 1,200-person regional health system in Ohio, with a team of nine recruiters who fill roughly 600 roles a year. Eight months ago her team adopted an AI resume-screening assistant to triage the 40,000 applications they receive annually. Then a candidate filed a complaint alleging the screening had unfairly rejected her, and Bianca pulled the file to see how the decision had been made. The note in the applicant tracking system read, in full: "Not a fit." No mention of the tool, no rationale tied to the job, no record of who reviewed the AI's recommendation. Bianca had no way to defend a decision her team had made eight months earlier, and no way to prove a human had even looked at it. That single empty field is what this lesson is about: not how to write one record, but what standard the whole function should be held to, and how much detail is enough.
The Just-Enough Principle
There are two ways to get documentation wrong, and they point in opposite directions. Under-document, and you cannot defend a decision when someone questions it. Bianca learned this the hard way: a "Not a fit" note tells a regulator, a plaintiff's attorney, or even her own team absolutely nothing about why a qualified-looking candidate was screened out. Over-document, and you create a different problem in two forms. Every speculative aside and offhand impression becomes a record that lives in the file forever, discoverable in litigation, and capable of making a perfectly lawful decision look biased. And a standard so heavy that nobody can meet it collapses into no standard at all, which puts you back where you started.
The just-enough principle resolves both failures. Document the job-related rationale for the decision, the AI tool and version that assisted it, the human who reviewed and any override they made, and the date. That is the load-bearing core. Everything beyond that should earn its place. If a note does not help you explain why this decision was job-related, and it is not required by law, ask whether writing it down creates more risk than it resolves. Good documentation is not the longest documentation. It is the documentation that lets you reconstruct and defend a decision without recording things you wish you had not, at a level of effort your team can sustain across every requisition in a busy quarter.
The Four Fields That Always Belong
For any AI-assisted recruiting decision, Bianca's team now captures four things, regardless of stakes. First, the job-related rationale: the specific, role-relevant reason for the outcome, tied to a criterion in the job posting. "Does not meet the posted minimum of three years of acute-care nursing experience" is job-related. "Felt junior" is not. Second, the AI tool and version: which system assisted the decision and which version or model, because an audit a year from now needs to know what was actually running. Third, the human reviewer and any override: who looked at the AI output, whether they accepted or overrode it, and the reason for an override. Fourth, the date.
These four fields do double duty. They are what makes a decision defensible under Title VII and the EEOC's uniform guidelines, which expect employers to articulate a job-related, business-necessity rationale for selection decisions that produce adverse impact. And the human-reviewer field is what keeps AI from becoming an unaccountable decision-maker. A recorded human review with a named reviewer is the difference between "our tool rejected her" and "our recruiter, having reviewed the tool's recommendation against the job requirements, declined to advance her for this reason." One of those is defensible. The other is a headline.
Scaling Detail to the Decision
Just-enough is not a fixed length. The right amount of detail scales with how consequential and how contestable the decision is, and the cleanest way to make a standard usable is to write that scale down explicitly rather than leaving each recruiter to guess. Bianca's team built a four-tier matrix mapping decision type to documentation requirement, so that nobody has to decide in the moment how much effort a given call deserves.
| Decision type | Detail level | What it looks like | Why |
|---|---|---|---|
| Routine rejection, screening out an obvious non-fit | Brief | A single field or one-line note: "Does not meet minimum experience requirement; requires 5 years, resume shows 2." | The criterion is objective and the decision is clear-cut, so extensive writing adds nothing a reader could not already verify. |
| Borderline rejection | More detailed | Two or three sentences naming the borderline nature and the rationale: "Candidate has 4.5 years of relevant experience against a 5-year requirement, with a strong learning trajectory and related skills. Chose not to advance given the current pipeline of candidates who meet the exact requirement and screening time constraints." | The call is close, so you need to be able to explain why you declined. If the pipeline changes or the decision is questioned, you can point to your reasoning. |
| Interview-stage decision, advance or reject after interview | Comprehensive | A structured rubric with multiple dimensions and a justification for each: technical skills 8 of 10, strong understanding of architecture; communication 7 of 10, clear explanation but slower to articulate complex ideas; collaboration 6 of 10, asked questions but did not actively build on others' ideas. Overall: strong technical fit, some concern about communication in roles requiring real-time technical discussion. Decision: advance to final interview with a focus on collaboration. | High-stakes calls involve judgment, and multiple people are evaluating, so you want each assessment captured rather than averaged into an impression. |
| Offer or rejection after final interview | Very comprehensive | A full narrative from each interviewer, a comparison against the other candidates who reached this stage, the business rationale for the final call, and any concerns raised along with why they were outweighed. | The highest-stakes decision, the most likely to be questioned, and the point of final commitment. |
The matrix is what turns a principle into a standard. A recruiter facing a routine screen-out does not have to wonder whether they are under-documenting, because the tier tells them. A panel wrapping up a final-round debrief knows the bar is higher and why. Crucially, the depth is set by the type of decision and never by the individual candidate, which is the single most important fairness property of the whole scheme. Depth that varies with how interesting the reviewer found someone is exactly the asymmetry a discrimination claim is built from.
The same logic extends up the funnel when AI is in the loop. A routine screen-out against an objective, clearly posted minimum needs little more than the four core fields and one job-related sentence, because the criterion is unambiguous and the tool's role is narrow. When several interviewers assess a finalist and AI summarizes their notes, capture each evaluator's job-related assessment against the same dimensions, the AI's role in synthesizing them, the human who reviewed that synthesis, and the business rationale for the final call. The more people involved and the closer to a final commitment, the more the record needs to show that humans, not the tool, owned the outcome.
What Goes Inside the Record
Format aside, six elements carry the substance, and a standard that names them gives recruiters something concrete to write against. Criteria evaluated states which dimensions of the candidate were assessed at this stage, drawn from requirements written down before screening began. Rating or assessment records where the candidate landed on each criterion, whether on a verbal scale of strong, adequate and weak or on a numeric one. Evidence is the specific example from the interview or resume that supports the rating: "when asked about system design, the candidate articulated clear trade-offs and scaling considerations." Reasoning connects assessment to outcome, explaining why this pattern of ratings produces this decision for this role.
Two more elements apply at the higher tiers. Comparables matter for high-stakes decisions, because a final-round call is inherently relative: "stronger technical skills than candidate B, similar communication ability, weaker leadership presence" tells a later reader what the choice actually turned on. And concerns and mitigations belong on any advance made despite a reservation, naming what you plan to assess further: "concern about communication; advancing to the final interview to observe in a group setting." That last element is quietly one of the most valuable, because it converts a vague hesitation into a testable question and gives the next interviewer something specific to do.
Choosing a Format
The elements above can be captured several ways, and the right choice depends on your team's size, discipline, and tooling rather than on which format is theoretically best. Each has a real cost.
| Format | What it is | Strengths | Costs |
|---|---|---|---|
| Structured templates | Forms with named fields: criteria, assessment, reasoning. | Consistency by construction, easy to audit, standardized across evaluators. | Can feel rigid, may not capture nuance, can slow decision-makers down. |
| Open-form notes with structure | Narrative notes that follow a consistent shape every time. | Flexible enough to capture nuance, faster for experienced evaluators. | Requires discipline from every user, harder to audit for consistency. |
| Scored rubrics | Quantified assessment with a brief justification per score. | Easy to compare across candidates, easy to analyze for patterns. | Can miss nuance and create false precision; needs clear definitions of what each score means. |
| Audio recording | The screener records their assessment verbally, transcribed later. | Faster than writing, captures the texture of the reasoning. | Needs a transcription system, harder to search, storage-heavy, slower for asynchronous review. |
| Hybrid | Structured fields for the basics, open text for nuance, optional scoring. | Consistency on the key information plus flexibility where it matters. | Requires clear guidance on what goes where, or people put things in the wrong place. |
Whatever you choose, the non-negotiable is that everyone uses it the same way. Inconsistent documentation defeats the entire purpose, because the value of a standard comes from being able to compare decisions across evaluators to spot bias and verify fairness. Two recruiters writing to two different shapes produce records that cannot be read against each other, which means the fairness analysis you were building the standard to enable is not available to you.
What to Leave Out
Just as important as what to capture is what to keep out of the record, and this is where the over-documentation failure does its real damage. Three categories create liability without adding defensibility, and they should be trained out of your team's habits explicitly.
The first is protected-class information and anything that proxies for it. Notes about a candidate's age, the year they graduated, their apparent pregnancy, their accent, their national origin, their religious observance, or a disability have no place in a screening note unless they are a bona fide occupational qualification, which they almost never are. Even when the underlying decision is lawful, a note referencing a protected characteristic hands a complainant the appearance of bias. The second is speculative or "vibe" notes: "seemed nervous," "not sure she'd fit the culture," "gut says no." These feel honest in the moment, but they are exactly the kind of subjective, unmoored impression that plaintiffs use to argue a decision was pretextual. If an impression matters, translate it into a job-related observation. "Could not describe a specific example of leading an incident response when asked" is defensible. "Seemed timid" is not. The third is editorializing about the AI itself: "the tool is probably wrong but I'll go with it" is a sentence you never want read aloud in a deposition.
The goal here is not to sanitize records into dishonesty, and the distinction is worth being precise about because teams get it wrong in both directions. Documenting the real, job-related reason for a decision is essential, and a false record is far riskier than an honest one. The goal is to record the job-related substance of the decision and to stop recording the parts that are neither job-related nor required.
A Tale of Two Records
Consider a single decision on Bianca's team: a recruiter, assisted by the AI screener, declines to advance a candidate named in the file as J. Okafor for a senior clinical-systems analyst role. Here is the same decision documented three ways.
The too-thin version, the one that sank Bianca eight months ago, reads: "Not a fit. Screened out." Twelve characters of rationale. There is no job-related reason, no record of the tool, no named human reviewer, no override note, no date beyond the system timestamp. If this candidate complains, Bianca cannot show the decision was job-related, cannot show a human exercised judgment, and cannot even confirm which version of the tool ran. This is the under-documentation failure: the decision may have been perfectly lawful, but it is indefensible because nothing was preserved.
The right-sized version reads: "Declined to advance for senior clinical-systems analyst (req #4417). Job-related rationale: posting requires five-plus years integrating EHR systems; resume and screening responses show two years, limited to a single read-only reporting module, no integration work. AI screener v3.2 flagged below experience threshold; recommendation reviewed and accepted by recruiter D. Reyes, who confirmed the gap against the posted requirement. Date: 2026-03-14." That is roughly seventy words. It names the job-related criterion, ties it to the posting, identifies the tool and version, records the named human reviewer and that she accepted the recommendation, and stamps the date. It says nothing about the candidate's age, name-implied background, demeanor, or anything speculative. If a complaint arrives, this record defends itself.
Now consider the over-documented version, the failure in the other direction: "Declined. AI flagged low experience. Honestly the tool seems aggressive lately and I'm not sure I trust it. Candidate also came across as much older than our usual hires and I worried about long-term fit with our young team, plus she mentioned needing time off for a religious holiday during onboarding. Two years experience though, so that's the official reason." This version contains a defensible fact buried in a minefield: it impugns the tool, references age and religion, and admits the stated reason is a cover for impressions. It is worse than the thin version, because the thin version is merely empty while this one is actively incriminating. The right-sized record is not a compromise between these two. It is a different discipline entirely: complete on the job-related substance, silent on everything else.
The Legal Floor Under Your Standard
Some documentation is not optional, and the clearest example is New York City Local Law 144, in effect since July 2023. If you use an automated employment decision tool to substantially assist or replace discretionary hiring or promotion decisions for a role located in New York City, the law requires that the tool undergo an independent bias audit within the prior year, that a summary of the most recent audit results be published, and that candidates receive notice at least ten business days before the tool is used. The records that prove you met these obligations are mandatory: which tool you used, the audit date and summary, and evidence that notice was given.
This matters for Bianca's Ohio health system the moment it posts a role based in its New York City satellite clinic, because Local Law 144 follows the job location, not the employer's headquarters. The lesson here is that just-enough has a floor set by law. You can exercise judgment about how much rationale to write for a routine screen-out, but you cannot exercise judgment about whether to keep the bias-audit summary or the candidate-notice record for a covered tool. Build those required records into your system so they are captured automatically, and treat them as non-negotiable. Everywhere else, the EEOC's general expectation that selection decisions be job-related and defensible is what your four core fields are designed to satisfy.
Designing a Standard People Actually Follow
A standard that nobody follows protects no one, and the fastest way to guarantee non-compliance is to make documentation burdensome. If Bianca required a 500-word narrative for every screen-out, her recruiters would either rebel or quietly write nothing, leaving her back at "Not a fit." Five things separate standards that survive contact with a busy quarter from standards that quietly die.
Clarity. Make the expectation unambiguous and illustrate it with examples rather than assuming people will interpret an abstract standard the way you intended. Most inconsistency in documentation is not defiance; it is people guessing differently. Efficiency. Build the four core fields into the applicant tracking system as structured prompts so the rationale, the tool and version, the reviewer, and the date are captured at the moment of decision rather than reconstructed later. Pre-populate the tool and version automatically. Give recruiters a short library of job-related rationale phrasings so they are translating impressions into defensible language by default. Accountability. Track whether documentation is actually being completed and address it when it is not, because a standard that some evaluators meet and others ignore produces exactly the uneven record that is hardest to defend. Training. Walk new team members through the standard with worked examples instead of assuming they will infer it from the template. Feedback. Show the team how the documentation gets used: "we reviewed last quarter's records to check for fairness and found this pattern, which is why this matters."
That last point is the one most teams skip, and it is the one that determines whether the standard is alive in a year. Documentation that is never reviewed degrades into theater, and a team that senses its notes go nowhere will stop investing care in them. Bianca's team now runs a quarterly review: they sample recent AI-assisted decisions, check that the four fields are present and job-related, calibrate by asking whether different evaluators applied the same standards to comparable evidence, and look for patterns that might signal adverse impact across groups. When recruiters see that the documentation is read, calibrated against, and used to catch problems before they become complaints, the standard stops being a chore and becomes the thing that protects them.
What a Good Standard Buys You
It is worth being explicit about the return, because the cost of documentation is felt immediately and the benefit arrives later. A standard that is followed lets you do five things you otherwise cannot. You can defend decisions when they are questioned, which is the compliance case. You can identify patterns and potential bias across a body of decisions, which is the fairness case. You can calibrate assessments across evaluators, because comparable records are what make it visible that one interviewer's 8 is another's 6. You can learn from experience, checking whether the criteria you weighted actually predicted who succeeded in the role. And you can revisit decisions when circumstances change, which is the mundane benefit teams appreciate most: when the pipeline thins in month three, a borderline rejection with two sentences of reasoning is a candidate you can go back to, while "not a fit" is a dead end. The goal is never documentation for its own sake. It is documentation that serves those five purposes at a cost your team will keep paying.
Anti-Patterns
Documentation that is defensive rather than honest. A recruiter records "candidate did not meet technical requirements" when the real reason was that the candidate seemed nervous and the recruiter doubted they would fit the culture. The note is a sanitized version that looks better and misses the actual thinking. It happens because honest documentation feels like it might invite disagreement or look biased, so the recruiter writes what seems safest. What goes wrong is that dishonest documentation is more legally risky than honest documentation, not less: if the record does not match the real reasons and that surfaces, it reads as concealment. The fix is to document what actually influenced the decision, including the uncomfortable parts, translated into job-related terms wherever possible. "Candidate demonstrated the required technical skills but appeared anxious in the interview; we chose to continue with other candidates and would re-evaluate if the pipeline thins" is both honest and defensible.
Documentation so detailed it never gets done. The standard says every rejection requires a 500-word explanation and a rubric with fifteen dimensions. Recruiters hate it. They document less than required, or they leave. The standard looks excellent on paper and does not survive a real week. It happens when requirements are designed for an ideal case rather than an actual workload. What goes wrong is that nobody follows the standard, documentation becomes inconsistent or disappears entirely, and the cure turns out worse than the disease. The fix is to make standards realistic, because a modest standard everyone follows beats a perfect standard nobody does.
Documentation with no follow-up. Screeners document their assessments as required. The records sit in the applicant tracking system. Nobody ever reads them, analyzes them, or acts on them. It happens because there is no process for review, only for capture. What goes wrong is that documentation becomes theater; once people sense their notes go nowhere, the quality degrades and the practice erodes. The fix is to schedule the review, monthly or quarterly, and to use it for something visible: calibration across evaluators, fairness analysis across groups, and improvements to training and process. Documentation that informs real decisions is documentation people take seriously.
Practice Prompts
- Build the decision-type matrix. Write your own version of the tier table: which decision types your team makes, and what documentation each requires. Be specific about the minimum for a routine screen-out and the expectation at offer stage.
- Create three templates. Draft templates for a screening rejection, an interview rejection, and an offer. Make them simple enough that your team will actually use them, then test with a handful of recruiters and see which fields they skip.
- Audit a sample. Pull twenty recent decisions and assess the records for consistency, sufficiency for compliance, and usefulness for improvement. Where records are thin, note whether the gap tracks the decision type or the individual reviewer.
- Design the review process. Decide how documentation will be used beyond storage: who reviews it, on what cadence, and what they look for on calibration, fairness, and improvement. Put the first session on the calendar.
- Pressure-test with the team. Walk your draft standard through with the people who have to follow it and ask specifically what they would skip under time pressure. Adjust before finalizing rather than after adoption fails.
Reflection
- What is your organization's current documentation requirement, and does anyone actually follow it? Is it useful for both compliance and improvement, or only nominally in place?
- What documentation do you most regret not having from a past hiring decision, particularly one that was later questioned?
- If you could implement exactly one documentation standard tomorrow, what would it be, and what specific problem would it solve?
- How would you motivate your team to document consistently? Which incentives or system changes would do more than exhortation?
- How would you actually use the documentation you collect: what would you review, how often, and what would you change as a result?
Glossary
- Documentation standard. An established expectation for what gets documented, at what level of detail, and in what format.
- Calibration. Using documented decisions to align standards across decision-makers, checking whether different evaluators apply the same criteria to comparable evidence.
- Proportionality. The principle that documentation depth should scale with the stakes and contestability of the decision, set by decision type rather than by candidate.
- Automated employment decision tool. A tool that substantially assists or replaces discretionary hiring or promotion decisions, and the category NYC Local Law 144 regulates.
Related Lessons
This lesson sets the standard across the function. These lessons cover the pieces it governs.
- Documenting Decisions: Clear Records for Legal and Fairness Review works through a single record in depth: how to phrase a rationale so it survives a hostile reader, and what a defensible artifact contains.
- Decision Logging: Recording Human Decisions, AI Input, and Reasoning covers the per-decision log entry, and why the AI's input and the human's decision belong in separate fields.
- Documentation and Evidence: Building a Trail for Compliance covers the surrounding evidence trail, including retention periods, filing, and answering a regulator's request.
- Compliance and Legal Review: Documentation for FCRA, EEO, and GDPR maps the specific regulatory expectations your standard has to satisfy in each regime.
- Auditing AI-Assisted Decisions: Sampling Methodology and Fairness Metrics covers the quarterly review your records make possible, and how to sample them properly.
Closing
Good documentation standards are the foundation of defensible, fair, improvable recruiting, and they are ultimately a design problem rather than a discipline problem. The temptation is to solve documentation by asking for more of it, and that reliably produces less. What works is deciding in advance how much each type of decision warrants, capturing the four load-bearing fields at the moment of decision rather than afterward, keeping out everything that is neither job-related nor required, and then actually reading what you collect. Bianca's team did not become more thorough than the average recruiting function; they became more deliberate about where thoroughness was worth spending. The result is a standard her recruiters can meet in a bad week, records that would have answered the complaint that started all this, and a quarterly review that finds problems while they are still cheap to fix.
Key Takeaways
- Document just enough to defend the decision, not everything you observed. Under-documentation leaves you unable to explain a decision when challenged; over-documentation creates a permanent record of speculative or protected-class notes and a standard nobody can sustain. The target is the job-related substance and nothing that is neither job-related nor required.
- Four fields always belong: rationale, tool and version, human reviewer and override, and date. The job-related rationale ties the outcome to a posted criterion, the tool and version make the decision auditable, and the named human reviewer is what keeps AI from becoming an unaccountable decision-maker.
- Set detail by decision type, in writing. Routine screen-outs get a line, borderline calls get a short narrative, interview decisions get a structured rubric, and offer-stage decisions get the full picture including comparables. Depth must track the kind of decision and never the individual candidate.
- Pick a format and enforce it uniformly. Structured templates, structured narrative, rubrics, audio, and hybrids all work; inconsistency does not, because comparing decisions across evaluators is the whole point.
- Keep protected-class details and vibe notes out of the record. Age, graduation year, religion, national origin, disability, and impressions like "seemed nervous" add liability without adding defensibility. Translate any impression that matters into a job-related observation.
- Honest beats sanitized, but job-related beats both. A record that states a formal criterion while the real basis was subjective is more dangerous than an honest one, because it contradicts the decision-making the moment it is examined.
- Some records are mandatory under NYC Local Law 144. For automated tools assisting covered NYC hiring or promotion decisions, the bias-audit summary, the audit date, and proof of candidate notice are required, and the obligation follows the job's location rather than the employer's.
- Make the standard easy, then actually use it. Clarity, efficiency, accountability, training, and visible feedback are what make a standard survive; a quarterly review for completeness, calibration, and adverse impact is what turns the records from theater into protection.
Frequently Asked Questions
How do we set the minimum without inviting the whole team to document less? Write the tiers down and tie each one to a decision type rather than to effort. The minimum for a routine screen-out is a line because the criterion is objective, not because screen-outs matter less, and the standard should say so. Once the tiers are explicit, "less" is not a judgment call a recruiter makes under pressure; it is what the standard already specifies, and anything below it is a gap you can see in an audit.
Is a scored rubric better than narrative notes? Neither is better in the abstract. Rubrics compare cleanly across candidates and analyze well, which makes them strong at interview stage where several evaluators assess the same dimensions. They also create false precision if the scale levels are not defined, and they miss reasoning that does not fit a dimension. Narrative captures nuance and is faster for experienced evaluators, but it only supports analysis if everyone writes to the same shape. Many teams end up hybrid: fixed fields for the load-bearing facts, short narrative for the reasoning.
Our recruiters say the standard slows them down. Do we relax it? Relax the format before you relax the substance. If the four core fields are taking real time, the problem is usually capture rather than content: the tool and version should be pre-populated, the reviewer should come from the logged-in user, and the rationale should have a phrasing library behind it. If the substance genuinely cannot be met in a busy quarter, the standard is too heavy and will fail, so cut it to what the team will sustain and enforce that consistently. A modest standard everyone meets is worth more than an ambitious one that quietly turns into "not a fit."
Skill.re