Documenting Decisions: Clear Records for Legal and Fairness Review
Tara is a TA partner at a 240-person regional hospital network, and she leads a recruiting pod of four people filling roughly 60 clinical and administrative roles a quarter. Her team adopted an AI resume-screening assistant last year, and most days it saves real time. Then a rejected nurse-manager applicant filed a complaint alleging the screening process disadvantaged candidates over 50. Tara's general counsel asked one question: "Show me the records for those decisions." What came back was a mix of thorough notes, two-word shrugs, and several decisions where nobody could say whether a human had reviewed the AI's output at all. That gap, not the AI itself, was the problem. This lesson is about the records Tara wished she had built from day one: factual, job-related, and detailed enough to survive both a legal challenge and an honest fairness audit.
Why the Records, Not the Decision, Decide the Case
Documentation never feels urgent until the moment you need it. Picture the ordinary version of the problem. You are in an interview debrief with a hiring panel. Three people voice strong opinions about the same candidate: one says outstanding, another says not quite ready, a third says interesting but uncertain. How do you reconcile those perspectives, and more to the point, what actually gets written down? Six months later, if that candidate alleges discrimination, the panel's impressions are gone and only the file remains. What the file says about job-relevant factors is the entire substance of your position.
Under the framework U.S. courts apply to employment discrimination claims, a candidate first has to raise an inference that a protected characteristic influenced the outcome. The burden then shifts to the employer to articulate a legitimate, non-discriminatory, job-related reason for the decision, and the candidate can still try to show that reason is a pretext. At every step, the employer's evidence is its contemporaneous documentation. A defensible hiring decision is not the one that was secretly well-reasoned in someone's head; it is the one whose reasoning was written down at the time, in job-related terms, and can be produced months later. A well-documented decision showing specific, job-relevant reasons with supporting data is often the difference between defending successfully and losing an expensive case.
Legal defense is only half of what the record buys you; the other half is fairness. Without documentation you cannot audit, because a status field reading "rejected" reports the outcome and nothing about the reasoning. Teams that document minimally give up their legal protection and their ability to improve in the same stroke, and both are served by the same discipline.
Tara's network now sits in two regulatory worlds at once. The federal standard rewards records that demonstrate job-relatedness and consistency. And because the network recruits for several roles based in New York City, it falls under NYC Local Law 144, which governs automated employment decision tools: it requires bias auditing of those tools, candidate notice that an automated tool is being used, and disclosure of the job qualifications the tool assesses. Local Law 144 does not tell Tara what to write in a single candidate's file, but it raises the baseline expectation that when an AI tool touches a decision, the organization can account for what the tool did and how a human used it. Good per-decision records are how that accountability actually shows up in practice.
The Decision Points Worth Documenting
Before you can write a good record you have to know where records are owed. The points worth documenting are the ones where bias could plausibly influence an outcome and where you might later need to defend a call. In most funnels that is six places: the initial resume screen, where the question is why a candidate was rejected or advanced; the phone screen, where the question is what you learned about job match; the technical or skills assessment, where the question is which competencies were assessed and how the candidate performed; individual interview feedback, where the question is what behavior or skill was observed; the panel debrief, where the question is what consensus emerged and on what basis; and the final hire or rejection, where the question is the rationale for the commitment.
Every one of those moments takes the same shape, which is what makes the practice teachable rather than a matter of individual style. Start with the decision itself: was the candidate advanced, held, or rejected? Then state the rationale in specific, job-relevant terms. Then attach the supporting data, meaning assessment scores, competency ratings, and interview feedback from the people who did the evaluating. Then close with a consistency note, the field most teams skip and the one that does the most work: how does this decision align with the decisions you made about similar candidates? Four elements, same order, every time. The repetition is what makes a hundred records comparable to each other, and comparability is the precondition for both audit and defense.
The Five Fields Every AI-Touched Record Needs
When an AI tool participates in the decision, the four-part shape above still applies, but Tara's team layers five specific fields on top of it. Miss one, and the record develops the same blind spot that exposed her during the complaint.
The job-related rationale. The specific, role-tied reason for the outcome, stated in terms of competencies the job actually requires. Not "strong candidate" but "meets the posted requirement of three years of charge-nurse experience and ACLS certification; advanced to phone screen." The rationale must trace back to the published job requirements, because that is exactly the link a fairness review and a legal defense both depend on.
The AI tool and version. Which tool produced the recommendation and which version. "Screening assistant v3.2, run on 2026-03-14" is a record. "The AI suggested" is not. Versions matter because model and configuration changes alter behavior, and a fairness audit that cannot tell which version scored which candidate cannot interpret its own results. This is where Local Law 144's expectation of accountable tooling becomes concrete in the file.
The human reviewer. The named person who reviewed the AI output and owns the decision. A record showing that a tool ranked a candidate but never showing a human reading that output is the worst of both worlds: it looks automated and unaccountable. Naming the reviewer establishes that a person, not a model, made the call.
The override decision. Whether the human agreed with the AI's recommendation or departed from it, and why. Overrides are where the real reasoning lives. When Tara's recruiter advanced a candidate the tool ranked low, the note explaining why is the most valuable line in the file, because it proves human judgment was applied to job-related facts rather than rubber-stamped from a score.
The date. When the decision was recorded. Contemporaneous records carry weight; records reconstructed weeks later read as rationalization. A timestamp landing within a day of the decision is part of what makes the rest of the record credible.
A Worked Record: One Decision, Done Two Ways
Tara's team screened 47 applicants for a charge-nurse opening. One of them, an applicant the network will call Applicant 19, became the kind of decision that gets scrutinized later. The screening assistant scored Applicant 19 at 62 out of 100 and tagged the file "below threshold (70)." The recruiter, Dana, read the full resume and disagreed: the score was dragged down by a two-year employment gap the tool treated as a negative signal, but the gap was a documented family-leave period, and the applicant met every posted requirement. Dana advanced the candidate anyway.
Here is the liability-creating version, the kind Tara found too often during the complaint review: "Applicant 19 - advanced. Seemed solid despite low AI score. Good background." Three sentences, zero defensible content. It names no requirement met, records no tool or version, names no reviewer, explains no override in job-related terms, and carries no date. If a fairness audit later finds older applicants were disproportionately scored below threshold by v3.2, this record cannot show whether Dana's override was a principled correction or a one-off favor.
Here is the defensible version: "Date: 2026-03-14. Tool: Screening Assistant v3.2, score 62/100, flag 'below threshold (70).' Reviewer: Dana R. Decision: advance to phone screen (override of tool recommendation). Job-related rationale: candidate meets all posted requirements, with 4 years charge-nurse experience against a 3-year requirement, current ACLS and BLS certification, and level-1 trauma background. Override basis: tool's low score driven by a 24-month employment gap; gap is documented family leave and is not job-related. Consistency: aligns with prior overrides where the tool penalized non-job-related gaps." Every field is present, every claim ties to a posted requirement, and the override is reasoned rather than asserted. If Applicant 19 is later compared against a rejected applicant with a similar profile, this record can defend the difference or expose an inconsistency worth fixing.
Specificity Is the Whole Craft
The difference between a record that works and one that does not is almost always specificity. "The candidate was good" is not documentation; it is a feeling with punctuation. "Demonstrated the required communication competency at 3.5 out of 5, evidenced by clear explanation of design decisions but a lack of proactive clarification questions in a complex technical exchange" is documentation, because a stranger can read it, understand what was assessed, and check it against what other candidates were assessed on. Specificity matters legally, because it is what an articulated non-discriminatory reason looks like, and analytically, because vague records aggregate into vague data no audit can interpret.
Consider a second worked case, this one from a technical pipeline rather than a clinical one. Sarah is a mid-level software engineer candidate. She completed the coding assessment with a score of 78 out of 100, placing her at the 68th percentile for the role. Two engineers rated her problem-solving at 4 out of 5, noting solid algorithm thinking with some inefficiency in implementation. Her system design was rated 3.5 out of 5: she understands distributed systems concepts but misses critical reliability trade-offs. Her communication in a technical setting was rated 3 out of 5, with clear explanation but a need to articulate design decisions more explicitly.
Written up properly, that decision reads: "Decision: hold, advance only if stronger technical candidates are unavailable. Rationale: technical assessment score of 78 percent is adequate but not exceptional for senior engineer placement. Interview feedback shows solid problem-solving and system understanding but gaps in efficiency and reliability considerations. This candidate would succeed as a mid-level engineer, but the role requires senior-level optimization thinking. Communication is adequate. Consistency: this holds the pattern for candidates with assessment scores in the 75 to 80 percent band, who are typically held and advanced only if the pipeline is weak." That record states the outcome, ties the reasoning to the level of the role rather than to an impression, carries the supporting numbers, and locates the decision inside a pattern that can be checked. A reviewer who has never met Sarah can evaluate whether she was treated the way comparable candidates were.
Getting that quality out of a whole team requires a standardized format rather than individual diligence. A shared format ensures everyone captures the same information, makes audit tractable, and reduces the appearance of inconsistency if a challenge arrives. Use rating scales consistently, with a common definition of each level, so a 4 from one interviewer means roughly what a 4 from another does. Then replace the blank text box with a form that asks whether the required competency was demonstrated, requests a rating on that scale, and requires behavioral evidence for it. Templates lower the cognitive load at the end of a long interview day and produce records that are consistent by construction rather than by willpower.
Writing Factual Records Instead of Liability
The line between a record that protects Tara and one that sinks her is usually the line between observation and impression. Factual records describe job-related behavior and measurable qualifications. Liability records capture impressions, demographic observations, and speculation about a candidate's personal life. "Candidate could not describe a time she escalated a patient-safety concern, a required competency for the charge role" is factual. "Didn't seem like a leader" is an impression that, repeated across a demographic group, becomes a plaintiff's exhibit.
The categories that create liability are predictable, and they should be trained out explicitly rather than left to instinct. Personal impressions unrelated to job requirements belong nowhere in a file: "seemed really nice," "reminded me of my brother," "doesn't seem like a fit," "great energy." Demographic observations of any kind are worse: age, family status, accent, appearance, or any variation on where someone "looks like" they belong. Speculation about personal situations is equally dangerous: "probably has family obligations," "might be job-hopping," "likely to retire in a few years." And vague culture-fit language hides bias behind a friendly word.
These notes are costly because they supply the factual basis for a claim you would otherwise be able to rebut. If your hiring data shows you advance women at a lower rate than men, and your files are dotted with "doesn't seem like a fit" or "not our culture," you have handed a complainant the connection between the disparity and a non-job-related standard. Observations that felt innocent when written become problematic once a pattern emerges across demographic groups, because a pattern is exactly what an adverse-impact analysis looks for.
The practical fix is a translation habit: every time a recruiter reaches for a soft phrase, they answer one question instead. Instead of "doesn't seem like a fit," ask what specific job skill is missing and write that. Instead of "great cultural fit," ask what job-relevant behavior demonstrated alignment with the role's requirements; "proactively proposed a shift-handoff checklist during the scenario exercise" survives review, while "culture fit" does not. Instead of demographic speculation, ask what work experience or skill the observation relates to, and if the honest answer is none, it does not go in the file.
This matters even more once an AI tool is in the loop, because the tool's outputs become part of the record. If the assistant surfaces a rationale that drifts toward non-job-related territory, the reviewer's job is to correct it before it is saved, not paste it through. A record is not improved by being machine-generated; it is improved by being job-related and true.
Using AI to Draft the Record, Not to Decide It
Tara's team uses the AI assistant to draft documentation, which solves the most common failure mode: the "I will write it up later" gap that produces post-hoc, fuzzy records. After an interview, a recruiter feeds the assistant the posted requirements and their raw notes and asks for a structured draft. The draft enforces structure and forces specificity, which is most of the battle for an interviewer who knows what they saw but not how to phrase it defensibly.
The prompt is short and mostly context. "We interviewed a candidate for a backend engineer role. The role requires distributed systems thinking, code quality standards, and technical mentorship ability. Interviewer feedback: strong algorithm background; sometimes writes inefficient code; good at explaining thinking; didn't ask good questions about requirements. Generate structured decision documentation with job-relevant rationale and suggested competency ratings." What comes back is a scaffold: algorithm and problem-solving at 4 out of 5, strong foundational algorithm thinking evidenced by rapid pattern recognition, slightly below expectations on implementation efficiency; code quality at 3 out of 5, functional code but missed opportunities to discuss design patterns or maintainability; technical communication at 4 out of 5, clear explanation but could proactively ask clarifying questions about requirements and constraints; technical mentorship readiness at 2.5 out of 5, with no evidence in the interview of thinking about knowledge transfer or team development.
That last line is the most useful thing in the draft. The assistant did not invent a mentorship assessment; it flagged that a required competency was never assessed at all. A human reading their own notes rarely notices the absence of something. A structured draft against the posted requirements makes gaps visible, which turns documentation from a record-keeping chore into a check on the interview itself.
The order of operations is the safeguard. The human decides; the AI helps write down a decision a human already owns. The recruiter reviews the draft against what actually happened, corrects anything the tool inferred rather than observed, names themselves as the reviewer, notes any override, and dates it. A record generated and saved without a named human reviewer reading it is precisely the record that fails both a Local Law 144 accountability question and an EEOC-style defense, because it cannot show that job-related human judgment governed the outcome. Tara's rule is blunt: no AI-drafted record enters the file until a named person has read it, corrected it, and signed it.
Building Documentation Into the Process
If documentation is friction, it will not happen, and no amount of exhortation changes that. Teams that document well have made it part of the workflow rather than a task waiting at the end of one. Interviewers complete structured feedback immediately after each interview, while the evidence is fresh and before the next conversation overwrites it. A named recorder captures the decision and the basis for consensus during the panel debrief. And the decision-maker reviews the file for completeness before an offer or rejection goes out, the last cheap moment to notice a required competency was never assessed or a reviewer never named.
Training is the other half. Show your team both kinds of record side by side and let the contrast teach. "Good candidate, hired" helps nobody, not in an audit, not in a defense, and not even the person who wrote it when they are asked about the decision a year later. The same call written properly reads: "demonstrated required technical competencies of algorithm thinking at 4 out of 5, distributed systems at 3.5 out of 5, and communication at 4 out of 5, all adequate to strong for a mid-level role; mentorship readiness was not assessed, a gap to cover in future interviews." Recruiters adopt the standard faster once they see the good version is not longer, just more specific, and that it protects them personally as much as the organization.
From Single Records to a Fairness Audit Trail
Tara insists on the same five fields for every decision because consistent records turn into an auditable trail almost for free. Once a quarter, her pod pulls the full set of screening decisions and looks for two things. First, aggregate patterns: are advancement rates noticeably different for any group, or for any characteristic the tool's scores correlate with, such as employment gaps? Second, paired consistency: did two applicants with materially similar job-related profiles receive different outcomes, and if so, do the records explain the difference in job-related terms?
The aggregate half is straightforward arithmetic and uncomfortable reading. Suppose your data shows you hire men at 35 percent and women at 25 percent. That gap may have a legitimate explanation: the candidate pools may differ in relevant ways, different roles may require different competencies, or the records may show job-relevant reasons that happen to distribute unevenly. But the explanation has to exist somewhere you can point to. If your documentation does not show different job-relevant criteria being applied, the rate difference is a finding rather than a coincidence, and it is one you would much rather make yourself.
The paired half catches problems the aggregate numbers hide. Suppose you hired candidate A, who had a technical assessment score of 75 and interview ratings of 3.5, 5 and 3.5, and rejected candidate B, who had a score of 76 and ratings of 3, 5 and 4. On the numbers alone, B looks at least as strong. Why did those decisions diverge? If the records show a job-relevant difference, the divergence is defensible. If they do not, you have found a consistency problem in a quarterly review rather than in a deposition.
When Tara ran her first real audit, the records did their job. The assistant's v3.2 scoring penalized employment gaps heavily enough that applicants with gaps advanced at a visibly lower rate, and many of those gaps were caregiving leaves. Because every override carried a job-related rationale, Tara could see her recruiters had been correcting the pattern by hand, override by override. That drove two fixes: a configuration change to stop treating documented gaps as a negative signal, recorded with a version note, and a standing instruction to flag gap-driven low scores for human review. Without per-decision records the pattern would have stayed invisible until a second complaint. With them, it became a maintenance task, which is the quiet payoff of disciplined documentation.
Anti-Patterns
The vague rationalization. Reasons written so generally that they provide no decision basis and no audit trail: "didn't feel like a fit," "good candidate," "not quite ready." It happens because vagueness is fast and translating an intuition into a specific observation is genuinely hard. What goes wrong is that the record provides zero insight in an audit or a challenge, and where disparities exist it invites the inference that something other than job-relatedness drove the calls. The fix is templates requiring a competency rating with behavioral evidence: "problem-solving competency at 3 out of 5, below the required 4, evidenced by difficulty decomposing a complex problem into sub-components; this gap led to the hold decision."
The demographic note. Recording demographic observations, personal impressions, or culture-fit assessments: "seems very young," "not sure she's committed to a career," "doesn't seem like our type of person," "great cultural fit because he's outdoorsy like us." It happens because these observations feel innocent in the moment. What goes wrong is that they create an explicit factual basis for a discrimination claim, in your own handwriting. The fix is to document job-relevant observations only, translating anything that sounds like fit into the behavior behind it: "demonstrated alignment with the required collaborative work style, proactively suggested team approaches to problems, and asked for feedback."
The post-hoc write-up. Writing the record weeks later, when memory has faded and the motivation is to justify a decision already made. It happens because the recruiting calendar is relentless and the decision already feels settled. What goes wrong is that the record becomes a reconstruction rather than an observation, and it reads that way; inconsistencies between what happened and what was written damage credibility on every other point in the file. The fix is timing: feedback forms completed immediately after the interview, debrief notes captured during the debrief, and the record finalized before the outcome is communicated.
The personality profile. Filling the record with observations about temperament rather than capability: "great attitude, energetic, seems motivated, friendly." It happens because those things are easy to notice and pleasant to write. What goes wrong is that none of it relates to job performance, so it displaces the assessment that would have been useful while adding subjective material a challenge can work with. The fix is to ask what job requirement each observation speaks to.
The uneven file. Documenting some decisions in detail and others minimally, or capturing communication for some candidates and only technical skills for others. It happens without anyone deciding to do it, because attention follows interest. What goes wrong is the worst inference available: that the well-documented candidates were the ones the reviewer liked, and if depth tracks a demographic line, that asymmetry is very hard to explain. The fix is a standard applied by decision type, so depth varies with the kind of decision and never with who the candidate is.
Practice Prompts
- Build a documentation template. Design a structured feedback form covering what must be captured at each decision point: the decision, competency ratings on a defined scale, behavioral evidence for each rating, the consistency note, and the job-relevant rationale. Pilot it on real interviews and refine it where people got stuck.
- Audit ten recent decisions. Pull ten recent hires or rejections and read only the documentation. Is it specific and job-relevant, or vague? Could someone unfamiliar with the candidate understand the decision? Mark the weakest records and look for what they have in common.
- Run a consistency check. Find pairs of candidates with similar backgrounds, scores, and feedback who received different decisions. Is the difference explained by a documented job-relevant factor, or does the pair suggest inconsistency worth investigating?
- Analyze rates by group. Calculate advancement rates by demographic group across your recent decisions. Where rates differ, compare the documented reasons given for the underrepresented group against those given for the overrepresented group. Are the differences job-relevant?
- Draft with AI, then compare. After your next interview, draft structured documentation from your raw notes and the posted requirements. Compare it to what you would have written unaided. Is it more specific? Does it surface a required competency you never assessed?
Reflection
- How much of your current documentation would hold up to legal scrutiny, and would it clearly show job-relevant reasons?
- What actually prevents your team from documenting thoroughly: time, uncertainty about what to write, or systems that make it awkward?
- Have you wanted to audit fairness but lacked the data? What would better records have let you check?
- Reading back through your own documentation, do you find patterns that surprise you or inconsistencies that concern you?
- How would your decisions change if you knew you would have to explain each one to a lawyer six months later?
Glossary
- Behavioral evidence. The specific observed behavior supporting a competency rating, for example: "asked clarifying questions about requirements before implementing a solution."
- Competency rating. A standardized assessment of whether a candidate demonstrates a required skill or behavior, usually on a numerical scale with a clear definition for each level.
- Decision documentation. The complete record of a hiring decision: the decision itself, the job-relevant rationale, the supporting data, the competency ratings, and the consistency note.
- Demographic parity. A condition in which hiring or advancement rates are similar across demographic groups, adjusted for legitimate differences in candidate qualifications.
- Fairness audit. A systematic review of hiring decisions and their documentation to identify patterns, inconsistencies, and possible bias.
- Job-relevant. Directly related to performing the job's requirements or predicting on-the-job success. Technical skills, required competencies, and relevant experience qualify; personality traits unrelated to the job, demographic characteristics, and personal situations do not.
- Post-hoc rationalization. Filling in reasoning after a decision has been made, rather than recording the reasoning that actually drove it.
- Structured feedback. A standardized form that requires specific information rather than free-form text, producing consistency and completeness by construction.
Related Lessons
This lesson is about the record of a single decision. The lessons around it cover the layers above and below that record.
- Decision Logging: Recording Human Decisions, AI Input, and Reasoning covers the log as a system, and why the AI's input and the human's decision belong in separate fields on every entry.
- Documentation Standards: What to Document and How Much Detail sets the function-wide standard: which decisions get documented at all, and how much detail each type warrants.
- Documentation and Evidence: Building a Trail for Compliance covers the surrounding trail, including retention periods, filing, and what a regulator asks for.
- Bias Audit Trails: Creating Evidence of Fairness Considerations turns pattern findings from your records into a documented cycle of investigation and correction.
- Structured Evaluation: Avoiding Halo Effects and Confirming Bias covers the assessment discipline that produces documentation worth writing down.
Closing
Documentation is not bureaucratic overhead; it is the foundation of fair, defensible recruiting, and it does several jobs at once. It creates a legal record showing decisions rested on job-relevant factors. It generates the data a fairness audit needs. It forces precision about what a candidate can and cannot do, which improves the decision itself and not merely the record of it. And it builds institutional knowledge, so a year of decisions shows which criteria predicted success and where bias had room to hide. The investment is front-loaded and the return arrives later, which is why it gets skipped, and exactly why the teams that build it in are the ones who still have it when they need it.
Key Takeaways
- The record, not the reasoning, is what survives review. Courts ask the employer to produce a legitimate, job-related reason, and the only evidence that counts is what was documented at the time. A well-reasoned decision with no contemporaneous record is, for legal and fairness purposes, an undocumented one.
- Documentation is legal protection and fairness data at once. The records that answer a discrimination claim are the same ones that let you audit patterns and improve. Teams that skip documentation give up both.
- Every AI-touched decision needs five fields. The job-related rationale, the AI tool and version, the named human reviewer, the override decision, and the date. Drop any one and the record develops a blind spot a complaint or audit will find.
- Specificity is what makes a record defensible and useful. "Good candidate" protects nobody and teaches nothing. A rating on a defined scale, the behavioral evidence behind it, and a note on how the decision compares to similar ones does both jobs.
- Write observations, never impressions. Demographic observations, personal speculation, and vague culture-fit phrasing create liability. Translate every soft impression into the job-related behavior actually observed, or leave it out.
- Immediate records are accurate records. Documentation written weeks later becomes rationalization rather than observation, and the gaps between it and what happened undermine credibility on everything else in the file.
- Let AI draft the record but never let it own the decision. Drafting closes the "document it later" gap and surfaces competencies you never assessed, but no machine-generated record enters the file until a named human has reviewed, corrected, and signed it.
- Consistent records are a fairness audit waiting to happen. The same five fields on every decision let you spot disparate rates and inconsistent pairs early and fix them as maintenance instead of meeting them as litigation.
Frequently Asked Questions
How specific does a routine rejection really need to be? One sentence tied to a posted requirement is usually enough, plus the five fields when a tool was involved. The test is not length but checkability: could a reader who never met the candidate identify which requirement was not met and verify it against the posting? What you cannot do is let depth vary candidate by candidate according to how interested the reviewer was. Depth should track the type of decision, never the identity of the person.
Isn't documenting more just creating more evidence against us? Only if you document the wrong things. The decisions happen whether or not you write them down, and any disparity exists whether or not you can see it. What a job-related record adds is the explanation, which is precisely what an employer is asked to produce. The dangerous file is the one full of impressions and speculation, or the one that is empty and forces a factfinder to guess.
Our interviewers say they cannot articulate why a candidate felt wrong. What do we do? Treat that as a training problem, because an intuition nobody can translate into a job-related observation is one you should not be acting on. Give interviewers a form that asks which required competency was in question, what specifically they observed, and how they would rate it on the shared scale. The exercise either produces a legitimate, documentable concern or reveals there was not one.
Skill.re