Decision Logging: Recording Human Decisions, AI Input, and Reasoning
Marcus runs recruiting operations for a 600-person fintech company, and last quarter his team processed 250 applicants for a backfilled platform-engineering role using an AI screening tool to triage resumes against the job description. In April, a rejected applicant filed an EEOC charge alleging the automated tool screened her out on a protected basis. Marcus's general counsel asked a deceptively simple question: "For each of those 250 people, can you show me what the AI recommended, what your recruiter actually decided, and the job-related reason for the difference?" For most of the applicants, Marcus could not. The AI scores lived in one system, the advance and reject decisions were buried in ATS status changes, and the reasoning lived in recruiters' memories. A decision log would have answered the question in an afternoon. Without one, Marcus was reconstructing intent from email fragments under litigation pressure. The lesson he took from that quarter is the subject of this one: a decision log is not paperwork, it is the evidence that decides whether you win or lose when someone challenges a hiring decision.
What a Decision Log Actually Is
A decision log is a structured, contemporaneous record of each significant hiring-stage decision that separates three things the law now insists be kept distinct: what the AI tool suggested, what the human decided, and why. The "why" must be tied to job-related criteria, not to impressions. A status field that reads "Rejected" is not a decision log. A note that says "not a fit" is not a decision log. A defensible entry tells a reader who was not in the room exactly how a specific candidate was evaluated, what role they were evaluated for, what the AI contributed, what a named human concluded, and which job-related criterion drove the outcome.
Two audiences make this matter: the external one, meaning a candidate, a regulator, or a court asking why a decision was made, and you, six months later, trying to remember whether the standard you applied to one candidate was the standard you applied to the next. It matters more now because the decision is no longer purely human. When an AI tool produces a match score, regulators treat it as part of the decision-making process and expect you to explain it and show that a human reviewed it. The log is where that explanation lives, and it is the difference between saying "our process is fair" and proving it.
What to Log for Every Significant Decision
Log at each decision that changes a candidate's status: advance to phone screen, advance to onsite, extend an offer, reject. The contemporaneous part matters, because a log written the day a charge arrives looks like, and may be treated as, after-the-fact justification. Start with candidate identity, meaning name or applicant ID and the specific requisition they applied for or were sourced for, since the same person considered for two roles generates two separate entries against different criteria. Record the stage, because "job-related" means something different at a resume screen than at an offer. Record the decision in plain terms: advanced, rejected, offer extended, offer declined. Record the timing as a real timestamp. Name the decision-maker, the individual or panel accountable, because accountability has to attach to a person rather than to a process.
Then come the fields that carry the evidentiary weight. Criteria assessed records what the candidate was evaluated against at this stage, drawn from requirements written down before screening began. Assessment records how they performed against each criterion in specific terms: "strong relevant experience, weaker communication skills" is an assessment, "good candidate" is not. AI input captures the score, the recommendation, and ideally the factors the tool cited; "AI match score 70 percent, flagged missing Kubernetes experience" is usable, "AI was used" is not. Human review of AI is a distinct field, and the one teams most often collapse into the previous one: if the tool made a recommendation, record whether a human actually reviewed it and what that person concluded on their own reading. Finally, the rationale is the heart of the entry, expressed against the requirements in the job description rather than against impressions: "Candidate meets technical requirements and has strong relevant experience. However, the communication skills assessment showed gaps that are critical for this role." The test is simple. This is objective assessment against defined criteria. This is not subjective impressions, gossip, or gut reactions, and any of those three appearing in a log is a liability rather than a defense.
Why You Separate What AI Suggested From What the Human Decided
The single most important structural feature of a good decision log is that it keeps the AI's input and the human's decision in two distinct fields. Collapsing them, by logging only the final outcome, destroys exactly the information a regulator or plaintiff's attorney will ask for first. Consider the two scenarios a log has to handle. In the first, the AI recommends advancing a candidate and the recruiter agrees. In the second, the AI recommends rejecting a candidate and the recruiter overrides that recommendation and advances them anyway. From the outside, both candidates simply show "advanced." Only a log that records the AI input separately from the human decision can show that in the second case a human exercised independent judgment and documented a job-related reason for departing from the tool.
That override, properly documented, is one of your strongest defenses, because it demonstrates the human was not rubber-stamping the algorithm. The inverse is equally important. If every recruiter advances exactly the candidates the AI scores highest and rejects the ones it scores lowest, with no independent reasoning recorded anywhere, your log reveals that "human review" in your process is a formality. Under several of the frameworks discussed below, that is a problem. The log does not create the problem; it surfaces it while you can still fix it. This is why a useful log captures the actual decision-making rather than a sanitized version.
A Worked Example: Two Entries From Marcus's Backfill
Here is how two of those 250 applicants should have been logged. Entry 1, the override that protects you. Candidate: A. Okafor (Applicant ID 4471). Role: Platform Engineer II (REQ-2291). Stage: Resume screen to phone screen. AI input: Match score 70 percent; tool flagged "missing Kubernetes (required skill)." Human review of AI: Recruiter read the flag against the posted requirements and judged it a token match rather than a skills gap. Human decision: Advanced to phone screen. Rationale, job-related: The job description lists "container orchestration experience" as the requirement and names Kubernetes only as a preferred example. The candidate shows five years operating Docker and AWS ECS in production, which is directly transferable container-orchestration experience. Decision-maker: J. Lin, Senior Technical Recruiter. Timestamp: 2026-03-12 10:42.
That entry records that a human reviewed the AI output, identified a specific limitation in how the tool scored the candidate, and advanced them for a documented, job-related reason. If Okafor belongs to a protected group and the tool systematically under-scores transferable skills, this entry is the recruiter catching it.
Entry 2, the rejection you have to defend. Candidate: R. Mehta (Applicant ID 4502). Role: Platform Engineer II (REQ-2291). Stage: Phone screen to reject. AI input: Match score 82 percent; no flags. Human review of AI: Score reviewed and set aside as resume-derived only; the screen tested criteria the resume could not show. Human decision: Rejected. Rationale, job-related: The phone screen assessed the two must-have criteria for this role, production on-call ownership and incident-response leadership. The candidate confirmed individual-contributor work only, with no on-call rotation or incident ownership in the last three roles, both of which are required for this position's primary responsibility. Decision-maker: J. Lin, Senior Technical Recruiter. Timestamp: 2026-03-14 15:08.
Note what makes the rejection defensible: a human rejected a candidate the AI scored highly, and the reason is tied to a specific, must-have, job-related criterion assessed at that stage, not to a vague sense of fit. Both entries diverge from a naive reading of the AI score, and both document the job-related "why." Not every entry needs this much text. A shorter narrative form works for routine calls: "AI scored 8 out of 10 on technical fit. I agreed and advanced the candidate. Communication concern noted for the next interview." That is still a complete entry, because it records the tool's output, the human's position on it, the reason, and the open question being carried forward.
Choosing a Logging System
Different systems work for different organizations, and the choice matters less than the discipline. ATS-based logging is strongest when your applicant tracking system has real fields for decision, criteria, assessment, and rationale, because every decision gets logged as the workflow progresses and nobody has to remember a separate step. Structured forms, digital or on paper, work when the ATS cannot be configured, filled out at the time of the decision and attached to the candidate record. Free-form notes with an imposed structure can work for small teams provided everyone writes to the same shape every time. What fails is free-form notes with no structure at all, because the fields you did not think to write are exactly the ones you will be asked about.
Whichever system you pick, the key is consistency, because a log is only evidence if the same fields are captured the same way for every candidate. A decision is logged identically whether the candidate arrived as a referral or a direct applicant, whether they are in one location or another, and whether or not they belong to a protected group. Inconsistency is itself a liability: if detailed reasoning is recorded for some candidates and "not a fit" for others, a plaintiff's attorney will ask why the well-documented and thinly documented candidates fall along a demographic line, and there is rarely a good answer. Standardize the fields, build them into the workflow so logging is not a separate chore performed at the end of a long week, train everyone on the standard, and name one person responsible for sampling entries and reporting on quality. Decide also on scope and stick to it. A workable default is to log decisions at the interview stage and beyond in full narrative depth, and to log screen-stage decisions in the shorter criteria-and-rationale form rather than trying to write an essay for every one of 250 resumes. What you cannot do is vary the depth candidate by candidate according to how the reviewer felt that day, because uneven depth is precisely what a plaintiff's attorney reads as differential treatment.
The Legal Frameworks Your Log Answers To
NYC Local Law 144. New York City's law on automated employment decision tools requires employers using an AEDT to screen candidates in the city to conduct an independent bias audit and to provide notice to candidates. The premise underlying the law is that automated screening decisions must be auditable and subject to human oversight. Your decision log is the internal record that lets you show, candidate by candidate, that the tool's output was reviewed rather than applied mechanically. When the audit asks how the tool affected real decisions, the log is your data.
EEOC expectations. The Equal Employment Opportunity Commission has made clear that using an algorithmic or AI tool does not relieve an employer of responsibility under Title VII. If an AI tool produces a disparate impact on a protected group, the employer is on the hook. The EEOC's guidance contemplates that automated tool decisions should be explainable and that a human remains accountable. A log that separates AI input from documented human reasoning is how you demonstrate the human accountability the EEOC expects, and how you spot a disparate-impact pattern before it becomes a charge.
FCRA adverse-action documentation. When a third-party background or assessment report contributes to a decision not to advance or hire a candidate, the Fair Credit Reporting Act imposes adverse-action obligations, including pre-adverse and adverse-action notices and a record of the basis for the decision. To the extent an AI-driven screening report functions as a consumer report in your process, your decision log supplies the documented, job-related reason that the adverse-action framework expects you to be able to produce.
GDPR Article 22 and the right to an explanation. For candidates in the EU or UK, Article 22 of the GDPR restricts decisions based solely on automated processing that produce legal or similarly significant effects, and it entitles the individual to human intervention and to contest the decision. Recital 71 describes a right to an explanation of an automated decision. Your decision log is the artifact that proves the decision was not solely automated: it shows the human in the loop and records the reasoning you would provide if a candidate exercised that right.
The common thread across all four is that the decision log is the evidence. In a discrimination claim or an audit, the question is never whether your process was well-intentioned; it is whether you can produce a contemporaneous, job-related record of how each decision was made. A log built into the workflow answers yes. A log reconstructed under pressure answers maybe, and maybe loses.
Using Logged Decisions for Improvement
Decision logs are not only for compliance, and treating them that way wastes the more valuable half of the asset. Calibration comes first: pull a set of decisions and review them as a team, asking whether different members apply the same criteria to the same evidence. Where they do not, discuss the divergence and align, because unaligned standards are how the same candidate gets two different answers depending on who picks up the file. Accuracy checking looks backward at hires: read the original decision notes for people who have now been in the role a year, and ask whether the criteria you assessed actually predicted success. If a criterion you weighted heavily has no relationship to performance, refine or drop it. That is a criterion you have been screening people out on for no reason.
Fairness analysis disaggregates the logged decisions by demographic group and by stage: are women advanced at the same rate as men from screen to interview and from interview to offer, and are candidates from certain locations or sources treated differently at the same gate? Disaggregation is the only way to see a pattern that is invisible one candidate at a time. Bias detection reads the rationales rather than the outcomes. If rejections of women overwhelmingly cite "weak communication skills" while rejections of men cite "insufficient experience," that asymmetry is a signal worth investigating, because the same underlying evidence is being labeled differently by group. Improvement opportunities emerge from the same reading: if candidates are repeatedly rejected for gaps that a few weeks of training would close, your criteria may be too strict for the market you hire in. Analyze quarterly rather than only when something goes wrong, and pay particular attention to overrides. Are humans overriding the AI often, rarely, or never, and when they do, is the documented reasoning sound? That data improves both the tool and the people using it.
Three Anti-Patterns
Logging theater. A company implements a logging system, the forms get filled out, and the forms capture nothing. "Rejected" with no explanation. Over time nobody puts real thought into it and the entries become reflexive. It happens because logging feels like overhead, so the minimum viable amount of information gets captured. What goes wrong is that the logs exist but are not defensible and do not improve decision-making, which is arguably worse than having no log at all, because they document that your stated process was hollow. The fix is to design logging around what you actually need, make it efficient but substantial, require the job-related rationale field, and reject entries that leave it empty.
Logging bias instead of correcting it. A team's logs show that candidates are rejected for "weak culture fit" disproportionately when they are women. Instead of investigating whether the assessment itself is biased, the team simply keeps logging it, effectively accepting that women have weaker culture fit and continuing to decide on that basis. It happens when logging is treated as documentation rather than as a trigger for improvement. What goes wrong is severe: you now have written proof that you saw the pattern and did nothing, which increases rather than reduces your legal exposure. The fix is to treat a pattern in the logs as the first step of a process, not the end of one. Identify it, investigate it, correct it, and document the correction.
Inconsistent logging. Some recruiters write detailed notes explaining their decisions; others log the bare minimum. Some assess against the defined criteria; others record subjective impressions. It happens because logging is not consistently required or enforced, and everyone finds their own level. What goes wrong is that the logs become hard to learn from and hard to defend, since neither analysis nor comparison works across records that do not share a shape. The fix is to standardize the format, train everyone on it, and give one person responsibility for auditing consistency.
Practice
- Design a decision logging template for your organization. Decide exactly which fields must be captured for each decision, and write the one-line instruction that tells a recruiter what a sufficient rationale looks like.
- Implement logging for one week of real decisions, then read back what you logged. Is it sufficient for compliance? Is it sufficient for improvement? Which field did you find yourself skipping under time pressure?
- Pull a sample of logged decisions from across your team and analyze them for consistency. Do different members log at different depths, or against different criteria, for comparable candidates?
- Disaggregate your logged decisions by demographic group and by stage. Are decisions being made at consistent rates across groups, and do the stated rationales differ in kind between groups?
- Use the logs to find one improvement opportunity. Identify a pattern you can see only in aggregate, and write down exactly what you would change, who owns the change, and how you would know whether it worked.
Reflection
- If you were questioned about a hiring decision you made six months ago, could you explain why it was made, and could you prove it with a record rather than a recollection?
- What information do you most regret not having logged about a past decision, and what would having it have changed?
- How would you motivate your team to log decisions consistently, given that logging always feels like overhead in the moment?
- What decision pattern have you noticed in your own hiring that surprised you, and did you find it by looking or by accident?
- How would you use a quarter of logged decisions to improve hiring in the next quarter? What specifically would you look at first?
Glossary
- Decision log. A documented record of hiring decisions, including who decided, what the decision was, what criteria were assessed, what the AI contributed, and why the decision was made.
- Calibration. The process of aligning standards across multiple decision-makers, using logged decisions to check whether the same criteria are being applied the same way.
Related Lessons
Decision logging is one layer of a wider documentation discipline, and these lessons cover the layers around it.
- Documentation and Evidence: Building a Trail for Compliance covers the record around the decision, including tool version, retention, and how a complete trail answers a regulator.
- Documenting Decisions: Clear Records for Legal and Fairness Review works through the writing itself, meaning how to phrase a rationale so it survives a hostile reader.
- Bias Audit Trails: Creating Evidence of Fairness Considerations turns the fairness analysis of your logs into a documented cycle of finding, investigating, and correcting.
- Compliance and Legal Review: Documentation for FCRA, EEO, and GDPR goes deeper on the four frameworks named here and what each expects you to produce.
- Avoiding Automation Bias: Staying Active and Skeptical addresses the behavior your override data measures, which is whether human review is real.
Closing
Comprehensive decision logging enables both compliance and continuous improvement, which is why it is best understood as infrastructure rather than overhead. It costs a few minutes per decision and buys two things you cannot buy later: a contemporaneous record when someone challenges a decision, and a dataset showing how your team actually decides rather than how it believes it decides. Marcus rebuilt his process around a fixed set of fields captured inside the ATS, and the next time counsel asked his question, the answer took an afternoon instead of a quarter. Good logging is also how you learn, because the team that reads its own logs finds the criterion that never predicted anything and the rationale pattern that splits along a demographic line while there is still time to fix it. That is the argument to make to a team that resists logging: you are not writing this down for the lawyers, you are writing it down so that next quarter's decisions are better than this quarter's.
Key Takeaways
- A decision log is evidence, not paperwork. When a candidate challenges a decision or a regulator audits your process, the question is whether you can produce a contemporaneous, job-related record of how each decision was made. Memory and email fragments reconstructed under pressure do not answer it.
- Log objective assessment against defined criteria, never subjective impressions. Gossip, gut reactions, and "not a fit" are liabilities in a log. Assessment against a criterion the job description names is a defense.
- Separate what the AI suggested from what the human decided. Keep AI input, human review of that input, and the human decision in distinct fields. A documented override for a job-related reason proves human review is real rather than a rubber stamp.
- A complete entry has a fixed shape. Candidate and role, stage, decision, timestamp, decision-maker, criteria assessed, assessment against those criteria, AI input, human review of the AI, and the rationale, which is the heart of it. Capture the same fields with the same rigor for every candidate regardless of source, location, or group, because uneven documentation that tracks a demographic line is what a discrimination claim is built from.
- Four legal frameworks expect this record. NYC Local Law 144 demands auditable, human-reviewed automated screening; the EEOC holds you accountable for AI-driven disparate impact; FCRA requires documented adverse-action reasons; and GDPR Article 22 entitles candidates to human review and an explanation.
- Use logs for calibration, accuracy checking, fairness analysis, and improvement. Review them quarterly to check whether overrides are justified, whether criteria are applied consistently, and whether your criteria predict success. When the logs reveal a bias pattern, investigate and correct it, because logging a pattern and continuing to decide on it increases your exposure.
Frequently Asked Questions
Does a decision log create legal risk by documenting decisions that could be second-guessed? The logic runs backwards. The decisions happen whether or not you write them down, and the disparities in them exist whether or not you measure them. What the log documents is the quality of your reasoning, and a contemporaneous, job-related rationale is far stronger evidence than an empty file that forces a factfinder to guess. The genuinely dangerous record is the one that logs a bias pattern and shows no response, which is why anything your log reveals, you investigate and act on.
What does a sufficient rationale look like in one line? It names the criterion, states the evidence, and connects them. "Rejected: no forklift certification, which the posted requirements list as required" works. "Advanced despite the Kubernetes flag: five years of Docker and ECS is the container-orchestration experience the role actually asks for" works. "Not a fit," "seemed sharp," and "team liked her" do not, because none of them can be checked against anything.
Who should own the logging standard, and how often should the logs be analyzed? One named person, not the team collectively, owns the template, trains new joiners on it, and samples entries each cycle for completeness and consistency. Analyze quarterly for calibration, override patterns, and fairness disaggregation, with accuracy checking on a longer cycle since it requires hires to have been in role long enough to assess. Analyze on a schedule rather than in response to a problem, because a review that only happens after a complaint arrives too late to prevent one.
Skill.re