Hands-On Project: Design a Quality Audit Plan
Devon manages recruiting for a 600-person healthcare network across four facilities, and his team leaned hard on AI to keep up with a flood of clinical and administrative applicants. The AI drafted summaries of every screened candidate, and for months Devon assumed the summaries were fine because no one complained. Then a hiring manager forwarded him a summary that confidently described a nurse as having ICU experience the resume never mentioned. Devon went looking and could not answer a basic question: how often was that happening? He had no sample, no error rate, no threshold, and no record. This project is about replacing that blind spot with a quality audit plan, and the deliverable is concrete: a one-page document you could hand to a colleague and say, this is how we audit the quality of our recruiting.
What a Quality Audit Plan Actually Is
Quality systems are not theoretical concepts. They are operational processes that run week after week, gathering data, surfacing problems, and driving improvement, and the defining trait of a good one is that it survives contact with a busy quarter. A plan that requires a day of analysis nobody has time for is a plan that lives on a shelf, and a plan on a shelf provides exactly as much quality assurance as no plan at all. That is the standard this project holds you to: not comprehensiveness, but implementability.
A complete audit plan answers seven questions in plain language. What are we measuring, expressed as specific metrics across accuracy, fairness, consistency, and candidate experience? How often do we measure it? What data do we need, meaning what information must be captured in your applicant tracking system and your process for auditing to be possible at all? How will we sample the decisions we audit? How will we analyze what we collect, and what comparisons will we make? Who sees the results, and in what form? And if we find a problem, what specifically will we do about it? If your draft cannot answer all seven, it is not yet a plan. It is an aspiration.
Work through the nine steps that follow in order, writing your answers down as you go. At the end you will have the deliverable, and the last section gives you a minimal starter template if you would rather begin from a skeleton and adapt it than build from nothing.
Step 1: Choose Four or Five Metrics
Start with metrics, and start small. Ask which four or five measurements would give you genuine confidence that quality is being maintained, taking at least one from each of the four dimensions rather than loading up on whichever dimension is easiest to count. Easy measurement is the trap that produces audits full of numbers about the wrong things.
For accuracy, pick both a lagging and a leading indicator if you can. The percentage of hires who stay at least one year is a lagging measure, honest but slow. The phone-screen-to-interview conversion rate moves much faster and tells you sooner whether the screen is advancing the right people. For fairness, the workhorse is the disparate impact ratio calculated for each demographic group at each stage, which answers whether any group is screening out at a higher rate than others. For consistency, measure the share of candidates rated by the same AI who receive the same decision when re-screened, which tests whether the system is stable. For candidate experience, measure the share of rejected candidates who report feeling respected and understanding why they were rejected, alongside a net promoter score on the overall process.
Before you finalize the list, apply one filter: what would most damage your hiring if it were wrong? Start there. It is genuinely easy to measure what is easy to measure and end up with an audit that is precise about something that does not matter much, while the thing that would actually hurt you goes uncounted.
Step 2: Set the Frequency for Each Metric
Not everything needs the same cadence, and forcing one schedule onto every metric is how plans become unaffordable. Some measures suit near-real-time monitoring, particularly consistency checks that can run automatically as decisions are made. Fairness analysis typically works monthly, because disparate impact ratios need enough decisions behind them to be meaningful. Candidate experience usually reads best quarterly, where you are watching a trend rather than reacting to a week.
Frequency depends heavily on volume. A high-volume function generates enough decisions to support frequent analysis, while a team making a handful of hires a month will produce ratios that swing wildly week to week and mislead more than they inform. Devon's team runs a weekly audit of AI-generated summaries because the volume supports it, and reports monthly. Match your cadence to the number of decisions you actually make, not to how often you would like to feel informed.
Step 3: Specify the Data You Need
An audit plan is only executable if the underlying data exists, so write down explicitly what must be captured in your systems and process. That list typically includes the AI scores or outputs themselves, the human reviewer notes attached to each decision, candidate demographic information, dates and timing for cycle-time analysis, and candidate feedback.
One detail carries disproportionate weight. You need demographic information across all candidates, not only those who advanced, because rates cannot be calculated without a denominator. An audit that captures demographics only for hires can tell you the composition of your hires and nothing at all about whether your screen was fair. Where demographic data comes from voluntary self-identification, which is the cleanest source, note the response rate, and be explicit in the plan about which portions of your demographic assignment are self-reported. Handle all of it under the privacy rules that apply to you, keep access limited to the people running the audit, and do not repurpose demographic data collected for compliance into anything that touches an individual hiring decision.
Step 4: Design the Sampling Method
You do not need to audit everything. Smart sampling gives you reliable insight for a fraction of the effort, and that efficiency is the only reason the plan is sustainable at all. The arithmetic is simple: if you make 500 decisions per month, auditing 50 of them gives you a ten percent sample, which is a defensible starting point. Devon runs the same ten percent on a weekly cadence because his volume supports it, and the consistency of the slice matters more than the exact percentage, since comparable samples are what make trends visible.
Do not sample blindly. Stratify by demographic group and by risk level so that a single high-volume requisition does not dominate the sample and so that each group is represented well enough for a group-level rate to mean anything. Then deliberately oversample borderline decisions, meaning the cases where a candidate scored near a decision boundary, because that is where an error does the most damage. A wrong assessment of an obvious reject changes nothing. A wrong assessment of a borderline candidate advances the wrong person or screens out the right one.
Step 5: Write the Rubric That Defines a Defect
An audit is only as good as its definition of an error. Devon writes a rubric precise enough that any two reviewers scoring the same summary reach the same verdict, because an audit whose results depend on who happened to run it is measuring the reviewer rather than the system. That property has a name worth knowing: inter-rater reliability, the extent to which multiple evaluators agree on the same assessment.
His rubric checks each sampled summary against the source material on four points. Factual accuracy: does every claim trace to something actually in the resume or screen notes, with no invented experience, credentials, or dates? Completeness: are the qualifications that matter for the role present, or did the summary drop a disqualifying gap or a key certification? Fairness: does the summary avoid inferring or surfacing protected characteristics, and avoid penalizing non-traditional paths? Usability: is the output in a form a hiring manager can act on without rewriting it? A summary failing any one of the four is logged as an error.
Devon is careful here about a trap that has sunk other audits: measuring consistency without measuring correctness. An AI can be perfectly consistent and consistently wrong, and a plan that only checks whether the system produces stable output is measuring precision rather than accuracy. Every rubric needs at least one check that asks whether the output is right, verified against the source, not merely whether it is stable across runs.
Step 6: Set the Threshold Before You Start
Counting errors is only useful against a line drawn in advance. Devon sets his acceptable error rate at 5 percent and treats any weekly sample exceeding it as a trigger for action rather than a topic for next quarter's agenda. The threshold is what converts a number into a decision. Without it, a 12 percent error rate is an interesting fact; with it, 12 percent is an instruction to investigate before the next batch of candidates flows through.
Set the fairness threshold the same way and set it explicitly: a disparate impact ratio below 0.8 for any group at any stage triggers investigation. That figure is the four-fifths rule, the long-standing rule of thumb for adverse impact, and it belongs in your plan as a written trigger rather than as background knowledge. Track every rate week over week rather than judging a single week in isolation, because one bad sample can be noise while a rising trend across three weeks is a signal. The threshold catches acute failures; the trend line catches slow drift as an underlying model updates or the role mix shifts.
Step 7: Define the Analysis
Write down which analyses you will actually run, because a plan that collects data without naming the comparisons produces a spreadsheet nobody knows how to read. For fairness, calculate the disparate impact ratio for each demographic group at each stage, taking each group's selection rate and dividing it by the highest group's rate. Run it stage by stage rather than across the funnel as a whole, because a funnel-wide pass can conceal a single stage that fails badly.
For consistency, sample decisions and check whether the AI made the same assessment on re-screen, and check whether different human evaluators rate comparable candidates similarly. For accuracy, compare the audited decision against what happened next: did the candidate the screen advanced succeed or fail at the following stage? For candidate experience, analyze feedback patterns rather than individual comments, looking for the themes that repeat.
One discipline governs all of it. A failed ratio is a flag, not a verdict. Disparate impact analysis identifies where outcomes differ; it does not by itself establish that the cause is unlawful or even that the criterion is wrong. What it obligates you to do is investigate whether a genuine job-related factor explains the difference, and to be able to show that any criterion producing the gap is job-related and consistent with business necessity. What you cannot do is compute the ratio, notice it is below 0.8, and carry on as though you never looked.
Step 8: Decide Who Sees the Results
Name the audience before you design the report. In most organizations the list includes recruiting leadership, whatever audit or governance committee oversees the AI tooling, and sometimes organizational leadership, and each of those audiences needs a different level of detail. Then pick the format deliberately: a standing dashboard, a monthly written report, a quarterly business review, or some combination.
Devon sends a short monthly summary covering the weekly error rates, any thresholds breached, the root causes found, and the actions taken. The format is a one-page trend rather than a forensic report, because a report nobody finishes is a report nobody acts on. He keeps the underlying detail available for anyone who wants it, which matters for a different reason: the audit record itself becomes evidence of your quality commitment if a decision is ever challenged, so document what was audited, what was found, and what was done.
Step 9: Connect Findings to Decision Authority
An audit that surfaces a problem and reports it into a void is wasted effort, and the failure is usually structural rather than personal. Before designing a single metric, confirm who decides what to do about findings, how quickly they can decide, and what the escalation path is if something urgent surfaces. A team that audits monthly and routes findings to a council that meets quarterly has built a system in which biased decisions continue for months while everyone behaves reasonably.
Devon confirmed his authority first. A breach above threshold routes to him with standing authority to pause or revise a prompt immediately, and a systemic or fairness finding escalates to the committee that governs the AI tooling. He then wrote a simple decision tree so responses are not improvised under pressure: a prompt-level error pattern triggers a prompt revision; a model regression from the vendor triggers a vendor conversation; a disparate impact ratio below 0.8 for any group triggers a deeper investigation of that stage, examining both the criteria and how consistently they were applied; low consistency triggers retraining or reconfiguration of the tool. The audit's value lives entirely in this loop. Auditing without action does not merely waste effort, it signals to the team that the findings do not really matter.
Worked Example: A Weekly Sample That Trips the Threshold
One Monday, Devon's weekly audit pulls a sample of 20 AI-generated summaries, stratified across his open requisitions. A reviewer scores each against the four-point rubric. Three come back as errors: one invented a certification the candidate never listed, one omitted a licensing gap that mattered for a clinical role, and one inferred a candidate's likely age from graduation dates and let it color the tone. Three errors in 20 summaries is a 15 percent error rate.
Fifteen percent sits well above the 5 percent threshold, so the plan triggers. Devon does not shrug and move on, and he does not throw out the AI. He treats the failures as a diagnosis. All three trace back to the same screening-and-summary prompt, which gave the model too much latitude to fill gaps and infer context. He revises the prompt to forbid inferring credentials or demographic traits and to require that any missing qualification be flagged explicitly rather than smoothed over. He re-runs the audit on a fresh sample of 20 the following week and the error rate drops to 1 in 20, which is 5 percent, back at the line. He logs the whole episode: the sample, the three errors, the root cause, the prompt revision, and the follow-up result.
Note what the third error triggered beyond the prompt fix. A summary that inferred age from graduation dates is a fairness defect, not merely a quality defect, so it escalated under the decision tree to a review of whether that inference had influenced advancement decisions, which is exactly the question a disparate impact analysis at that stage is designed to answer. The lesson Devon draws is that the threshold did the real work. The 15 percent figure was alarming, but it only mattered because he had committed in advance to acting above 5 percent. A measured error rate with no threshold attached is a metric that makes you feel responsible without making you do anything.
A Minimal Starter Template
If you would rather adapt a skeleton than start from a blank page, here is a minimal plan small enough to actually run. Write your own version of each line, then extend it only after it has survived a full quarter.
- Metric 1, accuracy. Are the right candidates advancing? Frequency: monthly. Sample: 20 interviews per month. Measurement: did the candidate succeed or fail at the next stage?
- Metric 2, fairness. Is disparate impact present? Frequency: monthly. Sample: all candidates screened. Measurement: disparate impact ratio by demographic group, with any ratio below 0.8 triggering investigation.
- Metric 3, consistency. Are all evaluators applying the same standards? Frequency: monthly. Sample: 5 decisions per evaluator. Measurement: do evaluators rate comparable resumes similarly?
- Reporting. A monthly dashboard to leadership showing trends rather than isolated figures, with breaches and actions taken called out explicitly.
Notice what the template does not include. It has no twenty-metric matrix, no weekly deep analysis, and no dimension added because it seemed thorough. Three metrics, three samples, one report. That is the version that gets run in a quarter when two people are out and a requisition catches fire.
Making the Plan Stick
An audit plan is only useful if it is actually implemented, and five habits separate the plans that run from the ones that quietly stop. Assign responsibility by making the audit a named person's job; if it is everyone's job it is no one's job. Build it into the process by scheduling audits at fixed intervals, putting them on calendars, and treating them as non-negotiable rather than as work that happens when there is time. Use consistent methodology, keeping the same sampling method, analysis approach, and documentation each cycle so results are comparable over time; a change in method mid-year destroys your ability to read a trend.
Document everything, meaning what was audited, what was found, and what was done about it, which creates the evidence of your quality commitment and the institutional memory that stops the same failure from recurring. Communicate findings to leadership and to the teams whose work was audited, and use them to drive visible improvement. That last habit is what makes the practice self-sustaining. When a team sees that audits lead to prompt fixes and better summaries, the audit stops feeling like overhead and starts feeling like the thing that keeps their work defensible.
Anti-Patterns
The overly ambitious audit plan. A team designs a plan to measure twenty metrics with weekly analysis. It is comprehensive, it is admired in the meeting where it is presented, and nobody has time to execute it, so it sits on a shelf. It happens because teams want to be thorough and because thoroughness is easier to demonstrate than discipline. What goes wrong is that the plan is never implemented and quality goes entirely unmonitored, which is a worse outcome than a modest plan that runs. The defense is to start simple and doable: four or five key metrics, run for a quarter, expanded only if the simple version proves insufficient.
The audit plan with no decision authority. You audit, you find disparate impact, and you report it to leadership. But the council that approves changes to the AI tooling does not meet until next quarter, so the finding sits for months while the affected decisions continue. It happens because the audit process was designed separately from the decision-making structure. What goes wrong is that findings do not drive action and quality problems persist while everyone follows the process correctly. The defense is to clarify decision authority before designing the plan: who decides, how fast they can decide, and what the escalation path is for an urgent finding.
The audit that measures the wrong thing. You audit and find that the AI scores are highly consistent, congratulate yourself on the quality, and never check whether those consistent scores are actually correct. You have measured precision rather than accuracy, and the system could be consistently wrong. It happens because it is easy to measure what is easy to measure. What goes wrong is that you invest real effort in auditing and learn something less important than what you needed to know. The defense is to define the most important questions before defining metrics, starting from what would most damage your hiring if it were wrong.
Practice Prompts
- Choose your metrics. Define four or five quality metrics you would measure if you had perfect data, taking at least one from each dimension. Then rank your priorities across accuracy, fairness, consistency, and candidate experience, and be honest about which one you have been avoiding because it is hard to count.
- Design the sampling plan. Decide how many decisions you would audit per month and how you would stratify them. Work out what percentage of your actual monthly volume that represents, and where you would oversample borderline cases.
- Build the reporting view. Sketch an audit dashboard. What does it show, who sees it, and at what frequency? Limit yourself to what fits on one page.
- Write the decision tree. Specify what you would do if you discovered disparate impact, and separately what you would do if you discovered inconsistent AI decisions. Name who has the authority to act in each case and how quickly.
- Draft the full plan. Write an audit plan for one specific role or decision, such as engineering screening or phone-screen assessment, covering metrics, frequency, sampling methodology, analysis plan, reporting, and action plan. This is the deliverable.
Reflection
- What is the quality concern that keeps you up at night, and what specific metric would give you confidence it is not happening?
- If you discovered a quality problem through auditing, how would you communicate it to your leadership? What story would you tell with the data?
- What would prevent your organization from acting on quality audit findings? Get specific about the barrier: time, resources, or willingness.
- How would you build buy-in for an audit plan that requires investment and regular attention? What would convince your leadership this matters?
- What is the minimum audit plan you could implement in the next three months, and what is the smallest viable starting point you would actually run?
Glossary
- Audit plan. A documented approach to monitoring quality across multiple dimensions, including metrics, sampling methodology, analysis plan, and reporting.
- Disparate impact. A hiring practice that disproportionately affects members of a protected group, whether or not anyone intended that result.
- Disparate impact ratio. A group's selection rate divided by the highest group's rate at the same stage. Under the four-fifths rule, a ratio below 0.8 is a flag warranting investigation.
- Inter-rater reliability. The extent to which multiple evaluators reach the same assessment on the same material, which is what a well-written rubric is designed to produce.
- Stratified sampling. Drawing your audit sample deliberately across groups, requisition types, and risk levels rather than at random, so each segment is represented well enough to analyze.
- Leading and lagging indicators. Leading measures move early and let you react, such as a conversion rate; lagging measures confirm outcomes slowly, such as one-year retention.
- Threshold. A limit committed to in advance that converts a measured rate into a required action.
Related Lessons
This project sits at the center of the quality systems chapter and draws directly on several neighbors.
- Quality Dimensions: Accuracy, Fairness, Consistency, and Candidate Experience defines the four dimensions this plan measures and is worth revisiting before you choose metrics.
- Auditing AI-Assisted Decisions: Sampling Methodology and Fairness Metrics goes deeper on the sampling and measurement mechanics behind steps four and seven.
- Metrics and Monitoring: Tracking Efficiency, Quality, and Fairness covers the wider measurement system your audit plan reports into.
- Hands-On Project: Audit a Recruiting Workflow for Bias is the companion project, applying the disparate impact analysis in step seven to a full funnel.
- Documentation and Evidence: Building a Trail for Compliance explains why the audit record itself is a compliance asset, not merely internal housekeeping.
- Governance Structures: Committees, Roles, and Decision Authority addresses the decision-authority question in step nine, which is the step most audit plans skip.
Closing
Quality audit plans operationalize quality systems. They are the infrastructure that maintains confidence in recruiting decisions as AI is introduced, and they are how you move from hope, meaning the sense that you are probably hiring well, to measurement, meaning the evidence that you are. That shift is what makes a recruiting function defensible and continuously improving rather than merely well-intentioned.
Devon's plan is not impressive on paper. Four rubric checks, a ten percent weekly sample, one threshold, one monthly report, one decision tree. Its virtue is that it runs, week after week, and that when it found something it produced a fix within days rather than a finding within quarters. Build the practice into your normal operations from the start, start small if you need to, and let the discipline of measurement establish itself before you make it more elaborate.
Key Takeaways
- An effective audit plan is simple, measurable, and connected to decision-making. Complex plans do not get implemented, and an unimplemented plan provides no quality assurance at all.
- Measure all four dimensions. Accuracy, fairness, consistency, and candidate experience, with at least one metric from each. Do not audit in silos; look at how the dimensions interact, since a fast process that screens unfairly is not a quality process.
- Specify the data before you specify the analysis. AI outputs, reviewer notes, demographics across all candidates rather than only hires, timing data, and candidate feedback. Rates cannot be calculated without a denominator.
- Use risk-based and stratified sampling. Auditing 50 of 500 monthly decisions is a workable ten percent slice; stratify by group and risk, and oversample borderline cases where an error does the most damage.
- Write a rubric that defines a defect. Score each item on factual accuracy, completeness, fairness, and usability so two reviewers reach the same verdict. Check correctness, not just consistency: an AI can be consistently wrong.
- Set thresholds in advance. An acceptable error rate turns a measured rate into a decision, and a disparate impact ratio below 0.8 at any stage is a written trigger for investigation, not a background fact.
- A failed ratio is a flag, not a verdict. It obligates you to investigate whether a genuine job-related factor explains the gap; what you cannot do is measure it, see it below 0.8, and proceed as though you never looked.
- Report to decision-makers promptly. Monthly or quarterly reporting allows faster response than annual review, and a one-page trend gets read where a forensic report does not.
- Connect findings to authority before designing metrics. Confirm who decides, how fast, and what escalates. Auditing without action wastes effort and signals that the findings do not really matter.
- Make it stick. Name an owner, calendar it as non-negotiable, keep the methodology constant so trends stay readable, document every run, and share results so the practice earns its place.
Frequently Asked Questions
How many metrics should I actually start with? Four or five, with at least one from each dimension, and no more until the simple version has run for a full quarter. The instinct to be comprehensive is the single most reliable predictor of an audit plan that never runs, because a twenty-metric plan requires analysis time nobody has budgeted. A plan with four metrics that produces a finding and a fix every month is worth more than a thorough plan that produced one impressive document. Expand after the practice is established, and expand toward whichever question your existing metrics keep failing to answer.
We do not have demographic data on most candidates. Can we still audit fairness? Partly, and you should be explicit about the limits rather than quietly proceeding. Voluntary self-identification is the cleanest source, so improving your response rate on it is a legitimate first project. Where coverage is thin, a disparate impact ratio computed on a small or unrepresentative subset can swing on one or two individuals and mislead badly, so report the coverage rate alongside the ratio and treat a marginal result on sparse data as a reason to gather better data rather than a finding you would defend. In the meantime, the accuracy, consistency, and experience metrics do not depend on demographics and can run immediately.
What error rate should I set as my threshold? There is no universal number, and the honest answer is that the threshold should reflect what an error costs you in that specific process. Devon set his at 5 percent for AI-drafted summaries because a defective summary at his volume reaches a hiring manager quickly. A step with higher stakes per decision warrants a tighter line. What matters far more than the exact figure is that the line is written down before you collect data, because a threshold set after seeing the results will always be set just above them. Choose a number, commit to acting on it, and revise it deliberately rather than in the moment.
Our audit found disparate impact. What do we do first? Investigate before you conclude anything, and do not stop using the number as your starting point. Identify the specific stage where the ratio falls below 0.8, then examine what criteria are applied at that stage and whether they are applied consistently, reading the actual decision notes rather than reasoning about the process in the abstract. Ask whether a genuine job-related factor explains the gap, and be prepared to show that any criterion producing it is job-related and consistent with business necessity. Escalate under your decision tree in parallel, since a fairness finding usually needs an owner above the person who ran the audit, and document the investigation and its outcome regardless of what you conclude.
Who should run the audit? Should it be independent of the recruiting team? Someone must own it by name, which is the non-negotiable part, and for routine internal quality monitoring that owner is usually inside the recruiting function because they understand the process well enough to interpret findings. Independence becomes essential as the stakes rise: fairness findings should escalate to a governance body outside the team whose work is being measured, and where regulation requires an independent bias audit of an automated employment decision tool, that requirement is not satisfied by an internal review no matter how rigorous. Treat this project as the internal discipline that keeps you honest between formal audits, not as a substitute for one.
How do I get leadership to fund the time this takes? Lead with the exposure rather than the process. Devon's opening was not a request for headcount, it was a question he could not answer: how often does our AI put a claim in front of a hiring manager that the resume does not support? An organization that cannot answer that has no evidence of quality and no defense if a decision is challenged, and audit records are precisely the evidence that demonstrates the process was monitored and corrected. Then make the ask small. Ten percent of one week's decisions, four rubric checks, and a one-page monthly report is a modest recurring cost against the alternative of finding out from a candidate, a hiring manager, or a regulator.
Skill.re