Fairness Metrics: Defining and Measuring Bias in Outcomes
Marcus runs talent acquisition for a 1,200-person regional health system, and the question that keeps him up at night is no longer "are we hiring enough nurses?" It is "could I prove, if a regulator or a plaintiff's attorney asked, that the way we hire is fair?" His team screens roughly 9,000 applicants a year across forty req types, and somewhere in that funnel an AI resume-ranking tool, three structured interview rubrics, and a dozen recruiters are all making decisions that produce outcomes. Marcus does not get to claim fairness. He has to measure it. This lesson is the toolkit he built so that fairness stopped being a feeling and became a number he reviews every month, along with the discipline that keeps a number from being over-read.
Why Outcomes, Not Intentions
The first thing Marcus had to unlearn is that a fair process produces fair outcomes automatically. It does not. You can remove names from resumes, train every interviewer on unconscious bias, and still end up advancing one group at a meaningfully lower rate than another, because bias hides in proxies: zip code, school name, a gap in employment, the phrasing a screening model rewards. Fairness metrics measure what actually happened to people as they moved through the funnel, not what anyone intended. That distinction matters legally as much as ethically. Under the EEOC Uniform Guidelines on Employee Selection Procedures, a selection practice can be challenged on its results even when there was no discriminatory intent, through the doctrine of disparate impact, also called adverse impact. The numbers are the evidence. So Marcus measures outcomes by demographic group at each decision point, and treats his own good intentions as irrelevant to the audit.
This is also why fairness is teachable as a set of computations rather than as a disposition. Three instruments do most of the work, and the rest of this lesson builds them in order: the selection rate, which is the raw measurement; the disparate impact ratio, which turns two selection rates into a comparison against a legal benchmark; and equity analysis, which extends the same arithmetic beyond demographic categories to any dimension along which the funnel might be quietly sorting people. Each answers a different question, and none is sufficient alone.
Selection Rate: The Foundation
Every fairness metric starts with one simple quantity: the selection rate. For a given group at a given stage, it is the number of people from that group who advanced divided by the number who were eligible to advance. If 400 women applied to a req and 88 were moved to phone screen, the selection rate at the application-to-screen stage is 88 divided by 400, or 22 percent. Compute the same rate for every group you can responsibly track, and you have the raw material for comparison. The quantity is deliberately unglamorous, and that is its strength: it depends on counts your applicant tracking system already holds, it can be recomputed by anyone, and it does not require you to model anything.
The comparison that follows is a gap in percentage points. Suppose 65 percent of men are selected at a stage and 55 percent of women are selected. That is a 10-point gap. Is a 10-point gap significant? Marcus works to a rule of thumb the program teaches: gaps above 5 percentage points warrant investigation. A 5-point trigger is not a legal standard and it is not a finding, it is an attention threshold, chosen low enough that real disparities surface early and high enough that ordinary quarter-to-quarter noise does not swamp the review queue. A gap is not yet proof of anything; it is a flag that says investigate.
Marcus computes selection rates at each funnel stage separately, because a funnel that looks fair in aggregate can hide a stage that is not. A disparity introduced at the resume screen propagates to every stage after it and gets harder to see the deeper you go: by the time it reaches the offer decision it has been laundered through several rounds of apparently reasonable human judgment, made by people choosing from a pool that had already been narrowed unfairly.
The Disparate Impact Ratio
A gap in percentage points is intuitive but it is not the benchmark regulators use. The disparate impact ratio is, and it is a ratio rather than a difference. The four-fifths rule, also called the 80 percent rule and set out in the EEOC Uniform Guidelines, holds that the selection rate of a protected group should be at least 80 percent of the selection rate of the favored group. Divide the lower rate by the higher one and compare the result to 0.80.
Take the simplest case. If men are selected at 60 percent and women at 45 percent, the ratio is 45 divided by 60, which is 75 percent, or 0.75. That is below 80 percent, and it can indicate disparate impact. Notice how differently the ratio and the gap behave. The gap here is 15 percentage points; the ratio is 0.75. But if the same 15-point gap sat between rates of 90 percent and 75 percent, the ratio would be 0.83 and would clear the threshold. Ratios are sensitive to the base rate in a way that differences are not, which is exactly why the Uniform Guidelines chose a ratio: a 15-point shortfall matters far more when few people are advancing at all. Marcus computes both, because the gap tells him how many people are affected and the ratio tells him where he stands against the standard.
The Four-Fifths Rule: A Worked Example
Here is the calculation exactly as Marcus runs it for one of his high-volume nursing reqs over a quarter. The numbers below are illustrative, chosen to show the arithmetic, and are not drawn from any study.
Step 1: Count applicants and advances at the resume-screen stage for each group. Group A: 500 applied, 150 advanced to screen. Group B: 300 applied, 72 advanced to screen. Marcus records the raw counts, not just the rates, because an auditor will want to reproduce every number and because the counts tell him whether the sample is large enough to interpret.
Step 2: Compute each group's selection rate. Group A's rate is 150 divided by 500, which is 0.30, or 30 percent. Group B's rate is 72 divided by 300, which is 0.24, or 24 percent. The gap is 6 percentage points, already past his 5-point attention threshold.
Step 3: Identify the highest rate and compute the impact ratio. Group A has the highest rate at 30 percent, so it becomes the reference group. The impact ratio is 0.24 divided by 0.30, which equals 0.80, exactly 80 percent. The rule always compares every group against the most-selected group, so the reference can change from quarter to quarter as rates move.
Step 4: Compare to the four-fifths threshold. An impact ratio of 0.80 sits right at the line. Now suppose the next quarter Group B advances only 60 of 300 applicants: the rate becomes 0.20, the impact ratio becomes 0.20 divided by 0.30, which is 0.67, and that is well below 0.80. A ratio of 0.67 signals adverse impact. It does not automatically prove unlawful discrimination, but it shifts the burden: Marcus now has to either find and fix what is driving the gap or demonstrate that the screening criterion is job-related and consistent with business necessity. That is the structure the Uniform Guidelines impose, and it is why Marcus computes the impact ratio for every group at every stage and keeps the worksheet.
One practical caution Marcus learned: the four-fifths rule is unstable with small numbers. If only 12 people from a group applied to a req, a single decision swings the ratio wildly. For low-volume reqs he aggregates across similar roles and a longer time window before drawing conclusions, and he notes the sample size next to every ratio so no one over-reads a flag built on five applicants. A ratio without a denominator beside it is not a measurement, it is a rumor with a decimal point.
Equity Analysis Beyond Demographic Categories
Demographic groups are where the legal exposure sits, but they are not the only place a funnel sorts people unfairly. Equity analysis applies the same selection-rate arithmetic to other dimensions of the candidate pool and looks for patterns that indicate bias. Are candidates from certain geographies selected at different rates? Candidates from certain schools? Candidates from certain industries? Marcus runs these cuts precisely because they often expose the proxy that is driving a demographic disparity he has already flagged.
The mechanism is worth being explicit about. A screening criterion almost never says "prefer this group." It says "prefer candidates from these three feeder schools," or "prefer candidates whose prior employer was a large health system," or, in a model, it silently weights a feature that behaves the same way. If the school list, the employer profile, or the region correlates with a protected characteristic in your pool, a criterion that looks neutral produces a demographic gap. So when Marcus cuts selection rates by school and sees one cluster advancing at half the rate of another, he has a candidate explanation for a demographic ratio he could not otherwise account for. He treats geography, prior industry, and educational background as standing dimensions on the monthly review and applies the same 5-point and 0.80 triggers to them.
Demographic Parity Versus Equal Opportunity
Here is where Marcus had to get precise, because "fairness" is not one definition and the metrics measure genuinely different things. Choosing the wrong one can make a fair process look biased or a biased process look fair.
Demographic parity asks: are groups advancing at equal rates? The four-fifths rule is a demographic-parity test. Its strength is that it needs no information about who was actually qualified; it only looks at advancement rates, which means it can be computed from data every recruiting team already has. Its weakness is the same thing: if two groups genuinely differ in the qualifications a job legitimately requires, demographic parity can flag a gap that reflects real, job-related differences rather than bias.
Equal opportunity, also called true-positive-rate parity, asks a sharper question: among the people who were actually qualified, did each group get advanced at the same rate? It compares selection rates conditioned on qualification. If 80 percent of qualified Group A candidates advance but only 60 percent of qualified Group B candidates advance, that is an equal-opportunity violation even if overall advancement rates happen to look balanced. This metric targets the failure mode recruiters most care about: qualified people from one group being passed over.
The catch, and Marcus is honest with his team about this, is that equal opportunity requires a defensible measure of "qualified," and that label can itself carry bias. He uses post-hire performance ratings and structured-rubric scores as imperfect proxies, documents their limits, and never treats either metric as the whole truth. Demographic parity and equal opportunity can also point in opposite directions for the same data; they cannot always be satisfied at once. That is not a bug to engineer away but a real trade-off to decide deliberately, in writing, with legal and the business at the table.
Using Multiple Metrics Together
No single metric is perfect, and Marcus does not let any one of them stand alone. Each instrument has a characteristic blind spot, so he triangulates: run them all, treat agreement as strengthening a finding, and treat disagreement as a question to resolve rather than a result to pick from. A passing ratio is not an all-clear, because the metric that passed may be blind to the failure you have, and a failing ratio is not a verdict, because it may be reacting to a thin sample or a real qualification difference.
| Metric | What it answers | Where it misleads |
|---|---|---|
| Selection rate and percentage-point gap | How many people from each group actually advanced at this stage, and how large the shortfall is in human terms | A given gap means something very different at high base rates than at low ones, so gaps are not comparable across stages |
| Disparate impact ratio (four-fifths rule) | Whether the shortfall meets the regulatory benchmark for adverse impact | Unstable on small samples, and it says nothing about whether the criterion causing the gap is job-related |
| Equity analysis on non-demographic dimensions | Which neutral-looking criterion, school, region, or industry background may be producing the demographic pattern | Correlation between a dimension and an outcome does not identify the mechanism, only a place to look |
| Equal opportunity (true-positive-rate parity) | Whether qualified candidates from each group advanced at similar rates | Depends entirely on a defensible definition of "qualified," which can itself encode the bias you are hunting |
Monitoring by Funnel Stage
A single funnel-wide number tells Marcus almost nothing about where to intervene. So his monthly fairness dashboard breaks the metrics out by stage: application to resume screen, screen to interview, interview to offer, and offer to accept. For each stage and each group he records the selection rate, the impact ratio against the reference group, and the sample size.
The pattern he hunts for is the stage where a ratio first drops below threshold, because that is the stage where the practice introducing the disparity lives. In his system the impact ratio held above 0.80 at every stage except screen-to-interview, which isolated the problem to a single rubric one panel was applying inconsistently. Without stage-level monitoring, the funnel-wide ratio looked acceptable and the rubric problem would have stayed invisible. He also tracks the metrics over time, not as a one-time snapshot, because AI tools drift, req mixes change, and a process that was fair last quarter can slip without anyone touching it.
Setting Thresholds and the Regulatory Context
Thresholds turn metrics into decisions. Marcus uses a two-tier rule. An impact ratio between 0.80 and roughly 0.90 is a "watch" zone: log it, monitor the trend, look for early drift. A ratio below 0.80, or any single-stage selection-rate gap above 5 percentage points that persists across a meaningful sample, triggers a formal review with a documented owner and a deadline. The thresholds are written into policy so that a flag produces an action, not a shrug. A metric that nobody is obligated to act on is just decoration, and a fairness policy that is never enforced teaches the organization that fairness reporting is a formality.
The regulatory backdrop sharpened all of this. New York City's Local Law 144 requires that automated employment decision tools undergo an independent bias audit before use, that the audit compute selection rates and impact ratios across sex and race or ethnicity categories, and that the results be published. Because Marcus's resume-ranking tool touches candidates who could be hired into covered roles, his team treats the four-fifths impact-ratio computation not as an internal nicety but as a disclosure they must be able to stand behind publicly. Other jurisdictions are moving in the same direction, so he designs his monitoring to satisfy the strictest standard he is exposed to rather than the loosest, and keeps the underlying counts so an external auditor could reproduce every ratio.
Causation Versus Correlation
Suppose the dashboard produces a fairness gap. Does the gap indicate bias, or a legitimate difference? This is the question that separates a competent fairness program from an anxious one, and the honest answer is that the metric cannot tell you. Perhaps women in Marcus's candidate pool have less experience on average for a particular senior role. If he is selecting on job-relevant qualifications, and experience is genuinely required to do that job, that is not bias, and forcing the rates to converge would mean advancing people who cannot do the work. The instruction is to investigate causation, not just correlation.
What makes this dangerous is that the same sentence is also the most common excuse for doing nothing. "The pools are just different" is the reflex explanation for every gap, and it is unfalsifiable unless someone does the work. So Marcus holds the explanation to a standard. The qualification has to be named specifically, its distribution across groups has to be shown from the actual data rather than asserted, and the qualification itself has to survive the job-relatedness question: does this requirement actually predict performance in the role, or is it a habit inherited from how the job was always posted? Years-of-experience thresholds, degree requirements, and continuous-employment preferences frequently fail that test, at which point the gap they produce is bias wearing the costume of a qualification.
When a Metric Flags Disparity
A flag is the beginning of work, not the conclusion. When the impact ratio drops below 0.80 at a stage, Marcus runs a fixed sequence. First, verify the data: a miscoded stage or a duplicated applicant record produces phantom disparities, and he checks the counts before he checks anything else. Second, ask whether the gap reflects a job-related difference or a process artifact, applying the causation discipline above rather than accepting the first plausible story. Third, if the gap is not justified by a job-related factor, find the specific practice introducing it: a screening criterion, a rubric item, an AI feature, a sourcing channel that skews the inbound pool.
Crucially, Marcus does not "fix" a disparity by setting group-specific selection quotas, which can create new legal exposure. Instead he repairs the practice: rewrite the criterion to measure the job rather than a proxy, retrain or recalibrate the rubric, adjust the AI tool's features or retire it, broaden sourcing so the inbound pool is more representative. Then he re-measures, on a fresh period rather than the one that produced the flag, because a fix that has not been verified against new data is a hypothesis. Every step gets documented with the date, the finding, the action, and the owner, because that record is both the engine of continuous improvement and the evidence that, when someone asks whether his hiring is fair, Marcus can answer with numbers and show his work.
Anti-Patterns
Treating a failed ratio as a verdict. This is reading an impact ratio of 0.67 and concluding that the process is discriminatory, or that the tool must be switched off today. It happens because the number feels definitive and because the four-fifths rule is described as a legal standard, which encourages people to hear it as a legal finding. What goes wrong is that the team skips straight to remediation without diagnosis, changes several things at once, and can never establish which change mattered or whether the mechanism was ever touched. A failed ratio is a flag, not a verdict: it shifts the burden to showing the criterion is job-related and consistent with business necessity, and that burden is met with an investigation, not a reflex.
Treating a passing ratio as an all-clear. The mirror error, and the more comfortable one. A ratio of 0.85 gets reported as "no adverse impact" and the review ends, because a passing number is what everyone in the room wanted and nobody is rewarded for looking harder at good news. What goes wrong is that the metric that passed may be blind to the failure you actually have: demographic parity can look fine while qualified candidates from one group are being passed over, which only equal opportunity would catch; a funnel-wide ratio can pass while one stage fails; and a ratio computed on 14 applicants means almost nothing in either direction. The counter is triangulation with sample sizes attached.
Measuring only at the hire decision. This is computing fairness metrics on offers or hires because that is the outcome that matters and the final-stage data is tidiest. What goes wrong is that a disparity introduced at the resume screen has already shaped the pool by the time the offer decision is made, so the offer stage can look scrupulously even-handed while operating on candidates who were unfairly narrowed two steps earlier. You end up certifying the fairness of the last decision in a chain of unfair ones. The counter is stage-level monitoring read in sequence, asking at each stage what the incoming pool looked like.
Explaining every gap away as a pool difference. This is answering each flag with "the pools are just different" and closing the item. It happens because the explanation is sometimes true, costs nothing to assert, and relieves everyone of work. What goes wrong is that the claim is never tested, so a real bias and a legitimate qualification difference receive identical treatment, and the fairness review becomes a ritual that produces no findings. The counter is to hold the explanation to evidence: name the qualification, show its distribution across groups from the actual data, and demonstrate that it predicts performance in the role. If the requirement cannot survive the job-relatedness question, it is not explaining the gap, it is causing it.
Practice
- Compute selection rates and impact ratios for one high-volume req at every stage. For each group you can responsibly track, divide the number who advanced by the number eligible, at application to screen, screen to interview, interview to offer, and offer to accept, writing the raw counts next to every rate. Then identify the highest-selecting group at each stage, divide every other group's rate by it, and mark each result against 0.80. You should end with a grid, not a number.
- Find your first failing stage. Read the stages in order and identify the earliest point where a ratio drops below threshold or a gap exceeds 5 percentage points. Then describe what practice operates at exactly that stage: which criterion, which rubric, which tool, which reviewer.
- Run an equity analysis on three non-demographic dimensions. Cut selection rates by candidate geography, by prior industry, and by educational background. Apply the same triggers. For any dimension that flags, state whether it plausibly correlates with a protected characteristic in your pool.
- Stress-test one gap against the causation question. Pick a gap someone has already explained as a qualification difference. Name the qualification specifically, pull its distribution across groups from your data rather than from memory, and write one paragraph arguing that the qualification predicts performance in the role. Notice whether you can.
- Write your threshold policy. Define a watch zone and a formal-review zone in numbers, name the owner who receives each flag, and set the deadline by which a review must produce a finding. A threshold with no owner and no deadline is not a policy.
Reflection
- If a regulator asked you today for selection rates and impact ratios by stage for your highest-volume role, how long would it take you to produce them, and could you reproduce the underlying counts?
- Which definition of fairness is your organization actually optimizing for, demographic parity or equal opportunity, and has anyone written that choice down?
- What is your working definition of "qualified," and whose judgment produced it?
- When was the last time a fairness flag in your organization resulted in a documented action with an owner and a date, rather than a discussion?
- Which of your screening criteria would you struggle to defend as job-related and consistent with business necessity if you had to do it in writing this week?
Glossary
- Selection rate. For a group at a stage, the number who advanced divided by the number eligible. The base quantity from which every other fairness metric is built.
- Selection-rate gap. The difference in percentage points between two groups' rates. A 65 percent rate against a 55 percent rate is a 10-point gap; gaps above 5 points warrant investigation.
- Disparate impact ratio. One group's selection rate divided by the favored group's rate. A 45 percent rate against a 60 percent rate gives 0.75.
- Four-fifths rule (80 percent rule). The EEOC Uniform Guidelines standard that a protected group's selection rate should be at least 80 percent of the favored group's rate. A ratio below 0.80 can indicate disparate impact.
- EEOC Uniform Guidelines on Employee Selection Procedures. The federal guidance under which a selection practice can be challenged on its results even absent discriminatory intent.
- Disparate impact (adverse impact). A substantially different outcome for a protected group produced by a facially neutral selection practice, judged on results rather than motive.
- Reference group. The most-selected group at a stage, against which every other group's rate is compared. It can change between stages and periods.
- Demographic parity. The definition asking whether groups advance at equal rates. The four-fifths rule is a parity test: it needs no qualification data, and for that same reason cannot distinguish a real qualification difference from bias.
- Equal opportunity (true-positive-rate parity). The definition asking whether qualified candidates from each group advanced at similar rates. Sharper than parity, but only as trustworthy as the definition of "qualified."
- Equity analysis. Applying selection-rate comparison to non-demographic dimensions such as geography, school, or prior industry, to locate the neutral-looking criterion behind a demographic pattern.
- Job-related and consistent with business necessity. The justification standard a criterion producing adverse impact must meet. It is the question separating a legitimate qualification difference from bias in a costume.
- Sample-size instability. The tendency of a ratio computed on few applicants to swing on a single decision, which is why every ratio should be published with its denominator.
- Local Law 144. The New York City requirement that automated employment decision tools undergo an independent bias audit before use, computing selection rates and impact ratios across sex and race or ethnicity categories, with results published.
Related Lessons
- Root Cause Analysis: Understanding Why Bias or Errors Occurred is what happens after a ratio fails, working from the finding this lesson produces to the specific mechanism that caused it.
- Data Infrastructure: Collecting, Storing, and Analyzing Recruiting Data builds the collection, storage, and baseline discipline without which none of these metrics can be computed at all.
- Remediation and Escalation: When and How to Act on Findings covers what a threshold breach obligates you to do, including severity classification and verifying that a fix held.
- Sources of Bias: Data, Algorithms, Humans, and Systemic Factors maps the categories of mechanism a flagged metric may be pointing at.
- Auditing AI-Assisted Decisions: Sampling Methodology and Fairness Metrics formalizes how to sample records for review when the population is too large to examine whole.
- Regulatory Landscape: GDPR, AI Act, Executive Orders, and Emerging Standards situates the bias-audit and disclosure obligations in the wider regulatory picture.
Closing
What changed for Marcus was not that his hiring became fair, but that fairness became something he could be wrong about in public and correct in private. Before the metrics, the only available claims were assertions: we train our interviewers, we care about this, we have never had a complaint. After the metrics, he had a monthly page carrying selection rates and impact ratios by stage and by group with sample sizes attached, an equity cut across geography and school and industry, a written threshold assigning each flag an owner and a deadline, and a file of closed investigations recording what was found and what changed.
No one of those instruments is sufficient alone. The selection rate counts, the impact ratio benchmarks, equity analysis locates, equal opportunity conditions on qualification, and causation analysis is what turns any of them into a finding. And when a ratio does fail, the next move is never a quota and never a shrug. It is an investigation into a specific practice, a repair to that practice, a re-measurement on fresh data, and a record with a date and a name on it.
Key Takeaways
- Measure outcomes, not intentions. Disparate impact under the EEOC Uniform Guidelines is judged on results, not motive. Track what actually happened to each group at each decision point and treat your good intentions as irrelevant to the audit.
- Selection rate is the foundation. For each group at each stage, compute advanced divided by eligible. A gap above 5 percentage points, such as 65 percent against 55 percent, is a flag to investigate, not yet proof of bias.
- Use the ratio, not just the gap. Divide each group's rate by the highest group's rate. A 45 percent rate against a 60 percent rate is 0.75 and can indicate disparate impact, while the same 15-point gap at higher rates would clear 0.80, because a shortfall matters more when few people advance at all. A failing ratio shifts the burden to showing the criterion is job-related and consistent with business necessity. Beware small samples, and publish the denominator beside every ratio.
- Extend the arithmetic beyond demographics. Equity analysis across geography, school, and prior industry is how you find the neutral-looking criterion producing a demographic gap, because criteria rarely name a group, they name a proxy for one.
- Demographic parity and equal opportunity are different definitions. Parity compares advancement rates outright; equal opportunity compares rates among the qualified. They can disagree and cannot always be satisfied at once, so choose the trade-off deliberately and in writing.
- No single metric is sufficient, so triangulate. Each instrument has a characteristic blind spot. Report them together with sample sizes, and treat disagreement between them as a question to resolve rather than a result to choose from.
- Monitor by funnel stage and over time. A funnel-wide ratio hides the stage where a disparity is introduced, later stages inherit an already-skewed pool, and a fair process drifts as tools and req mixes change.
- Investigate causation, not just correlation. A gap may reflect a legitimate job-related qualification difference, but that explanation must be held to evidence: name the qualification, show its distribution, and defend its job-relatedness.
- Set thresholds that trigger action. Define a watch zone and a formal-review zone in policy, with owners and deadlines. Local Law 144 requires published bias audits computing selection rates and impact ratios for covered automated tools, so design to the strictest standard you face.
- A flag starts an investigation, not a quota. Verify the data, separate job-related causes from process artifacts, repair the offending practice rather than imposing group quotas, re-measure on fresh data, and document every step.
Frequently Asked Questions
Our impact ratio is 0.78. Are we breaking the law? No, and that is not what the number establishes. A ratio below 0.80 is evidence of adverse impact under the EEOC Uniform Guidelines, which means it triggers an obligation to investigate and, if the disparity is real, to show that the criterion producing it is job-related and consistent with business necessity. A failed ratio is a flag, not a verdict. What would create genuine exposure is having the number, being unable to explain it, and having no record that anyone looked. The productive response is to verify the counts, check the sample size, identify which stage and which specific practice produced the gap, and document the inquiry.
Our reqs are too small for these ratios to mean anything. What do we do? This is a real constraint, not an excuse to skip measurement. The four-fifths rule is unstable with small numbers, and on a req where 12 people from a group applied, one decision moves the ratio dramatically. Aggregate across similar roles and a longer time window so the denominator becomes large enough to interpret, publish the sample size beside every ratio so nobody acts on a flag built on five applicants, and keep recording the per-req counts, because aggregation is only possible if that data exists.
Can we just adjust our selection rates until the ratios pass? No, and this is the most consequential warning in the lesson. Setting group-specific selection quotas to make a metric clear a threshold can create new legal exposure, and it leaves the practice that produced the disparity fully intact. The metric is an instrument for finding a broken practice, not a target to hit. The repair is always to the practice: rewrite the criterion so it measures the job rather than a proxy, recalibrate or retrain the rubric, change or retire the AI feature, or broaden sourcing so the inbound pool is more representative. Then re-measure on a fresh period and document the finding, the action, and the owner. Marcus reviews monthly rather than quarterly for the same reason, since a quarterly cadence lets a disparity run a full quarter before anyone notices, and he keeps the underlying counts rather than only the computed ratios so an external auditor can reproduce the arithmetic.
Skill.re