Anomaly Detection: Identifying Unusual Patterns That Signal Problems
Hana is a recruiting-operations analyst at a mid-market logistics company, and she owns the fairness dashboard for an AI resume-screening tool that processes roughly 8,000 applications per quarter across six job families. For the first year, her job was to pull a report every quarter, eyeball the pass-through rates by demographic group, and present a slide to the talent-acquisition VP. The problem with that rhythm became obvious the day the vendor pushed a silent model update in week two of a quarter: by the time Hana ran her quarterly report ten weeks later, the tool had already screened more than 5,000 applications under the new behavior. The damage was done before anyone looked. Periodic reporting tells you what happened last quarter; continuous anomaly detection tells you what is happening now, while you can still do something about it.
From Periodic Reports to Continuous Monitoring
Most recruiting teams treat fairness as a reporting exercise: someone runs the numbers monthly or quarterly, files a slide, and moves on. That model has a fatal latency problem. An AI screening tool can shift behavior overnight, whether from a vendor model update, a change in the applicant mix, or a configuration tweak, and a quarterly cadence means weeks of biased decisions accumulate before a human notices. Anomaly detection flips the posture: instead of asking what the numbers were last quarter, you ask whether anything unusual is happening today and whether it should page someone. The goal is shrinking the gap between when a problem starts and when Hana finds out, from ten weeks down to a few days.
It helps to be precise about how this fits the legal landscape. NYC Local Law 144 requires employers using automated employment decision tools to commission an independent bias audit annually. That audit is a point-in-time obligation: it certifies the tool's impact ratios as of a given date. Continuous monitoring does not replace it, and Hana should never pitch it as a substitute. The annual audit is the legal floor; her monitoring is the early-warning system catching drift in the months between audits. Frame the two as complementary, because a regulator or counsel will expect both.
Baselines: Deciding What Normal Looks Like
Anomaly detection is comparison, and comparison needs a reference point. Before a new tool goes live, establish baseline fairness metrics and write them down: the current selection rate for each demographic group, the current disparate impact ratio, time to decision, and whatever quality measures you already track. Without a baseline, every number the dashboard produces is simply the number, and no measurement can be called unusual because there is nothing for it to be unusual against.
A concrete baseline reads like a fact, not an impression. Suppose that before a screening tool is deployed, 40 percent of female candidates advance from screening and 45 percent of male candidates advance. The impact ratio is 40 divided by 45, or 0.89. Those three figures are the baseline. After the tool goes live you measure again and find 35 percent of female candidates advancing and 42 percent of male candidates advancing, a ratio of 0.83. Something has moved. Whether it moved for a bad reason is a separate question you cannot answer from the ratio alone: the pool may have changed, hiring managers may be using the tool differently, or it may be noise.
Notice what this example does and does not trigger. The ratio has fallen from 0.89 to 0.83, which is a real deviation from baseline worth investigating, and yet 0.83 is still above the four-fifths threshold discussed below. A monitoring system that watches only for four-fifths breaches would have said nothing here. That is the whole argument for baselines: they catch degradation on its way down, rather than at the moment it crosses a legal tripwire and becomes expensive.
Baselines belong in governance documentation, written so a successor could use them without asking you anything: the tool covered, the selection rate for each group, the impact ratio, the period the baseline was established over, the number of candidates it was computed from, and the date of the next review. Quarterly review is a reasonable default, because a baseline built against last year's applicant mix slowly stops describing this year's. Budget two to four weeks of data collection before deployment, plus analysis and documentation time, and validate the result with whoever owns fairness and legal review before you alert against it.
What to Monitor: The Signals That Actually Matter
Hana cannot monitor everything, so she monitors the handful of signals that move first when fairness degrades. The most important is the pass-through rate by group and by stage: of the candidates from a given group who enter resume screening, what fraction advance. She tracks this separately at each funnel stage, because a tool can look fine in aggregate while quietly suppressing one group at a single gate. Group here means the categories your audit and compliance obligations already use, typically sex and race or ethnicity, with age tracked where you have the data to do it responsibly.
The second signal is the score distribution. Most screening tools emit a numeric score, and the shape of that distribution for each group is a leading indicator: if the median score for one group slides while another holds steady, that shows up here before it shows up in pass-through counts. The third is sudden shifts after a known event, especially a vendor model update. Hana keeps a change log, and any date a vendor pushes a new model version becomes an annotation on her charts so she can see whether the line moved at that exact point. The fourth is drop-off spikes: an unusual jump in candidates abandoning the application at a particular step, which can signal a broken or unfair gate rather than a fairness ratio as such.
Alongside those four, keep the disparate impact ratio itself, time to decision by group, and whatever candidate-experience measure you collect, because a fairness problem often shows up as differential treatment in speed or communication before it shows up in outcomes. Cadence matters as much as the metric list: a dashboard that refreshes daily or weekly is a monitoring system, while one that refreshes quarterly is the report you already had. None of this requires exotic tooling. Hana built the whole thing as a small set of control charts in her BI tool, fed by a weekly export from the applicant tracking system.
Simple Statistical Signals You Can Trust
Hana resists the temptation to get fancy. Three simple signals carry most of the weight. The first is the control limit. On a control chart, she plots the weekly pass-through rate for each group against a baseline mean with upper and lower control limits set at roughly two to three standard deviations of normal week-to-week variation. A single point outside the limits, or a run of several consecutive points drifting one way, is the classic signal that something has changed beyond ordinary noise. A common alerting rule sits tighter than the limits themselves, flagging any metric that deviates more than one to two standard deviations from baseline, because it is meant to prompt a look rather than a conclusion.
The second is the week-over-week delta. A pass-through rate that wobbles by a point or two between weeks is normal; her applicant pools are not identical week to week. A jump of five or more points in a single week, especially in a high-volume job family, is worth a look even if it has not yet broken a control limit. The third, and the one with the most legal weight, is the four-fifths threshold. Under the four-fifths, or 80 percent, rule used in adverse-impact analysis, you compare the selection rate of the lowest-passing group to that of the highest-passing group; a ratio below 0.80 is the traditional tripwire suggesting potential adverse impact under Title VII. Hana treats any week where that ratio dips below 0.80 as an automatic anomaly that fires an alert, regardless of what the control chart says.
A word of caution she learned the hard way: the four-fifths rule is a monitoring tripwire, not a verdict. A ratio under 0.80 does not by itself prove unlawful discrimination, and a ratio above 0.80 does not guarantee a tool is clean, especially at small sample sizes where the ratio swings wildly. The rule's job is to tell her where to point an investigation, not to end one. The same caution applies to the deviation rules: a two-point drop from baseline is a reason to look, not a finding.
Setting Thresholds and Managing False Positives
The hardest part of anomaly detection is not detecting anomalies. It is detecting the right ones. Set thresholds too tight, alerting on every half-point change, and Hana drowns in alerts, the team learns to ignore them, and the one real signal hides in the noise. Set them too loose, waiting for a ten-point move, and a genuine problem slides past for months. Her calibration approach is to back-test against historical data: replay the last four quarters through a candidate threshold and count how many alerts it would have fired and how many were real. If almost every alert is noise, the threshold is too tight.
Calibration is not a one-time exercise, so track alert accuracy as a running metric. A quarter that produced 100 alerts, of which 70 warranted an investigation and 20 confirmed a real issue, tells you something a raw alert count never will. Roughly 20 to 30 percent of alerts confirming real issues indicates good calibration. Far below that band and you are training your own team to ignore the dashboard; far above it and you are almost certainly missing problems that never reached the threshold. Review the number quarterly alongside the baselines, and change the threshold rather than exhorting people to take alerts more seriously.
She also separates severity tiers so not every signal pages a human at midnight. A four-fifths breach or a control-limit violation in a high-volume family goes to her and the TA VP immediately. A modest week-over-week delta in a low-volume family that screens forty people a quarter is a low-severity note reviewed in her weekly sweep, because at forty applications the numbers are too small to trust a single week. Tying alert severity to sample size is the single biggest false-positive reducer she found: small samples produce dramatic-looking ratios that mean almost nothing.
Stakes deserve the same treatment as volume. A tool that drives a final hiring decision warrants tighter thresholds and faster escalation than one that only routes applications into a queue, because the consequence of an undetected shift is different in kind. Set the threshold per tool, write the reasoning next to it, and expect to defend that reasoning to an auditor who asks why two tools were watched differently.
Writing Triggers Down Before You Need Them
Thresholds decide what the dashboard notices. Triggers decide what a human is obliged to do about it, and they need writing down in advance for one reason: otherwise the decision to investigate gets made by whoever is looking, on whatever kind of week they are having. Anomalies surfacing during a quiet fortnight get chased; identical anomalies surfacing during a hiring surge get rationalized. Documented triggers remove that discretion, which is what makes the program consistent enough to describe to an auditor.
A workable trigger list is short, specific, and covers more than the headline ratio. Investigate if any demographic group's selection rate changes by more than two points from baseline. Investigate if the disparate impact ratio drops below 0.80. Investigate if the time-to-decision metric rises by more than 20 percent for any group, because differential speed is differential treatment even when outcomes match. Investigate if candidate feedback raises a fairness concern, which is a qualitative trigger that catches things no metric was designed to see. And investigate if a pattern emerges across multiple tools, since a shift appearing in two independent systems at once usually points at the data or the process rather than at either model.
A Standard Investigation Process
When an alert fires, run the same six steps every time. Confirm the anomaly is real rather than noise. Examine the affected group and time period closely rather than reasoning from the aggregate. Collect context: candidate pool demographics for the period, hiring manager behavior, tool settings and version. Generate hypotheses about root cause, plural, because the first explanation that occurs to you is usually the one you already believed. Test each against the data. Document the findings, including the hypotheses you ruled out, since that record is what makes the conclusion defensible six months later.
Here is that process on an anomaly where the tool turned out to be innocent. An alert fires: the female selection rate has dropped from 40 percent to 35 percent. Step one confirms it is real, consistent across more than 100 candidates rather than a one-week wobble. The first hypothesis is a pool shift, and the data partly supports it: female candidates make up 2 percent less of the June pool than of the baseline period, real but far too small to explain a five-point swing. The second hypothesis is a settings change, and the change log shows none. The third, a change in how humans are using the tool, lands. Hiring managers have started applying the "must have five years of experience" filter far more aggressively; female candidates in this pool average 4.2 years of experience against 4.8 for male candidates, and that gap is what the tightened filter is cutting on.
The root cause is a hiring-manager behavior change, not tool bias, and the recommendation follows from it: clarify that five years of experience is a guideline rather than an absolute requirement, then monitor whether the ratio recovers. This case is worth studying because the remedy has nothing to do with the model. Had the investigation stopped at hypothesis one, the team would have shrugged at a 2 percent pool shift and let a real, human-caused adverse effect run. Anomaly detection watches the system, and the system includes the people operating it.
A Second Worked Example: When the Tool Is the Cause
Here is the other kind of event the system exists to catch. In week two of Q3, the vendor pushes a model update; Hana annotates the date on her charts. In week three, her dashboard fires a high-severity alert on the warehouse-associate job family, the highest-volume family at roughly 1,200 applications a quarter.
The numbers: before the update, candidates from her reference group were passing resume screening at 58 percent, and candidates from the affected group at 52 percent. That gives a four-fifths ratio of 52 divided by 58, or 0.90, comfortably above the 0.80 tripwire. In week three, the reference group held at 58 percent, but the affected group's pass-through dropped to 41 percent. The ratio is now 41 divided by 58, which is 0.71. That is below 0.80, so the four-fifths tripwire fires, and the 11-point single-week drop also blows through the lower control limit on the affected group's chart. Two independent signals agree, on the exact week the vendor model changed.
Hana does not pause the tool on the spot, because a 0.71 ratio is a signal, not a conviction, and she first rules out the obvious confounder. She checks the applicant mix: the affected group's application volume and apparent qualifications look the same as prior weeks, so a pool shift does not explain the drop, and the timing lines up exactly with the vendor update. That correlation, plus a clean confounder check, is enough to act on. She escalates to the TA VP with a one-page summary, the tool is paused for the warehouse family while the vendor investigates, and the vendor confirms the update changed how the model weighted a credential that correlated with the affected group. The fix is rolled back, the ratio returns to its baseline range the following week, and the vendor-update date goes into her permanent change log so the next push gets watched from day one. The whole cycle took eight days instead of ten weeks.
The Response Playbook
An alert is useless without a rehearsed response, so Hana wrote a short playbook tiered by severity. When a high-severity alert fires, step one is confirm, not react: is the signal real, or a small-sample artifact or one-week blip. Step two is rule out the benign explanations, primarily a genuine shift in the applicant pool. Step three is check the change log for any vendor update, configuration change, or new screening rule that lines up with the timing. Step four, if the signal survives, is escalate to a named owner, in her case the TA VP, with a one-page brief and a recommendation. Step five is the containment decision: pause the tool for the affected family, fall back to human review, or tighten the gate, depending on volume and exposure. Step six is document everything, because that record goes to the next bias auditor and, if it comes to it, to counsel.
Response speed should be tiered along with response content, and the tiers written before an alert forces the question. Speed is not a preference here: bias compounds, and a slow response means the tool keeps deciding while you deliberate.
| Severity | What it looks like | Investigation target | Response |
|---|---|---|---|
| Minor | Noise, a one-time anomaly, no bias found on confirmation | Two weeks | Document the finding, keep monitoring, continue operating |
| Moderate | A pattern emerging, potential bias, real but limited impact | One week | Modify the tool settings or how it is used, increase monitoring frequency, plan remediation |
| Major | Confirmed bias, adverse impact, demonstrable harm | 24 to 48 hours | Pause the tool immediately, investigate root cause, remediate, test thoroughly before resuming |
For each tier the procedure must name four things: who decides the response, whether the tool owner, a fairness officer, or an executive sponsor, chosen so authority matches consequence; how quickly the decision must be made, with a stated alert-to-response commitment such as under 24 hours for a major issue and under a week for a moderate one; what escalation is required and to whom, including when legal is brought in rather than informed afterward; and how a change is validated before the tool resumes, because a fix deployed without re-testing has only moved the risk.
The two failure modes she designs against are over-reacting and under-reacting. Pausing the tool on every noisy ratio destroys the team's trust in the system and the recruiters' throughput. Waiting for certainty before acting lets a real problem compound across thousands of applications. The playbook exists precisely so that the response is proportional and pre-decided, not improvised in a panic at the moment an alert lands.
Scaling the System Across Multiple Tools
Most talent functions run several AI tools, and monitoring five the way you monitor one produces five times the alerts and a fifth of the attention per alert. Separate the metrics common across tools from the ones specific to each: selection rate by group and the impact ratio apply to anything that filters or ranks candidates, so define them identically everywhere and compare on a single view, while turnaround time, user satisfaction, and output-quality checks belong on each tool's own panel.
Against alert overload there are three levers: aggregate related alerts so one underlying shift does not fire six times, set thresholds per tool according to stakes and volume, and order the queue by severity rather than arrival time. Decide as well who investigates, whether a single team that builds pattern recognition across tools or tool-specific owners who know their system's quirks. Name it either way, because the failure mode is an alert everybody can see and nobody owns. Then build a path for sharing findings: if one tool penalizes a credential correlating with a protected group, the next question is whether any other tool uses that feature, and the check should be recorded even when it comes back clean.
Anti-Patterns
Deploying without a baseline. The tool goes live and monitoring starts the same week, so the first numbers the dashboard produces silently become the definition of normal. It happens because collecting baseline data delays a launch someone has already announced. What goes wrong is that six months later, when a colleague notices one group's selection rate looks low, nobody can say whether it was always low or fell recently, and the investigation that opens produces weak conclusions because it has nothing to compare against. The counter is two to four weeks of pre-deployment data collection, documented as selection rates, disparate impact, time to decision, and quality by group, stored in governance with a review date attached.
Alert fatigue from thresholds nobody calibrated. The dashboard alerts on every half-point change, so the team receives five to ten alerts a week and finds normal variation behind nearly all of them. It happens because a tight threshold feels conscientious and nobody wants to be the person who argued for looser alerting. What goes wrong is that people quietly stop investigating, and the first genuine problem arrives as one more line in a queue everyone has learned to skim. The counter is ongoing calibration: back-test thresholds on historical data before adopting them, use a deviation rule of one to two standard deviations rather than a fixed tiny delta, and treat an alert-accuracy share far below 20 to 30 percent as a defect in the threshold rather than in the team.
Detecting fast and responding slowly. The alert fires on schedule, the investigation opens, and then it competes with everything else on the analyst's desk. Weeks pass, other issues intervene, and by the time a remediation ships six weeks have gone by and the tool has processed more than a thousand additional candidates under the behavior that triggered the alert. It happens because detection is automated and response is not, so response inherits the priority of whatever else is happening. The counter is the tiered playbook with time commitments written in advance, a named decision owner per tier, and the willingness to pause on a major issue before the investigation is complete rather than after.
Treating the four-fifths ratio as a pass-fail gate. The dashboard shows 0.81 and the team relaxes; it shows 0.79 and the team panics. It happens because a single number with a bright line attached is enormously convenient. What goes wrong runs both ways: a tool degrading steadily but still above the line gets no attention, while a low-volume family whose ratio swings on a handful of applications generates emergencies every other week. The counter is to run the ratio alongside baseline deviation and control limits rather than instead of them, and to weight any ratio by the sample behind it.
Assuming an anomaly means the model is broken. The alert fires, the investigation opens with a single hypothesis, and the team goes straight to the vendor. It happens because the tool is the newest thing in the process and therefore the most suspicious. What goes wrong is that human-caused drift, a filter applied more aggressively or a change in how a req is written, goes uncorrected while everyone waits for a vendor response that will not explain it. The counter is the multi-hypothesis discipline: pool shift, settings change, and user behavior all get tested every time.
Practice
- Establish a baseline for one tool. Specify the baseline project end to end: which metrics, over what period and how many candidates, how demographic data will be collected responsibly, who analyzes, and who validates before it becomes the reference point. Write the finished baseline so a successor could use it unaided, including period, sample size, and next review date.
- Design the dashboard and its alerts. List the metrics, how each is visualized against baseline, and the refresh cadence. Then write the alert rules: which deviation fires, which ratio breach fires automatically, which alerts page a person versus drop into a weekly sweep, and the sample-size floor below which a single week's ratio does not fire at all.
- Write your investigation triggers. Draft the numbered list of conditions that oblige someone to investigate, including one qualitative trigger such as a candidate fairness complaint and one cross-tool trigger. Then ask a colleague to point at any trigger that leaves room for a judgment call about whether it fired, and tighten until none do.
- Build the investigation SOP and the response matrix. Turn the six-step process into a report template covering the anomaly and its confirmation, the context collected, each hypothesis and how it was tested, the surviving root cause, and the recommendation. Then, for each severity tier, specify what qualifies, the investigation target, the response actions, who decides, what escalation is required, and how a change is validated before the tool resumes.
- Scale it to your portfolio and audit your alert accuracy. Inventory every AI tool touching candidate decisions, separate the metrics common to all of them from the tool-specific ones, name who investigates across tools, and write the rule that a confirmed finding on one tool triggers a documented check of the others. Then pull the last quarter of alerts, count how many confirmed a real issue, and compare that share to the 20 to 30 percent band. If you are outside it, propose the threshold change rather than a process reminder.
Reflection
- If your screening tool changed behavior tomorrow, how many candidates would it process before anyone found out, and what is that number based on?
- Can you produce the documented baseline for your most consequential tool right now, including the period and sample size it was built from?
- When did your team last investigate an alert, and would that same alert have been investigated during your busiest week of the year?
- The last time a fairness metric moved, whose behavior did you investigate besides the model's?
Glossary
- Baseline. The documented reference point for a tool's fairness metrics, established before deployment: selection rates by group, impact ratio, time to decision, and quality measures, plus the period and sample size behind them and a scheduled review date.
- Pass-through rate. The fraction of candidates from a given group entering a funnel stage who advance to the next. Tracked per stage, because a tool can look clean in aggregate while suppressing one group at a single gate.
- Disparate impact ratio. The selection rate of the lowest-passing group divided by that of the highest-passing group. The headline fairness number in most bias audits.
- Four-fifths rule. The convention that an impact ratio below 0.80 signals potential adverse impact under Title VII and warrants investigation. A tripwire, never a verdict: a ratio under 0.80 does not by itself prove unlawful discrimination, and a ratio above it does not certify a tool as clean.
- Control chart and control limits. A time-series plot of a metric against its baseline mean with limits at roughly two to three standard deviations of normal variation. A point outside the limits, or a run of consecutive points drifting one way, indicates change beyond ordinary noise.
- Change log. The dated record of vendor model updates, configuration changes, and new screening rules, annotated onto the charts so a shift can be lined up against an event that might explain it.
- Investigation trigger. A written condition obliging someone to investigate: a two-point selection-rate change from baseline, a ratio below 0.80, a time-to-decision increase above 20 percent for any group, a candidate fairness complaint, or a pattern across multiple tools. Written in advance so investigation does not depend on who is looking or how busy the week is.
- Alert accuracy. The share of alerts that confirm a real issue on investigation. Around 20 to 30 percent indicates good calibration; well below that band the thresholds are too tight and the team is being trained to ignore the dashboard, which is alert fatigue.
- Severity tier. The classification of an anomaly as minor, moderate, or major, each carrying its own investigation target, response actions, decision owner, escalation path, and validation requirement.
- Confounder check. The step that rules out benign explanations before acting on a signal, principally a genuine shift in the applicant pool, alongside settings changes and changes in how humans use the tool.
- Automated employment decision tool. The regulatory category for systems that computationally screen or score candidates, and the category to which NYC Local Law 144's annual independent bias audit obligation attaches.
Related Lessons
- Fairness Metrics: Defining and Measuring Bias in Outcomes defines the selection-rate and impact-ratio measures this lesson monitors, and explains what each one can and cannot tell you.
- Hands-On Project: Design a Fairness Monitoring Dashboard is where you build the dashboard, thresholds, and alert routing described here as a deliverable.
- Root Cause Analysis: Understanding Why Bias or Errors Occurred goes deeper on the hypothesis-testing discipline that turns a confirmed anomaly into an explanation.
- Remediation and Escalation: When and How to Act on Findings develops the severity tiers, containment decisions, and escalation paths in the response playbook.
- Data Infrastructure: Collecting, Storing, and Analyzing Recruiting Data covers the exports, storage, and demographic-data handling that make weekly monitoring possible at all.
- Regulatory Landscape: GDPR, AI Act, Executive Orders, and Emerging Standards situates the annual bias-audit obligation that monitoring complements rather than replaces.
Closing
Hana's system is not sophisticated. It is a weekly export, a handful of control charts, a documented baseline, a short list of triggers, and a playbook that says what happens when one of them fires. What made it work was not the statistics but the decisions taken before any alert existed: what normal looks like, what obliges a human to investigate, who decides, how fast, and what has to be true before the tool goes back on.
The parts that will be tempting to cut are the ones carrying the weight. The pre-deployment baseline, because it delays a launch. The alert-accuracy review, because it invites the finding that your thresholds are wrong. The multi-hypothesis discipline, because checking hiring-manager behavior is politically harder than emailing a vendor. And the response time commitments, because they oblige you to pause something during your busiest quarter. Drop those four and you have built a dashboard that reports what already happened, which is the quarterly slide with a faster refresh rate.
Key Takeaways
- Continuous beats periodic. A quarterly fairness report can let weeks of biased decisions accumulate before anyone looks; anomaly detection shrinks the gap between when a problem starts and when you find out, which is the difference Hana measured as ten weeks versus eight days.
- Without a baseline there are no anomalies. Document selection rates by group, the impact ratio, time to decision, and quality before deployment, with the period, sample size, and next review date attached. Baselines catch degradation on the way down, before it crosses a legal tripwire.
- Monitoring complements the required audit, it does not replace it. NYC Local Law 144 mandates an annual independent bias audit for automated employment decision tools; continuous monitoring is the early-warning layer catching drift in the months between audits, and you should frame it that way to counsel and regulators.
- Watch the signals that move first, and refresh often. Pass-through rates by group and stage, score distributions, shifts after a vendor model update, and drop-off spikes change before final outcome counts do. A dashboard refreshing daily or weekly is monitoring; a quarterly one is the report you already had.
- Keep the statistics simple, and remember what the ratio is for. Control limits, week-over-week deltas, a one-to-two standard deviation rule, and the four-fifths ratio cover most real anomalies. A ratio below 0.80 flags potential adverse impact under Title VII and tells you where to investigate; it does not by itself prove discrimination, and small samples make it unreliable.
- Calibrate thresholds against history, sample size, and stakes. Back-test on past quarters, tie alert severity to volume, set tighter thresholds on higher-stakes tools, and track what share of alerts confirm real issues against a 20 to 30 percent band so the team never learns to ignore the dashboard.
- Write triggers down so investigation is not discretionary. Numbered conditions, including a candidate-complaint trigger and a cross-tool pattern trigger, prevent the failure where identical anomalies get chased in a quiet week and rationalized in a busy one.
- Test more than one hypothesis. Pool shift, settings change, and human behavior all get checked every time. A filter applied more aggressively by hiring managers can produce a real adverse effect that no vendor response will explain.
- Rehearse the response before the alert fires. A tiered playbook with time commitments, named decision owners, escalation paths, and a validation step before resuming keeps the reaction proportional and prevents both over-reacting to noise and under-reacting to a compounding failure.
Frequently Asked Questions
Our applicant volume is far smaller than Hana's. Does any of this work at low volume? The structure works; the statistics need honest handling. At forty applications a quarter in a job family, a single week's impact ratio is close to meaningless, because one or two decisions swing it dramatically. Two adjustments help. Lengthen the window, comparing rolling quarters rather than weeks, so each data point rests on enough decisions to mean something. And set a sample-size floor below which a ratio does not fire an alert at all, reviewing those families in a periodic sweep instead. What you must not do is apply a high-volume threshold to a low-volume family and then either panic at the noise or conclude from a clean-looking ratio that a barely measured process is fair.
We do not collect demographic data on every candidate. Can we still monitor for adverse impact? Partly, and the gaps need stating rather than papering over. Where demographic data is voluntarily self-reported, coverage is incomplete and your metrics describe only the responding population, which may not represent the whole. Say so on the dashboard, so nobody reads a ratio computed from a third of applicants as if it covered all of them. The signals that do not depend on demographic data still work and are worth running: overall score distributions, drop-off spikes at particular steps, sudden shifts after a vendor update, and time-to-decision changes. Any change to how demographic data is collected is a question for legal and privacy counsel before it is a question for analytics.
How is this different from the bias audit our vendor already runs? Different instrument, different question, different cadence. A bias audit is an independent, point-in-time assessment certifying impact ratios as of a date, and under NYC Local Law 144 it is an annual obligation for automated employment decision tools, with a published summary and candidate notice attached. Monitoring is internal, continuous, and diagnostic: it watches for change between audits and tells you where to look. Neither substitutes for the other, and a vendor-run audit does not discharge your own obligations. Monitoring is what makes the next audit unsurprising, because you already know what moved during the year and why.
Skill.re