Continuous Monitoring Fundamentals
Tariq Hassan, a data analyst at a city housing department, was the only person watching the AI tool that flagged rental applications for possible fraud. For four months it hummed along, and Tariq mostly checked that it was online. Then a caseworker mentioned, almost in passing, that the tool had been flagging an unusual number of applications from one ZIP code. Tariq pulled the data. The flag rate for that neighborhood had crept from 8 percent to 31 percent over six weeks, and not because of fraud.
An upstream data vendor had changed how it formatted income fields, and the model was misreading them. Six weeks of residents had been wrongly flagged and sent to manual investigation, delaying their housing. The tool never went down. Tariq was watching, and he was watching the wrong thing, which is the failure this lesson exists to prevent. His six weeks of harm is exactly what good monitoring catches in six hours.
Deployment is the beginning, not the end
Deploying an AI system is not the end of the work. It starts a new phase. Systems degrade, data changes, fairness drifts and accuracy declines, and if you are not monitoring you will not know when any of it begins. Continuous monitoring for AI is the ongoing practice of checking that a deployed model is still doing what it should, before its mistakes reach the public. It is not "is the server up," though that matters too. It is monitoring of quality, fairness and decision integrity, and as an analyst or project lead you are usually the one who has to build it.
The public-sector stakes are what make this urgent rather than tidy. Government AI systems serve the public, so when they degrade, citizens absorb it. A hiring system that becomes biased unfairly excludes people. A benefits system with declining accuracy denies legitimate claims. A criminal justice system with accuracy drift makes worse predictions about real lives. Without monitoring, those degradations can persist for months or years before anyone notices. Tariq's ran for six weeks and was found by accident, in a passing remark from a caseworker.
Monitoring is also how you demonstrate responsible governance to the people who will eventually ask. Congressional committees, inspectors general and civil rights agencies want evidence that systems are being maintained responsibly, and monitoring dashboards and alert logs are the form that evidence takes. It is required as well as expected: OMB memo M-24-10 requires continuous monitoring of AI systems, and the NIST AI Risk Management Framework includes monitoring as a core practice. Where those requirements apply, continuous monitoring is not optional. The honest way to put it is blunter than the policy language: deploy with monitoring, or do not deploy.
What you actually monitor
"Monitor the AI" is too vague to act on. There are four distinct signals, and each catches a different kind of failure. Watch only uptime, as Tariq did, and you will miss three of the four.
- Accuracy. Is the model still right? Are its predictions matching reality when you check them against verified outcomes?
- Drift. Has the input data changed? The income-format change was pure drift, the inputs shifting under a model trained on the old format.
- Bias. Are errors landing unevenly across groups? The ZIP-code spike was a fairness alarm, one neighborhood absorbing the harm.
- Performance. Is the system fast and available? Latency, errors, uptime. The traditional IT signal, and the only one most agencies watch.
The order matters. Performance is the easiest to monitor and the least likely to harm the public when it slips. Bias is the hardest to monitor and the most likely to cause real damage. Most agencies spend all their attention on the easy, low-harm signal and none on the hard, high-harm ones. Tariq's department was textbook, and so is almost every department that has never been forced to think about it.
There is a fifth family that agencies forget entirely, because it is nobody's job: usage and outcome metrics. How many decisions is the system processing? Is it being used the way it was intended? Are staff actually using it? And what happens downstream to its recommendations? A hiring AI approved for 100 positions is a case in point. Track how many of its recommendations led to actual interviews, and if that rate drops from 80% to 40%, the model may be fine while its credibility with the humans has collapsed. That is a real failure, and no accuracy metric will show it to you.
Monitoring accuracy when you do not know the right answer
Accuracy seems simple: compare the model's prediction to the truth. The catch in government is that you often do not know the truth at the moment of prediction. The fraud tool flags an application as risky, but you only learn whether it was truly fraud weeks later, after an investigation, if ever. Two practical techniques get you around that.
- Sampled human review. Have a person check a random sample of the model's decisions each week and record whether the model was right. A sample of 50 to 100 cases is usually enough to spot a real accuracy drop. This is your ground truth, with an important limit: a sample that size will surface a large decline and can easily miss a small one, and a clean week is evidence rather than proof. Treat a good sample as the absence of an alarm, not as a certificate.
- Proxy signals. Watch numbers that move when accuracy moves, even before confirmed outcomes exist. A sudden jump in the flag rate, a shift in the model's confidence scores, or a spike in caseworker overrides all signal that something changed.
Tariq's flag rate jumping from 8 to 31 percent was a screaming proxy signal. He had the data the whole time. He simply was not looking at it as a monitoring signal, because nobody had told him it was one. That is the ordinary shape of this failure: the number that would have saved you is usually already being collected.
Accuracy is also not one number. Track the overall trend, so you can tell whether the system is becoming more or less accurate over time. Track accuracy by demographic group, so you catch degradation that is happening to one population while the aggregate holds. Track it by case type, since accuracy can collapse for one category of decision while everything else is fine. And watch the shape of the change, because a slow three-week slide and a sudden overnight drop have different causes and different responses. A worked example: track overall accuracy weekly against a target of 85%, and alert if it falls below 82% or shows a consistent downward trend.
Watching for drift in the inputs
Drift monitoring asks a question that does not need the right answer: does today's input data look like the data the model was trained on? You can watch this continuously without waiting for outcomes, which makes it the earliest warning available to you. The practical move is to track the distribution of key inputs over time. What is the average and spread of the income field this week against the training baseline? What share of applications are missing a field? Are outliers appearing, or are whole categories disappearing?
Drift comes in three kinds, and naming them helps you pick the response. Covariate shift is a change in the input distribution, as when a pandemic shifts the economic patterns behind loan applications. Label shift is a change in the output distribution, as when a policy change moves hiring preferences. Concept drift is a change in the relationship between inputs and outputs, as when the skills a job actually requires move underneath a model that learned the old ones. All three degrade accuracy, and only the first is visible by looking at inputs alone.
Consider a benefits eligibility model trained in 2022. In 2023 a policy change alters the eligibility rules and the demographics of applicants shift with them. Incoming data no longer matches the training distribution, accuracy drops because the model was never trained on the new one, monitoring detects the drift, and the team retrains on 2023 data. That is the loop working. When drift is detected the response is either retraining on newer data or adjusting the operational process to match the new conditions, and detection is what makes either choice available.
When the income field's format changed, its distribution would have shifted overnight, days before the wrongful flags piled up. A weekly comparison of input distributions against a baseline, with an alert when they diverge past a threshold, would have caught Tariq's problem at its source. Know what that check cannot see, though. Drift detection catches changes in the shape of the data. A change that preserves the distribution while corrupting its meaning will pass straight through it, which is why drift monitoring is the earliest signal and not the only one.
Monitoring for bias after launch
A model that was fair at launch can become unfair in production, because drift and changing conditions do not affect every group equally. Bias monitoring means tracking your key metrics broken down by group rather than in aggregate. The single most important habit is this: never look at a citywide number without also looking at it by neighborhood, by demographic group, by language preference, and by whatever protected or vulnerable dimensions apply to your program.
Tariq's overall flag rate barely moved, because one neighborhood's spike was diluted across the whole city. The citywide number was reassuring and wrong. Only the per-ZIP-code breakdown revealed the harm. This is the rule that would have saved him, and it is worth memorizing: the average is where bias hides. A dashboard that shows only citywide averages is not a monitoring tool. It is a place for harm to hide in plain sight.
Four fairness measures are worth breaking out by group. Approval or acceptance rates, so you can see whether groups are approved equally often. False positive rates, so you can see whether one group absorbs more false alarms, which is exactly what Tariq's neighborhood was experiencing. False negative rates, so you can see whether one group is being wrongly rejected more often. And calibration, meaning whether the model's confidence scores are equally reliable across groups, since a score that means one thing for one population and something else for another will quietly mislead every human who relies on it.
This is not optional for government. The Office of Management and Budget's 2024 AI guidance, memo M-24-10, requires agencies to monitor rights-impacting and safety-impacting AI both for performance and for disparate impacts on protected groups. The NIST AI Risk Management Framework places ongoing measurement, including fairness, at the heart of trustworthy AI. Disaggregated bias monitoring is how you actually do what those documents require, and it is also the only version of monitoring that would have caught what happened to Tariq's residents.
Operational and usage signals
The operational family is the one agencies already know how to run, and it is still worth specifying rather than assuming. Watch system availability and uptime, processing time, error rates, and the quality of the incoming data itself. Concrete targets make these useful: a system uptime target of 99.5% with an alert if availability drops below 99%, and a processing time target under 24 hours with an alert if the median exceeds 48 hours, which is to say if the work starts taking twice as long as intended.
Data quality belongs in this family even though it feels like someone else's problem, because it is the most common upstream cause of the failures in every other family. Tariq's incident was a data quality event that presented as a fairness event. The share of records arriving with missing or malformed fields is one of the cheapest signals to compute and one of the most diagnostic when something goes wrong.
Building the dashboard and the alerts
A dashboard is only useful if someone looks at it, and alerts are only useful if they reach a human who can act. The goal is a single view a non-specialist can read in two minutes, plus automatic alerts so that nobody has to remember to check. The essential elements are current accuracy with its trend against last week and last month, fairness metrics by group with their trends, uptime and error rates, decision volume and processing time, and the alerts for any threshold violations.
Update daily for high-stakes systems and weekly at minimum for everything else. Then make the whole thing visible beyond the data science team. Operations staff, policy staff and leadership should all see these metrics weekly, because visibility is what creates the attention and the accountability that make monitoring real. A dashboard that only its author reads has the same practical effect as no dashboard at all, and it takes longer to build.
For each signal, define three things: the metric, the baseline that describes normal, and the threshold that triggers an alert. An alert should name a person, state plainly what changed, and say what to do first. "Flag rate for ZIP 90011 is 31%, baseline 8%, the twice-citywide threshold is exceeded; pause auto-flagging for that ZIP and notify the data lead" is something a person can act on at 8 a.m. "Anomaly detected" is not.
The monitoring plan and dashboard spec
Build one of these per production model. Each row becomes a dashboard panel and an alert rule. The worked column is Tariq's fraud tool, specified the way it should have been from the start.
| Signal | Metric | Baseline | Alert threshold | First action, and who |
|---|---|---|---|---|
| Accuracy | Model-correct rate on a weekly 75-case human sample | About 92% | Drops below 85% | Data lead reviews the sample for an error pattern |
| Accuracy proxy | Citywide flag rate | About 8% | Moves more than 3 points in a week | Analyst investigates the cause |
| Drift | Income-field distribution against the training baseline | Stable | Distribution shifts past a set divergence | Check the upstream data feed for a format change |
| Drift | Share of records with missing or malformed fields | Under 2% | Exceeds 5% | Analyst checks with the data vendor |
| Bias | Flag rate by ZIP code and language group | Within 1.5x of the citywide rate | Any group exceeds 2x the citywide rate | Pause auto-flag for that group; notify data lead and program director |
| Performance | Uptime and response latency | 99.5% and under 2 seconds | Below 99% or over 5 seconds | Notify operations and the vendor |
| Usage | Share of flags that survive caseworker review | Set from the first stable month | Sustained move in either direction | Analyst checks whether the model or the reviewers changed |
Run Tariq's incident against that table and it fails three ways at once, which is what a good spec should do. The flag rate climbed from 8 percent to 31 percent across six weeks, a pace that clears a three-point weekly trigger. The affected neighborhood sat at 31 percent against a citywide 8 percent, well past a two-times trigger. And the income-field distribution shifted the day the vendor changed its format. Any one of those alerts fires within days. He had none of them.
Thresholds that mean something
Thresholds are where monitoring either becomes operational or stays decorative. For accuracy, a workable pattern is a target with a warning level and an alert level beneath it: target 90%, warning at 87%, alert at 85%. Add a rule that fires when accuracy for any demographic group drops below 82% even while the overall figure looks fine, one for a consistent downward trend across three weeks, and one for a drop confined to a single decision type.
For fairness, set an explicit disparity limit, such as an approval rate gap between groups exceeding 10 percentage points, plus rules for a disparity in false positive rates, for confidence scores becoming miscalibrated for any group, and for any dramatic change in outcomes for a particular group. Write the number down. An unwritten fairness threshold is not a threshold, and the difference between an 8-point gap and a 10-point gap is a judgment somebody has to have made in advance rather than in the middle of an incident.
For operations, alert when downtime exceeds an hour in any day, when processing time doubles from a 24-hour target to 48 hours or more, when the error rate exceeds 1% of requests, or when the data quality score falls below what you agreed was acceptable. Then set all of these against the actual risk profile of the system. A criminal justice system warrants stricter accuracy thresholds than a hiring support tool, and copying one system's numbers onto another is how a threshold ends up meaning nothing.
What happens when an alert fires
Alerts are only useful if somebody responds, and the response should be a defined sequence rather than an improvisation. Eight steps cover it, and writing them down in advance is what keeps the third one from being skipped at 8 a.m. on a Friday.
- The alert triggers. A metric crosses its threshold and reaches a named person.
- Initial investigation. Is it real or a false alarm? Did the metric genuinely move, or is this measurement noise? Look at the data behind the alert rather than at the alert.
- Impact assessment. If it is real, how many cases are affected, how far has quality degraded, and is this urgent or can it wait?
- Root cause analysis. What caused it? Did the training data change, did the input distribution shift, did downtime corrupt state, or did a business process change upstream?
- Determine the fix. Retrain the model, update the configuration, repair a data quality problem, or address the process change.
- Implement the fix. Urgent problems within hours, less urgent ones within days, with the difference decided at the impact assessment rather than by whoever is free.
- Validate the fix. Verify it actually worked. Is accuracy restored? Has fairness recovered? If not, keep investigating rather than closing the ticket.
- Update the thresholds if needed. Did this threshold produce a false alarm, or fire too late? Adjust it and record why.
Attach a clock to the first step. When a metric exceeds a threshold, someone should be investigating within 24 hours, and the alert should say who that someone is. Review the alert history monthly to see which alerts fired and what happened next, and revisit the thresholds themselves quarterly against real experience and actual risk. A threshold left untouched for a year is either producing constant false alarms that everyone now ignores or sitting so loose that it will never fire at all.
Three investigations, three different causes
A federal agency's hiring AI held 88% accuracy for six months. Then the monthly figures read 88, 87, 86 and 84, and the fourth month crossed the 85% alert threshold. Investigation found that the agency had changed its hiring criteria eight weeks earlier, emphasizing different skills, while the model was still recommending for the old skill set. The team retrained on recent data reflecting the new criteria, accuracy came back to 87%, and the training data refresh schedule moved from annual to quarterly. The lesson is that organizational changes cascade into models, so you have to monitor business process changes alongside system metrics.
A state benefits system started with excellent fairness: Group A approved at 75%, Group B at 74%, a disparity of one percentage point. Six months of monthly monitoring later, Group A stood at 76% and Group B at 68%, a disparity of eight percentage points. Investigation found data quality degradation in Group B records, which were increasingly incomplete because the intake form and its requirements had changed, so the model was rejecting Group B applicants for missing information rather than on the decision criteria. Fixing the collection process and retraining on better data recovered the metric.
That second case is also a lesson in why the threshold number matters. An eight-point disparity does not exceed a ten-point rule. Whether that widening gap raises an alarm on the day it happens depends entirely on where the limit was set and written down beforehand, and a team that has not made that decision in advance will make it under pressure, while looking at the result.
A permit system averaged 3.2 days of processing time for twelve weeks, then reported 4.1 days in week 13 and 5.8 days in week 14 against a four-day alert threshold. Note where the line was actually crossed: week 13 already exceeded four days, so the alert should have fired then rather than in week 14, and a week of slower service went by because nobody was reading the number against the rule. The cause turned out to have nothing to do with the model. System performance was unchanged and volume was unchanged, but two specialists had left and their replacements were not yet fully trained. Operational metrics monitor the whole decision process, not just the AI inside it.
Monitoring is a job, not a setting
The last and most common failure is treating monitoring as something you configure once. Dashboards go stale, alert thresholds drift out of usefulness, and the person who built it moves on. Monitoring needs a named owner, a weekly rhythm of actually reviewing the signals rather than only the alerts, and a place where that review is reported so it cannot quietly lapse. The cheapest insurance in all of AI operations is one analyst spending an hour a week genuinely looking at disaggregated numbers.
After his ZIP-code lesson, Tariq rebuilt the tool's monitoring around all four signals, with the flag rate broken out by neighborhood and language group front and center. Three months later a different upstream change shifted the data again. This time the per-neighborhood drift alert fired the same morning. Tariq paused auto-flagging for the affected group, traced it to the vendor by lunch, and no resident was wrongly investigated. Same kind of failure. Six hours instead of six weeks, because he was finally watching the right things.
Anti-Patterns to Avoid
The first two of these are the ones that produce headlines. The rest are how the first two survive an audit.
- Deploy and forget. The most dangerous approach is deploying a system and assuming it will work indefinitely. Problems accumulate unseen, accuracy degrades silently, bias emerges gradually, and six months later you discover the system has been making bad decisions at scale, by which point hundreds or thousands of people have been affected. Monitoring is not optional. Deploy with monitoring or do not deploy.
- Monitoring theater. Metrics get collected, dashboards exist, and nobody looks. No alert triggers a response and no finding drives an action. This is worse than not monitoring, because it produces the artifacts of diligence without any of the effect, and it will read as evidence of care right up until someone checks what was done about the readings.
- Alerts with no procedure behind them. The alert fires and then nothing happens, because nobody knows whether it is real, who investigates, or what the response is. Define it in advance: who is alerted, within what timeframe they must respond, and who has the authority to change or pause the system.
- Frozen thresholds. A threshold set once and never revisited is either too strict, producing false alarms that train everyone to ignore it, or too loose, so it never fires. Review the alert history monthly and the thresholds themselves quarterly, against real experience rather than intention.
- Monitoring accuracy but not fairness. Fairness drifts exactly the way accuracy does, and a system can be accurate overall while being unfair to a specific group. Skipping fairness monitoring means the harm you are most likely to cause is the one you have chosen not to look for.
- Reading the average. Tariq's citywide flag rate barely moved while one neighborhood's tripled. Any metric worth watching is worth disaggregating, and an aggregate figure presented on its own should be treated as an unanswered question rather than a reassurance.
- Assuming monitoring means detection. Tariq was monitoring. Monitoring shortens the time to detection for the failure modes you chose to watch, and it is completely blind to the ones you did not. When you write the plan, write down what it would not catch, and revisit that list when the system changes.
- Treating a clean sample as proof. A weekly human-checked sample of 50 to 100 cases will surface a large accuracy drop and can easily miss a small or narrow one. A clean sample means no alarm was raised, not that the model is fine, and reporting it as the latter is how a slow decline stays invisible.
- Showing the dashboard as the answer. A dashboard and an alert log demonstrate that you watched. They do not demonstrate that the system was accurate or fair. When an oversight body asks how you know the system works, the answer is the problems you detected and what you did about them, not the existence of the panel.
Practice Prompts
Use a system your agency actually runs, and write the answers down rather than thinking them through.
- Design a dashboard. For a state benefits eligibility AI processing 10,000 applications monthly, define which metrics you would track, how often you would update them, what alert thresholds you would set, and how you would present the results to non-technical leadership.
- Write the response procedure. Your hiring recommendation AI alerts that accuracy dropped from 87% to 82%. Define who is alerted immediately, what investigation happens in the first 24 hours, what data you would examine, what the possible root causes are, and how your response would differ if the cause turned out to be data quality rather than model quality.
- Handle a drift finding. Incoming data has shifted significantly from the training distribution, with features in different ranges and a changed demographic composition. Work out how you would confirm the drift is real, assess its impact on accuracy, decide between retraining, reconfiguring or accepting it, and catch the next one earlier.
- Disaggregate one real metric. Take a single number your agency currently reports about an AI system and break it out by every group dimension your program touches. Note which breakdowns you could not produce, because those are the ones harm can hide behind.
- Write the blind-spot list. For a system you monitor, write down the failure modes your current plan would not catch, and put a date on when you will revisit the list.
Reflection
Answer these about a system you are responsible for, not about monitoring in general.
- If this system started harming one neighborhood tomorrow, which number would move, and is anyone looking at it?
- When was the last time one of our alerts fired, and what actually happened next?
- Could I show an inspector general a problem we detected ourselves and fixed, or only that we have a dashboard?
- Who is the named owner of this monitoring, and what happens to it when they leave?
- Which of our metrics have I only ever seen as an agency-wide average?
Glossary
- Continuous monitoring. Ongoing measurement of a deployed system's performance, fairness and operations, providing early warning of degradation.
- Data drift. A change in the input data distribution over time, which can leave a model trained on older data performing poorly on newer data.
- Covariate shift. Drift in which the input distribution changes, such as an economic shock altering the pattern of applications.
- Label shift. Drift in which the output distribution changes, such as a policy change moving what gets approved.
- Concept drift. Drift in which the relationship between inputs and outputs changes, such as the skills a job actually requires moving over time.
- Alert threshold. A metric value that triggers investigation when crossed, such as an alert when accuracy drops below 85%.
- False positive rate. The share of negative cases incorrectly classified as positive, which produces false alarms and wasted investigation.
- Calibration. Whether confidence scores match actual accuracy, and whether they do so equally across groups. Miscalibration is detectable through monitoring.
- Proxy signal. A metric that moves when quality moves, watchable before confirmed outcomes exist, such as a flag rate or an override rate.
- Disaggregation. Breaking a metric out by group rather than reporting it in aggregate, on the principle that the average is where bias hides.
Related Lessons
Monitoring is the operational half of several disciplines taught elsewhere in this level.
- Systematic AI Output Validation covers the individual checking discipline that the sampled human review in this lesson formalizes into a weekly routine.
- Bias Detection Tools and Methods goes deeper on the fairness metrics you disaggregate here, and on how to compute them properly.
- Quality Assurance for AI Work Products addresses the review practices behind the sample your accuracy metric depends on.
- Human-in-the-Loop: Design and Implementation covers the caseworker review step that both catches the model's errors and produces your override signal.
- AI Incident Documentation and Response picks up where an alert becomes an incident, including what you owe the people already affected.
- Building an AI Quality Culture is what keeps the weekly review from lapsing once the person who cared about it moves on.
Closing
Continuous monitoring turns deployment from a one-time event into ongoing stewardship. You deploy, then you watch, you measure, you detect problems, and you respond. Agencies that do this keep their systems healthy. Agencies that do not degrade gradually and eventually fail in public, and the difference between a trusted government AI system and a failed one is frequently nothing more sophisticated than vigilance applied to the right numbers.
It is also how public trust is maintained rather than asserted. When an oversight body or a resident asks how you know your system is working fairly, you can answer with the dashboard, the monthly metrics, and specific examples of problems you found and fixed. That last item is the one that carries the weight. A monitoring program that has never caught anything is either watching a perfect system or watching the wrong things, and Tariq's six weeks are what it costs to find out which.
Key Takeaways
- Watch four signals, not one. Accuracy, drift, bias and performance each catch a different failure, and uptime alone misses three of the four. Usage and downstream outcomes make a fifth that almost nobody tracks.
- Spend attention where the harm is. Performance is easy and low-harm; bias is hard and high-harm, so it deserves the most monitoring rather than the least.
- Use proxy signals and sampled review for accuracy. When you cannot know the right answer immediately, a weekly human-checked sample plus a watched flag rate stand in for ground truth, and a clean sample is evidence rather than proof.
- Catch drift at the inputs. Comparing input distributions to a training baseline warns you days before wrong outputs reach the public, and knowing the three kinds of drift tells you which response is available.
- The average is where bias hides. Never read an agency-wide number without the per-group breakdown, and disaggregate approval rates, false positives, false negatives and calibration.
- Write the thresholds down in advance. A disparity limit or accuracy floor decided during an incident is decided while looking at the answer. Set target, warning and alert levels, and match them to the system's risk profile.
- Make alerts actionable. Every alert should name a person, state what changed against the baseline, and say what to do first, with someone investigating within 24 hours.
- Define the response before you need it. Trigger, verify, assess impact, find the root cause, choose the fix, implement, validate, and adjust the threshold. Skipping validation is how a fix that did not work gets closed as one that did.
- Monitoring is not detection. It shortens the time to find the failures you chose to watch for and is blind to the rest. Write down what your plan would not catch.
- Monitoring is a staffed job. Assign an owner, review the signals weekly, report the review, revisit thresholds quarterly, or the dashboard goes stale and the harm comes back.
Frequently Asked Questions
We do not have a data science team. Can we still do this? Most of it, yes. The highest-value pieces are a weekly disaggregated flag or approval rate, a share-of-records-with-missing-fields check, and a small human-reviewed sample, and all three are spreadsheet work rather than machine learning work. What you cannot improvise is the ownership: someone has to have the hour, on the calendar, with the standing to raise what they find. Start with the numbers you already collect, because in most incidents like Tariq's the signal was already in the system and nobody had been told it was a signal.
How do I set a fairness threshold when I do not know what disparity is acceptable? Set a provisional one, write down the reasoning, and involve legal or civil rights staff before it becomes the operative rule, because a disparity limit is a policy judgment with legal weight and not a technical parameter. What matters most immediately is that a number exists in writing before an incident, since a threshold agreed while staring at a live disparity will be argued about rather than applied. Then revisit it quarterly with the alert history in front of you.
Our alerts fire constantly and everyone ignores them. What now? That is a threshold problem being experienced as a culture problem, and it is fixed at the threshold. Pull the alert history for the last few months, separate the alerts that turned out to be real from the ones that were noise, and loosen the rules that produced only noise while keeping the ones that caught something. Then set a review cadence so this does not recur. An alert nobody acts on is worse than no alert, because it produces a record showing the system warned you.
The model looks fine but caseworkers have stopped trusting it. Is that a monitoring issue? Yes, and it is the family most agencies never instrument. Track what happens to the system's recommendations downstream: how many are accepted, how many survive review, and whether that rate is moving. A collapse in acceptance can mean the model degraded, and it can equally mean a change in staffing, training or workload, which the model metrics will never show you. Either way it is a real failure of the system as deployed, and it is invisible if you only measure the model.
We found a problem that ran for weeks. Do we have to tell anyone? Treat that as an incident rather than a maintenance item, because people were affected while it ran. Establish who was affected and over what period, since that is the question everyone will ask and it is much harder to reconstruct later. Then follow your agency's incident process for notification and remediation rather than deciding case by case what deserves reporting. Agencies are judged on what they did after they found out, and a self-detected problem that was reported and fixed is a far better record than one that was found for you.
Skill.re