←
AI for Government
Capable · M3 · lesson 3 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI for Compliance Monitoring and Reporting
📖
now learning

AI for Compliance Monitoring and Reporting

15 min

Priya Desai-Lindqvist had spent the better part of a decade at an EPA regional office reading facility self-reports, the discharge monitoring reports and compliance certifications that regulated facilities are required to submit. There were thousands of them, and she and four colleagues read them by hand, watching for the discharge value that exceeded a permit limit or the report that arrived three weeks late. When her division proposed an AI tool to triage the incoming reports and flag anomalies, Priya was the natural skeptic in the room. She wanted the help. She had also seen what happened when an automated flag was treated as a finding, and a facility's lawyer asked, reasonably, "On what basis did your agency conclude my client violated the law?" This lesson is about building compliance AI that makes Priya faster without ever letting a machine make that conclusion for her.

Where AI Genuinely Helps in Compliance Work

Compliance monitoring is, at its core, the search for the few records that matter inside the many that do not. That is exactly the kind of work AI can accelerate, provided you are precise about what each technique does. Priya mapped the help into five categories.

  • Anomaly and outlier detection. Flagging a reported discharge value that is far outside the normal range for that facility or that industry. The AI is not deciding the value is illegal; it is saying "this one is unusual, look here first."
  • Document review and extraction. Reading a submitted report and pulling out the structured fields (pollutant, value, date, permit number) so a human is not retyping them from a PDF.
  • Pattern detection across records. Spotting that a facility's reports cluster suspiciously just under a permit limit, or that several facilities under one operator show the same gap, patterns a single-report review would never surface.
  • Deadline and timeliness tracking. Watching which required reports are late or missing, a deterministic check that nonetheless saves enormous manual effort.
  • Report drafting. Assembling routine summaries and first drafts of internal reports from the underlying data, leaving the analyst to verify and judge.

Notice the common thread. In every case the AI triages and surfaces; the human decides. Priya's office adopted the tool on exactly that understanding, and she wrote it into the design from the first meeting.

Across government the same techniques appear under six broader application headings, and it is worth knowing them because they are how an agency's compliance portfolio is usually described in a governance document. Document compliance reviews documents against regulatory requirements, classifying them, extracting key information and flagging potential issues, as when contracts are reviewed against procurement regulations. Data quality compliance checks that data meets completeness, accuracy and consistency standards where those standards are themselves regulated. Process compliance monitors whether processes are being followed, for example whether required approvals were obtained before funds were disbursed. Timeline compliance flags actions that did not happen inside a required window, such as a case that should have been resolved within 30 days and was not. Reporting compliance extracts required data and compiles it into regulatory reports. And audit support organises the relevant documents and evidence ahead of an audit review.

Match the technique to the task

Priya also learned not to reach for a complex model when a simple rule would do. Deadline tracking is deterministic: a report is either late or it is not, and a plain calendar check is more reliable, more explainable, and cheaper than any machine-learning model. She reserved the harder techniques for the genuinely fuzzy problems. Anomaly detection earns its keep when "unusual" cannot be written as a fixed rule, because what counts as a normal discharge value for a steel mill differs from a wastewater plant, and the model can learn each baseline. Pattern detection earns its keep when the signal spans many records that no single reviewer holds in their head at once. The discipline is to ask, for each task, whether a rule, a statistic, or a learned model is the lightest tool that does the job. Over-engineering a compliance check does not make it more rigorous; it makes it harder to explain when a regulated party asks why they were flagged.

What You Gain, Stated Precisely

The case for compliance AI is real, and it is routinely overstated in ways that matter. Six benefits are commonly claimed. Each one is genuine at a certain scope and false outside it, so the useful move is to state the scope alongside the claim.

  • Consistency. A human reviewer might miss a violation one day and catch an identical one the next; an automated check applies the same logic every time. What consistency does not mean is correctness. A system that has misread a requirement applies that misreading uniformly, to everyone, all day, which is worse than a human's intermittent error and much harder to notice. Consistency is a property of the process, not evidence about the rule.
  • Coverage. Reviewing ten thousand documents by hand takes months; an automated pass takes minutes. Coverage is the benefit least likely to be exaggerated and the one that most changes what an agency can attempt.
  • Speed. Continuous monitoring replaces periodic audit, so issues surface when they occur rather than months later.
  • Audit trail. The system can record what it found and on what basis, which creates a reconstructable record. A record is not the same thing as accountability. A log can show that a flag was raised, that a reviewer was assigned, and that a decision was entered; it cannot show that the reviewer read anything. Accountability is a named human who can be asked to explain the decision and who has to answer. The log is what that person is asked about.
  • Earlier detection. Problems can be caught internally rather than discovered by external auditors, for the categories of problem the system was built to detect. A monitoring system finds what it looks for, so early detection in one category is not assurance across the rest.
  • Resource efficiency. Compliance staff move from routine checking to the complex judgment calls and investigations that actually need them.

Priya's rule of thumb for briefing her own leadership was to state every benefit with its boundary attached in the same sentence. It made the briefings slightly less exciting and it made them survivable when the Inspector General asked what the system had not been looking at.

What Can Go Wrong

Badly implemented compliance AI does not just underperform; it creates new categories of problem that did not exist before. Seven failure modes recur, and they are worth naming separately because they are detected differently and remediated differently.

  • Misunderstanding. The system misinterprets a regulation, flagging conduct that complies or passing over conduct that does not.
  • False positives. The system flags things as violations when they are not, generating noise and burning out staff.
  • False negatives. The system misses actual violations, and regulators or auditors later find what it should have caught.
  • Bias. The system applies rules differently to different entities or populations, flagging certain groups more often not because they violate more but because of how the system was built or trained.
  • Opacity. Automated determinations lack transparency and the people affected cannot understand why they were flagged.
  • Over-reliance. Staff trust the system too much and stop exercising independent judgment, which converts a triage aid into an unreviewed decision-maker.
  • Brittleness. The system fails when regulations change or novel situations arise, continuing to enforce a rule that no longer exists in that form.

The last one deserves a moment against the first benefit in the previous section. A system is only consistent with respect to whatever it currently encodes. When the regulation moves and the system does not, its consistency becomes the mechanism by which an outdated requirement is applied uniformly across an entire regulated population. The same property is an asset and a liability depending on whether anyone is watching the underlying rule.

The Danger: Automated Enforcement Decisions

The line Priya would not cross was letting the AI's flag become an enforcement action on its own. An enforcement decision affects a regulated party's rights. It can lead to penalties, orders, and public findings of violation. Under OMB Memorandum M-24-10, AI whose output materially influences decisions affecting a person's or entity's legal rights or access to government benefits is treated as rights-impacting, which triggers specific protections: an impact assessment before deployment, ongoing monitoring, and meaningful human oversight.

Due process is the deeper principle underneath the memo. A regulated party is entitled to know the basis for an adverse action and to contest it. "The algorithm flagged you" is not a basis a hearing officer can defend and not one a court will accept. The flag is an input to Priya's judgment. The violation determination is Priya's, made by a named human who can explain it.

The reframe Priya gave her own team: an AI flag is probable cause to look, never a verdict. Treating a flag as a finding is how an agency converts a helpful triage tool into a due-process problem and an Inspector General finding in the same stroke.

This is also the point at which a common piece of advice needs correcting. Compliance monitoring is often recommended as a good first AI use case for an agency on the grounds that the rules are explicit, the outcomes are measurable, human judgment is still essential, and the risks are manageable with appropriate safeguards. Those four reasons hold. The reason sometimes attached to them, that compliance work does not involve high-risk decisions directly affecting citizens, does not. A system that contributes to determining whether someone is non-compliant is contributing to a determination about their rights, which is precisely the definition that triggers the safeguards. Compliance monitoring is a good starting point because the safeguards are learnable there, not because it is exempt from them.

Explainability and Contestability for Regulated Parties

If a facility is going to face scrutiny because of an AI flag, the facility, and Priya, must be able to understand why. A model that says "anomalous" with no reason is useless to a compliance analyst and indefensible to a regulated party. Priya required that every flag arrive with its reason attached: which value, against which threshold or which historical baseline, and how far outside the norm it fell.

Contestability follows from explainability. The regulated party needs a real path to respond before any adverse action, to say "that high reading was a one-time calibration error, here is the documentation." A compliance process that cannot absorb that response is not enforcing the law; it is automating a guess. Priya built the human review and the response window in as required steps, not optional courtesies.

Two obligations sit alongside this and are stated flatly in the guidance her office worked from. Make it clear when people are being evaluated by an AI system, and explain how the determinations are made. And where an entity or person is flagged for non-compliance, provide an opportunity for human review and appeal. Priya treated both as floors rather than ceilings, and she resisted every proposal to describe the appeal path in a way that implied the flag had already settled anything.

Implementation Considerations

Seven decisions determine whether a compliance AI deployment is defensible, and all seven are made before the tool is doing anything interesting.

  • Clear requirements. Document exactly what compliance means for this process. What should the system flag? What is optional against what is required? What is a violation against what is merely a warning? Ambiguity here becomes a defect in every downstream determination.
  • Training data. Train on examples of compliance and non-compliance that represent the full range of situations the system will meet, and have subject matter experts confirm that the examples reflect actual requirements rather than local habit.
  • Human review. Do not fully automate. Maintain human review of AI findings, especially for important decisions.
  • Transparency. Make it clear when people are being evaluated by an AI system, and explain how decisions are made.
  • Appeal process. If someone is flagged for non-compliance, provide an opportunity for human review and appeal.
  • Continuous monitoring. Monitor whether the system's findings match human judgment. If humans frequently disagree with the system, investigate why rather than adjusting the humans.
  • Updates. As regulations change, update the system. As you learn about its blind spots, fix them. Neither happens automatically and neither happens without an owner.

Audit Trails, GAO Review, and the Privacy Act

Everything the system does has to be reconstructable later. The Government Accountability Office (GAO) and the agency's Inspector General (IG) will eventually ask: what did the system flag, what did it not flag, who reviewed it, and what did they decide. An audit trail that records each flag, the reason, the reviewer, the decision, and the date is the evidence that answers those questions. Priya treated the audit log as a first-class output of the system, not an afterthought.

What the log cannot do is stand in for the review it records. A complete audit trail proves that a process ran and that someone's identifier was attached to a decision. It cannot establish that the decision was sound, and an agency that offers its logs as proof of diligence has offered proof of procedure. The value of the trail is that it makes a specific person answerable for a specific determination on a specific date. That is a precondition for accountability, not a substitute for it.

Where the records touch identifiable individuals rather than corporate facilities, the Privacy Act of 1974 applies. It governs how federal agencies collect, maintain, use, and disclose records about individuals retrieved by a personal identifier, and it requires accounting for those records. If a compliance system holds personal information, it must fit within a published system-of-records notice and honor the Privacy Act's access and correction rights. Priya looped in her office's privacy officer before any personal data entered the tool.

Measuring Whether It Actually Works

Do not assume the system is working because it is running. Seven measures answer that question from different directions, and an agency that tracks only the first will be surprised by the others.

  • Accuracy. Do the system's findings match human expert judgment?
  • Consistency. Does it treat similar cases similarly?
  • Completeness. How many actual violations does it catch, which is the false negative question and the hardest one to answer honestly, because you are measuring what you did not see.
  • Specificity. How many non-violations does it flag, which is the false positive rate.
  • Population fairness. Does it apply the rules fairly across populations, or are certain groups flagged at higher rates for reasons unrelated to their conduct?
  • Staff satisfaction. Does it reduce burden or create new problems for the people using it?
  • Audit outcomes. Do external audits confirm the system's findings, or contradict them? This is the measure an agency has the least control over and the most to learn from.

Handling False Positives Without Burning Out Staff

A compliance AI that cries wolf destroys itself. If the tool flags 500 reports a month and 400 are nothing, reviewers stop trusting it and either rubber-stamp or ignore the flags, both of which defeat the purpose. Priya tracked the false-positive rate as a headline metric and set a contractual expectation that flags her analysts overturned should feed back into tuning. In the first quarter the tool flagged about 220 reports a month; her team confirmed roughly 180 as genuine issues and cleared 40 as false positives. That 18 percent false-positive rate was tolerable and trending down as the harder categories got retrained. A rate twice that would have triggered a redesign.

A second deployment, for comparison

A procurement agency ran the same pattern against a different problem. It wanted assurance that contracts complied with federal requirements such as small business preferences and wage requirements, and built a system that extracted key information from contracts, checked the requirements against the regulations, flagged potential violations, and gave an explanation of why each item was flagged. The system flagged 200 contracts a month for human review. Reviewers found 180 were genuine compliance issues and 20 were false positives, and the agency reported roughly 90 percent accuracy for the main compliance categories, which is the same ratio those two numbers describe. Reviewers resolved the flags in substantially less time than manual review would have taken.

Three things came out of that deployment that generalise. False positives created noise at the start, and training reduced them. The system was stronger on some requirements than others, performing better on contract amounts than on small business documentation, and targeted training on the harder categories improved it. And reviewers wanted to understand the system's reasoning, so adding explanations raised their confidence in it. The reported outcome was improved compliance, reduced staff burden and lower cost, which is what that deployment achieved rather than what deployments of this kind produce by default.

Set the two side by side and the useful lesson is what they do not share. One office found 40 of its 220 monthly flags were nothing; the other found 20 of 200. There is no industry false-positive rate to aim at, and quoting one agency's figure as a target for another is how a tuning decision gets made on the wrong evidence. What transfers is the practice: measure the rate for your own system, decide in advance what rate would trigger redesign, and route every overturned flag back into tuning. What does not transfer is the number. Note too that a 90 percent accuracy figure on main compliance categories is a statement about the main categories. It says nothing about the requirements outside them, and those are usually the ones where the documentation is hardest and the errors are most likely.

Fitting Compliance AI Into Agency Governance

A compliance AI system is subject to the same governance as any other AI in the agency's inventory, and the classification step comes first. Ask whether the system is rights-impacting, and note that if it determines whether someone is non-compliant, it affects their rights. From that classification everything else follows: an impact assessment examining the fairness implications and which populations might be affected differently; continuous monitoring of both accuracy and fairness in the determinations; a working mechanism for flagged people and entities to contest what was found; and documentation of how the system works, communicated to the people it evaluates. Priya's office ran all four, and the artifact that made them coherent is the one in the next section.

The Artifact: An AI Compliance-Monitoring Control Matrix

Priya's deliverable was a control matrix: one row per monitoring task, stating the AI technique used, the trigger that forces a human into the loop, the evidence and audit trail captured, and how false positives are handled. This single table told her management, her privacy officer, and a future GAO reviewer exactly how accountability was preserved at each step.

Monitoring task AI technique Human review trigger Evidence / audit trail False-positive handling
Discharge value exceeds permit limit Threshold check plus anomaly detection against facility baseline Mandatory: every flag reviewed before any contact with facility Flagged value, limit, baseline, reviewer ID, decision, date Overturned flags logged and fed to retraining; rate reported monthly
Late or missing required report Deterministic deadline tracking Analyst confirms before any notice of noncompliance Due date, receipt date, gap, reviewer, action taken Check for waivers and extensions before flag is actioned
Suspicious pattern across a facility's reports Pattern / outlier detection over time series Senior analyst review; pattern alone never an enforcement basis Pattern description, records involved, analyst assessment Treated as lead for investigation, not as a finding
Extract fields from submitted PDF reports Document understanding / extraction Analyst spot-checks extraction against source document Source file, extracted fields, confidence, verifier Low-confidence extractions routed to manual entry
Draft routine compliance summary report Report generation from verified data Analyst reviews and signs every report before release Data sources, draft, edits, approving official Not applicable; human authorship is the control

The matrix made the governance story legible. Every row that could touch a regulated party's rights carries a mandatory human review trigger and a named decision-maker, every row produces an audit trail GAO can follow, and the false-positive column shows the system is monitored for the failure that would otherwise erode staff trust. When Priya's office presented the tool for its M-24-10 impact assessment, the matrix was the backbone of the documentation.

Keeping a Human Accountable for Enforcement

Priya's final principle was the simplest and the most important. For any action that affects a regulated party, a specific human being is accountable, by name, in the record. The AI can find the report, extract the values, and rank the queue. It cannot sign a notice of violation, and in Priya's office it never did. That is not a limitation of the technology. It is the design that keeps the agency on the right side of due process, the Privacy Act, and the next audit.

Report Drafting Without Outsourcing Judgment

The last use case, drafting reports, was the one Priya's staff loved most and the one she watched most carefully. An AI can assemble a quarterly compliance summary from verified data in minutes, which is a genuine gift to overworked analysts. The trap is that a fluent draft reads as authoritative even when it is wrong, a tendency people call automation bias, the human inclination to over-trust a confident machine output. Priya set two rules. First, the model drafts only from data the agency has already verified; it never pulls in unverified numbers or invents context. Second, a named analyst reads, corrects, and signs every report before it leaves the office, and that signature is the accountable act. The draft is a time-saver, not an author. When a report later informs a decision affecting a regulated party, the agency can point to the human who stood behind every sentence.

Anti-Patterns to Avoid

  • Selling the log as the accountability. The system produces a complete, timestamped audit trail, and the agency offers it as evidence that determinations were properly reviewed. A log records that a review was entered. It cannot record that anyone read the underlying report, and a rubber-stamped queue produces a log indistinguishable from a scrupulous one. The trail exists so a named person can be asked to explain a decision, not so the file can answer in their place.
  • Treating a flag as a finding. The flag becomes the basis for contacting a regulated party, and the language of the contact implies a violation has been established. This is the failure the whole design exists to prevent, and it is usually committed under workload pressure rather than by decision.
  • Mistaking consistency for correctness. "The system applies the rule the same way every time" is offered as assurance that the rule is being applied properly. Uniform application of a misread requirement is uniform error, and it lands on the entire regulated population at once. Validate the encoded rule against the actual regulation, and revalidate when the regulation moves.
  • False confidence at deployment. A system is built and assumed accurate without ever being validated against expert human judgment. The problem surfaces months later when an external audit contradicts it. Validate extensively before deployment, compare the system's findings against expert judgment, and fix the discrepancies before rollout rather than after.
  • Garbage in, garbage out. The system is trained on poor-quality examples and learns patterns that do not reflect the actual requirements. Have subject matter experts confirm that the training examples represent the requirements as written, not as locally practised.
  • Change blindness. The system is implemented and then never updated when regulations change, so it enforces requirements that have moved on. Establish a process that ties regulatory updates to a system review, and give someone the job.
  • Over-automation. Every compliance decision is automated and humans are removed from the loop, so there is no safety net when the system is wrong. Keep human review for consequential decisions and confine automation to routine tasks where the tolerance for error is genuinely higher.
  • Inequitable application. The system applies rules differently across entities or populations and nobody is looking. Run regular fairness audits, monitor whether certain groups are flagged at higher rates, and investigate and correct systematic bias rather than explaining it.
  • Quoting somebody else's false-positive rate as a target. One agency's tolerable rate is a fact about one system, one regulated population and one data pipeline. Measure your own, decide your own redesign threshold in advance, and do not let a benchmark from a different program do that thinking for you.
  • Reading a headline accuracy figure as coverage. "Ninety percent accurate on the main compliance categories" is a statement about the main categories. The requirements outside them are usually the ones with the messiest documentation and the highest error rate, and they are invisible in that number.

Practice Prompts

  • Inventory the compliance monitoring that happens in your organization. Which processes are most labor-intensive, and which of those could genuinely benefit from AI rather than from a better rule?
  • Pick one compliance process and design an AI system for it. What would it monitor? What would it flag? How would human review work, and who would be named as the decision-maker?
  • Design how you would validate that a compliance AI system works before deploying it. What would you compare its findings against, and what result would tell you it is not ready?
  • For your proposed system, assess the fairness implications. Could it apply rules differently to different populations, and how would you detect that if it did?
  • Map your proposed system into your organization's governance. Is it rights-impacting? What assessment, monitoring, appeal and documentation obligations follow from that answer?
  • Write the sentence your agency would use to tell a regulated party why they were flagged. If you cannot write it without referring to the model, the explainability requirement is not met yet.

Reflection

  • Where in your compliance work could AI improve things, and what would need to be true before you would let it run against real submissions?
  • If a regulated party challenged a determination your office made last month, could you name the person accountable for it and reconstruct the basis from your records?
  • Which of your current compliance checks would keep producing the same answer if the underlying regulation changed tomorrow, and who would notice?
  • Think about the last time an automated alert in your workplace was overturned by a human. Was that overturn recorded anywhere it could improve the system, or did it disappear?

Glossary

  • Compliance monitoring. Continuous assessment that processes, decisions or documents comply with applicable regulations and standards.
  • False positive. When an AI system incorrectly flags something as non-compliant when it actually complies.
  • False negative. When an AI system fails to flag actual non-compliance.
  • Accuracy. The proportion of AI determinations that match expert human judgment.
  • Rights-impacting AI. Under OMB Memorandum M-24-10, AI whose output materially influences decisions affecting a person's or entity's legal rights or access to government benefits, which triggers impact assessment, ongoing monitoring and meaningful human oversight.
  • Automation bias. The human tendency to over-trust a confident machine output, which is why a fluent draft reads as authoritative even when it is wrong.
  • Audit trail. The record of what was flagged, on what basis, by whom it was reviewed, what was decided and when, kept so a determination can be reconstructed later.
  • System-of-records notice. The published notice under which a federal agency maintains records about individuals retrieved by a personal identifier, and within which a compliance system holding personal information must fit.

Closing

The tool made Priya's office faster in exactly the way it promised: it read what nobody had time to read and put the interesting records at the top of the queue. What it never did was answer the lawyer's question. That answer still comes from a person who looked at the report, weighed the facility's response, and made a determination they can defend by name and by date. Every control in this lesson, the mandatory review trigger, the reason attached to the flag, the response window, the audit trail, the signature on the report, exists to keep that person in the position of having decided. None of them decides anything on their own, and an agency that starts describing them as though they do has quietly moved the determination onto the machine while keeping the paperwork that says otherwise.

Key Takeaways

  • AI triages; the human decides. Anomaly detection, document extraction, pattern spotting, deadline tracking, and drafting all surface candidates. None of them determine a violation.
  • A flag is probable cause, never a verdict. Treating an AI flag as a finding turns a helpful tool into a due-process violation and an IG finding at the same time.
  • Rights-impacting compliance AI triggers M-24-10 protections. Impact assessment before deployment, ongoing monitoring, and meaningful human oversight are required when AI output influences enforcement. Compliance is a good first use case because those safeguards are learnable there, not because it is exempt from them.
  • State every benefit with its boundary. Consistency is not correctness, an audit trail is not accountability, early detection covers only what the system was built to detect, and a headline accuracy figure describes the categories it was measured on.
  • Explainability enables contestability. Every flag must carry its reason, and every regulated party must have a real path to respond before any adverse action, including an opportunity for human review and appeal.
  • Build the audit trail as a first-class output. GAO and the IG will ask what was flagged, by what logic, reviewed by whom, and decided how. Log all five for every flag, and remember the log makes a person answerable rather than answering for them.
  • Respect the Privacy Act when records touch individuals. Personal information needs a published system-of-records notice and must honor access and correction rights; involve your privacy officer early.
  • Manage false positives or the tool dies. Track the false-positive rate as a headline metric, feed overturned flags back into tuning, set your own redesign threshold in advance, and do not borrow another program's rate as a target.
  • Measure across all seven dimensions. Accuracy, consistency, completeness, specificity, population fairness, staff satisfaction and external audit outcomes each catch a failure the others miss.
  • A named human is accountable for every enforcement action. The model never signs a notice of violation. Human authorship of the decision is the control that keeps the agency defensible.

Frequently Asked Questions

Is a compliance monitoring system rights-impacting? Ask whether its output materially influences a determination about a person's or entity's legal rights or access to benefits. A system that contributes to deciding whether someone is non-compliant is contributing to a determination about their rights, so the answer is usually yes, and the M-24-10 protections follow: impact assessment before deployment, ongoing monitoring, and meaningful human oversight. Make the classification explicitly and record it.

We have a complete audit trail. Does that satisfy our oversight obligations? It satisfies the reconstruction requirement and nothing beyond it. A log shows that a flag was raised, that someone was assigned, and that a decision was entered on a date. It cannot show that the assigned reviewer read the underlying report, and a queue that was rubber-stamped produces a log that looks exactly like a queue that was scrutinised. Treat the trail as what makes a named person answerable, and treat that person as the control.

Our system applies the rule identically every time. Is that not more reliable than human reviewers? More consistent, yes. More correct, only if the encoded rule is right. Uniform application of a misread requirement produces uniform error across the entire regulated population, and it is harder to spot than a human's intermittent mistakes because nothing looks anomalous. The same applies after a regulation changes, when consistency becomes the mechanism that keeps applying the old requirement. Validate the encoded rule against the regulation, and tie regulatory updates to a system review.

What false-positive rate should we aim for? There is no transferable number. One office found 40 of its 220 monthly flags were nothing; another found 20 of 200 in a different program with different data. Measure your own rate, decide in advance what rate would trigger a redesign, and route every overturned flag back into tuning. Borrowing a figure from another agency means making a tuning decision on evidence about somebody else's system.

If our system flags a facility, can we send the notice of noncompliance automatically? No. Any action affecting a regulated party needs a named human who reviewed the flag, considered the response, and made the determination. The flag is probable cause to look. Automating the notice would put the agency in the position of having no answer when someone asks on what basis it concluded a violation occurred, and "the algorithm flagged you" is not a basis a hearing officer can defend.

Do we have to tell regulated parties that AI is involved? The guidance this lesson works from says to make it clear when people are being evaluated by an AI system and to explain how determinations are made. Treat that as a floor, and check what your own agency's policy and any applicable notice requirements add on top of it. Be careful about the shape of the disclosure: telling someone an AI was involved does not answer the question of why they specifically were flagged, and it is the second answer that contestability depends on.

Our reviewers agree with the system almost every time. Is that good news? It is ambiguous, which is why the continuous monitoring measure is framed as "do findings match human judgment" rather than "do reviewers approve". Very high agreement can mean a well-tuned system or it can mean reviewers have stopped reviewing, and the two produce the same statistic. Sample the agreements as well as the disagreements, and investigate frequent disagreement by looking at the system rather than at the reviewers.