AI Policy Monitoring and Enforcement
Gerard Whitfield is the Deputy CIO for a federal civilian agency with about 8,400 employees across 14 regional offices, an average of roughly 600 people per office. Last fiscal year, his agency deployed three AI systems: a document triage tool for the legal division, a scheduling optimizer for field inspection teams, and a natural language interface for the agency's internal knowledge base. Each went through the agency's nascent AI governance review process. Each received approval. Eleven months after the last deployment, an Inspector General review found that one of the three, the legal document triage tool, had been operating for six months without the quarterly bias monitoring reports the agency's own AI use policy required. Two reporting cycles had passed. The missing reports were not falsified; they were simply never produced. Nobody had been assigned responsibility for producing them, and nobody noticed when the deadline passed. Under the agency's own severity scale the finding was a Level 2 compliance deficiency, requiring a written corrective action plan within 30 days and a re-review of the deployment authorization. Gerard's first call was to his general counsel, his second to the team lead for the triage tool, and his third was to himself, asking how this had happened.
The Gap Between Policy and Practice
Gerard's situation is the dominant failure mode in government AI governance: the policy existed, the system was approved under it, and then nobody monitored whether the policy's ongoing requirements were being met. Note what did not fail. The review board met, read the submission and reached a judgment. That judgment was a point-in-time assessment of a system as described on the day it was described, and nothing about it continued operating afterward. An approval is a decision, not a control.
Most agencies have learned, through Office of Management and Budget guidance, through the National Institute of Standards and Technology's AI Risk Management Framework, which is voluntary rather than binding, and through Government Accountability Office recommendations, that they need written AI policies. The harder problem is the distance between having a policy and operating a compliance function that checks whether the policy is followed. That distance is where most enforcement failures live, and it is not closed by writing a better policy.
Think of it as the difference between a building code and a building inspector. You can write an excellent building code. If nobody checks whether the construction followed it, the code is aspirational rather than protective. AI policy monitoring is the inspector function: the mechanism that converts policy language into operational accountability. And it is worth being precise about what that mechanism can deliver, because the surrounding literature routinely overstates it. Monitoring does not ensure compliance. It makes non-compliance visible, on the cadence you chose, for the things you chose to measure. Everything you did not think to watch remains exactly as invisible as it was before.
What Monitoring Requires
Effective AI policy monitoring has four operational components, and each must be designed before a system is authorized for production rather than assembled after a finding. They are an inventory of active systems carrying their specific obligations, defined metrics with stated measurement methods, clear escalation and reporting pathways, and a structured process for exceptions and remediation. Missing any one of them produces a recognizable failure: obligations with no owner, metrics nobody can evaluate, alerts that reach nobody with authority, or a pattern of undocumented gaps that looks identical to negligence when an auditor reads it.
The timing requirement is the part that gets waived under schedule pressure, and it should not be. Designing these components after a system is live means retrofitting obligations onto a team that already has its hands full, negotiating metrics with people who now have an interest in easy ones, and reconstructing a baseline from data collected before anyone knew it would be evidence. Everything is cheaper at the authorization boundary, when the deployment still depends on getting it right and nobody has yet formed an attachment to the current arrangement.
The Inventory and Its Obligations
This is the foundation. Every AI system in production should appear in a registry recording the system, its deployment date, its authorization status, and the specific ongoing compliance requirements attached to it, including who is responsible for each requirement and when each falls due. A quarterly bias report is not a requirement that monitors itself. It requires a named person and a calendar entry, and if it has neither, it will be produced exactly as often as somebody happens to remember.
Gerard's agency had approved three systems and had documented no compliance owner for any ongoing requirement on any of them. The policy stated that quarterly bias monitoring reports were required. The approval memo did not name who would produce them. That single gap between a policy that says what must happen and an authorization that says who must do it produced the Inspector General finding, and it is the most common gap in the field.
The fix is a discipline at the authorization boundary rather than a new system. No system reaches production until every ongoing obligation attached to it has a named individual, a stated frequency, a due date and a defined artifact that counts as having met it. Write those into the authorization itself, so the record of what was approved is also the record of what was promised. When the owner transfers, the obligation transfers with the role rather than lapsing with the person, which is the failure that turns a one-off gap into a chronic one.
Metrics That Can Actually Be Checked
Each compliance requirement must specify what will be measured, how, and what range of results is acceptable. A requirement to monitor for bias is not measurable. A requirement to run demographic parity analysis monthly across the six protected class categories named in the agency's own equity policy, using a stated statistical methodology recorded in a named appendix, and to flag any metric deviating from baseline by more than a defined threshold, is measurable and auditable. Those specifics are illustrative of the form rather than a standard to copy; what transfers is that every element of the requirement is checkable by someone who did not write it.
Metrics should cover at least three domains. Accuracy and performance asks whether the system is performing as specified or whether accuracy is degrading over time. Equity and fairness asks whether outcomes are distributed equitably across protected class categories, or whether the system performs significantly differently for different demographic groups. Policy compliance asks whether the procedural requirements in the AI use policy are actually being met: the reports, the reviews, the human oversight protocols. The third domain is the one agencies leave out, and it is the domain in which Gerard's failure occurred.
One caution belongs here because it is the difference between a monitoring program and a monitoring theater. Measurable is not the same as sufficient. A metric that is easy to compute and easy to audit can still be the wrong metric for the harm you are actually worried about, and a system passing a well-specified threshold every month is producing evidence about that threshold and nothing else. Choose the metrics by asking what harm would show up in them, then check periodically whether the harms you now worry about are still the harms the metrics can see.
Escalation and Reporting Pathways
When a monitoring metric falls outside its acceptable range, what happens next? The framework must specify who receives the alert, within what timeframe, and what actions are then required. Low-severity deviations might trigger an internal review by the system owner within a short stated window. High-severity deviations, such as significant accuracy degradation or a bias metric exceeding the acceptable threshold for a protected class, should escalate to the Chief AI Officer or equivalent and may require suspending the system pending investigation. Carry that last provision as stated: it errs toward protecting the people the system decides about, and an agency that has pre-authorized suspension will act faster than one debating whether it may.
An alert with no named recipient is not an alert. The commonest structural defect here is a monitoring report that goes to a distribution list, where every recipient assumes another is acting. Name the individual, name the deputy, and state the action each is expected to take rather than only what they are expected to receive. The same applies to the timeframe: an escalation with no clock attached will be handled at the convenience of whoever received it, which in a busy quarter means it will be handled after the next report has already been generated.
The pathway also flows upward. The Office of Management and Budget requires federal agencies to report their AI inventory and governance activities, and many state legislatures and city councils are beginning to require similar annual reports. The compliance monitoring function is the source of that data. If monitoring data is not being collected systematically through the year, the reporting obligation cannot be met accurately, and an agency in that position generally discovers it in the final week before a submission is due. Reporting is downstream of monitoring, always, and an agency that does not monitor cannot report honestly no matter how much effort it applies at the end.
Exception Management and Remediation
Not every compliance gap is a crisis. The framework needs a structured way to handle exceptions: situations where a requirement was not met for a documented reason, where a metric sits outside range but is trending back toward it, or where a genuine operational constraint made compliance temporarily impractical. Without such a route, teams facing an impossible requirement have only two options, neither good: comply nominally by producing an artifact that means nothing, or fail silently.
Exception management requires three things. Documentation: what the exception was, when it occurred, and why. Review: who evaluated whether the exception was acceptable, and on what basis. And a remediation plan: what is being done to return to compliance, by when, and who is accountable for it. An Inspector General will eventually ask about exceptions, and a documented process for managing them is the difference between a manageable finding and a determination of systemic failure.
Be careful about how much protection documentation buys. A documented exception is substantially more defensible than an undocumented gap, and it is not a permission slip. Writing down that a requirement was skipped does not make skipping it acceptable, and a file full of well-documented exceptions to the same requirement, quarter after quarter, is not evidence of good process. It is evidence that the requirement is wrong, unresourced, or being ignored with paperwork attached. Track exception rates by requirement, and treat a recurring exception as a finding about the policy rather than about the team.
The Enforcement Ladder
Enforcement mechanisms span a wide range, from warnings at the light end to funding penalties at the heavy end. Between those two endpoints, the sources for this lesson supply a sequence of intermediate steps, and setting them out as a ladder makes two things visible: that most non-compliance never needs to travel far up it, and that the heavy steps only work if the light ones were actually used first.
| Step | What it looks like | When it fits |
|---|---|---|
| Warning | The finding is communicated to the system owner with a deadline to respond | A first, isolated lapse where capability and intent are not in question |
| Corrective action plan | A written plan submitted within a stated period, naming actions, dates and owners | A documented deficiency that the unit can close with its own resources |
| Escalation to the Chief AI Officer | The alert leaves the program and reaches an official with cross-agency authority | High-severity metric deviations, or a corrective plan that was not delivered |
| Re-review of the authorization | The approval to operate is reopened and the system is reassessed | The basis on which the system was approved is no longer demonstrably true |
| Suspension pending investigation | The system stops being used while the question is resolved | Evidence of harm, or a bias metric out of range for a protected class |
| Funding penalties | Financial consequences applied to the non-compliant organization | The heaviest instrument, and the one most likely to punish constituents rather than decision-makers |
The bottom rung deserves a warning of its own. Funding penalties are the mechanism most often reached for rhetorically and least often useful, because the money removed from a struggling unit is generally the money it needed to comply, and the people who feel the consequence are the constituents that unit serves rather than the officials who made the decision. It belongs on the ladder because the source material names it, and it belongs at the bottom because almost every problem worth solving is solved higher up.
The reason to build the ladder deliberately rather than improvising each response is that enforcement only works when the agencies subject to it trust the process. A unit that expects a proportionate response reports a problem while it is small and cheap to fix. A unit that expects the heaviest available instrument regardless of cause conceals the problem until an external reviewer finds it, which is both later and worse. Fairness here is not a courtesy extended to colleagues; it is the mechanism that keeps information flowing to the people who need it.
Proportionality: Capacity Versus Conduct
Federal AI policy is new, and as policies roll out, agencies and sub-units will implement them with varying degrees of success. Some will comply fully. Some will struggle against real resource constraints. Some will misinterpret requirements. And some will find legitimate conflicts between requirements. Treating all four as the same event is the fastest way to destroy an enforcement function's legitimacy, because three of the four are not misconduct at all.
A proportionate framework distinguishes four situations. A capacity gap means the unit lacks the technical capability to implement the requirement and needs training or support. A resource gap means the unit has the capability but lacks the staff time or budget, and needs prioritization or funding. An interpretation gap means the requirement was unclear and the unit made a reasonable but incorrect reading, which calls for clarification and usually a policy amendment. A compliance failure means the requirement was clear, the unit had the capacity to meet it, and it was not met. Each warrants a different response: support, escalation for resources, clarification, and corrective action respectively.
Fair enforcement also requires consistency. If the same non-compliance is a Level 1 finding in one division and a Level 3 in another, the function loses legitimacy faster than any single missed report could damage it, and units start managing their relationship with the compliance office rather than their systems. Document the criteria, apply them uniformly, and publish the criteria internally so a unit can predict how a given lapse will be treated before it happens.
When Requirements Contradict
The fourth situation deserves separate treatment because agencies handle it worst. Sometimes a unit is not failing to comply; it is complying with something else that pulls the other way. Transparency obligations can sit awkwardly against security or privacy obligations. Retention requirements can sit against minimization requirements. A monitoring regime that has no route for reporting a genuine conflict forces the unit to pick one obligation privately and hope nobody notices, which converts a policy problem into a hidden compliance risk.
Build an explicit conflict route. The unit raises the conflict in writing, naming both obligations and the specific decision it cannot make without resolving them. Someone with authority over both, or with access to counsel on both, issues a written determination. That determination is recorded and becomes the answer for every other unit facing the same collision. The value compounds: an agency that has resolved a run of conflicts in writing has effectively built an interpretive layer over its own policy, and the next one arrives already answered. The alternative is every affected unit making its own unrecorded judgment, none of which the compliance function can see.
Continuous Improvement: Closing the Feedback Loop
Monitoring data is the most valuable input to policy improvement, and most agencies use it only to find fault. Patterns in compliance gaps, meaning requirements that are consistently unmet, metrics that frequently fall outside range, and exception categories that recur, are signals about where the policy itself is unrealistic, unclear or unresourced. A requirement that half the agency fails is more likely to be a bad requirement than evidence of half an agency acting in bad faith.
After the Inspector General finding, Gerard's agency implemented a quarterly AI compliance review doing two things in parallel: checking whether active systems meet their compliance requirements, and checking whether the requirements themselves are calibrated correctly. Two of the three existing monitoring requirements turned out to be under-resourced relative to what the policy specified. The policy was updated with realistic timelines, and a compliance coordinator role was added to each system's governance structure with explicit responsibility for each ongoing requirement. No system has missed a monitoring deadline in the 18 months since.
That last sentence is worth reading carefully rather than as a happy ending. It means no system missed a deadline the agency was tracking, on the requirements the agency had defined, over eighteen months. It is a genuine improvement and it is a statement about deadlines rather than about whether the systems are behaving well. The two questions are related and they are not the same question, and an agency that mistakes a clean compliance record for a safe portfolio has replaced one blind spot with a better documented one.
Anti-Patterns
- Treating approval as a continuing control. A governance board reviewed the system and approved it, so the system is governed. The approval was a judgment about a description on a particular day, and nothing about it operates afterward. All three of Gerard's systems were approved and one of them still ran six months out of compliance. Attach ongoing obligations to the authorization, with owners and dates, or the approval is a ceremony.
- Believing the monitoring regime ensures compliance. A monitoring function makes selected non-compliance visible on a chosen cadence. It does not prevent non-compliance, and it is completely blind to anything outside its metrics. State plainly what your regime does and does not cover, review that coverage when the risks change, and never let a leadership briefing imply that monitored means safe.
- Requirements with no named owner. The policy says quarterly bias reports are required, the authorization says nothing about who produces them, and the deadline passes unnoticed. This is the single most common cause of findings in this domain. No system goes to production until every ongoing obligation has an individual, a frequency, a due date and a defined artifact, recorded in the authorization itself.
- Unmeasurable requirements. Monitor for bias is not a requirement, it is a sentiment, and it cannot be audited or failed. Specify what is measured, by what method, and what range is acceptable. Then remember the converse: a highly specific metric is checkable, not automatically correct, and periodically ask whether it still points at the harm you are worried about.
- Exceptions as a permission slip. A team documents why it could not comply, and the documentation quietly becomes the substitute for complying. A documented exception is more defensible than a silent gap, and it is not approval. Track exception rates per requirement, and read a recurring exception as evidence that the requirement is unrealistic, unclear or unfunded.
- Uniform response to capacity and conduct. A unit that cannot comply and a unit that chose not to receive the same finding, and the enforcement function loses the cooperation of everyone in the first group. Distinguish capacity gaps, resource gaps, interpretation gaps and actual compliance failures, and respond with support, prioritization, clarification and corrective action respectively.
- Using monitoring data only to penalize. The compliance office reports violations upward and never asks whether the policy is the problem. Recurring gaps are the cheapest available signal that a requirement is miscalibrated, and an agency that only ever punishes will never receive that signal honestly again.
Practice Prompts
- Take one AI system in production and reconstruct its full obligation list from the policy and the authorization. For each obligation, name the individual responsible, the frequency, the due date and the artifact that proves it was met. Note every obligation for which you could not fill all four fields.
- Rewrite one vague monitoring requirement from your agency's AI policy as a measurable one: what is measured, by what method, at what frequency, and what range triggers what response. Then argue against your own metric by asking what harm it would fail to detect.
- Design the escalation pathway for one system: who receives an alert, within what timeframe, and what each recipient must do. Walk it up the enforcement ladder in this lesson and identify at which rung your agency currently has no defined process.
- Draft your agency's exception process: what must be documented, who reviews it, what a remediation plan must contain, and how exception rates per requirement will be tracked and reported.
Reflection
Sit with the AI system in your agency that carries the highest stakes for the people it decides about. Can you name, without looking, every ongoing obligation attached to it and the person responsible for each? If an Inspector General asked today for the last four periods of evidence on any one of those obligations, how long would it take to produce, and would the answer be retrieval or reconstruction? Then ask the harder version: of everything that could go wrong with that system, how much of it would your current metrics actually see, and what would you have to be watching to catch the rest?
Glossary
- Compliance monitoring. The continuing function that checks whether a system's ongoing policy obligations are being met, as distinct from the approval that authorized the system.
- AI system registry. The record of production systems with deployment dates, authorization status and the specific obligations attached to each, including owners and due dates.
- Ongoing compliance requirement. An obligation that continues after deployment, such as a periodic bias report or a human oversight protocol, and which lapses unless it has an owner and a date.
- Escalation pathway. The defined route by which an out-of-range result reaches a named person with authority to act, within a stated timeframe.
- Exception. A documented instance where a requirement was not met for a stated reason, subject to review and a remediation plan; a record, not an approval.
- Capacity gap. Non-compliance arising because a unit lacks the technical capability to meet a requirement, calling for support rather than sanction.
- Interpretation gap. Non-compliance arising because a requirement was unclear and a unit made a reasonable but incorrect reading, calling for clarification and usually a policy amendment.
- Proportionate enforcement. Matching the response to the cause of non-compliance rather than to its surface appearance, so that inability and unwillingness are treated differently.
Related Lessons
- AI Policy Impact Assessment is the upstream half: modeling whether a requirement can be complied with before it is adopted.
- Drafting Agency AI Policies is where the obligations this function monitors are written, and where enforceability should be designed in.
- Continuous Monitoring Fundamentals covers the technical monitoring of system behavior that feeds the accuracy and equity domains.
- AI Use Case Inventory Management is the registry practice the whole function depends on.
- Congressional and IG Reporting on AI covers the upward reporting that monitoring data ultimately supplies.
- Oversight Mechanisms: IG, GAO, Congress explains who is asking and on what authority.
- Your Agency's AI Governance Structure defines the roles the escalation pathway routes to.
- Federal/State/Local AI Alignment addresses the contradictory obligations that the conflict route exists to resolve.
Closing
Policies are only effective if they are enforced, and enforcement in government is a design problem rather than a disciplinary one. Nothing in Gerard's agency was corrupt or careless in any way an outsider would recognize. A policy said reports were required. An approval memo did not say who would write them. Six months passed. Every element of that sequence is ordinary, which is exactly why the failure is so common and why writing a stricter policy would not have prevented it.
The work is unglamorous: attach every obligation to a person and a date, specify metrics somebody else could check, name who receives an alert and what they must do, give teams a documented route for exceptions and for conflicting requirements, and match the response to whether a unit could not comply or would not. Then use what the monitoring finds to fix the policy as well as the units. An enforcement function that agencies trust is more effective than one they fear, because the trusted one gets told about problems while they are still small, and the feared one finds out when an Inspector General does.
Key Takeaways
- Approval is not governance. A well-written policy and a board sign-off provide no protection against the outcomes they were designed to prevent. The monitoring function is what makes a policy operational after the meeting ends.
- Monitoring makes non-compliance visible; it does not ensure compliance. It sees what you chose to measure, on the cadence you chose. State the coverage limits explicitly rather than letting monitored be heard as safe.
- Every ongoing requirement needs a named owner. Periodic reports, annual reviews and oversight protocols do not self-execute. Record the individual, the frequency, the due date and the qualifying artifact in the authorization itself.
- Metrics must be specific, and specific is not the same as sufficient. Define what is measured, how and within what range, then keep asking whether those metrics can still see the harms you now worry about.
- Exceptions need documentation, review and remediation. A documented exception is defensible; it is not permission. A requirement that generates the same exception every quarter is a finding about the requirement.
- Distinguish capacity from conduct. Capacity gaps, resource gaps, interpretation gaps and genuine compliance failures call for support, prioritization, clarification and corrective action respectively. Treating them alike costs the function its cooperation.
- Enforcement runs a ladder, and the top rungs are where the work is. Warnings, corrective action plans, escalation, re-review and suspension resolve almost everything. Funding penalties usually punish the constituents rather than the decision-makers.
- Use compliance data to fix the policy. Recurring gaps are the cheapest signal available that a requirement is unrealistic, unclear or unfunded, and an agency that only penalizes will stop receiving that signal.
Frequently Asked Questions
We have a governance board that approves every AI system. Is that not enough?
It is necessary and it is a different control from the one this lesson describes. A board approval is a judgment about a system as it was described on the day it was reviewed, and it stops operating the moment the meeting ends. Gerard's agency approved all three of its systems through exactly such a process and one still ran six months out of compliance, because nothing in the approval said who would produce the reports the policy required. The board decides whether a system may operate; monitoring establishes whether it is operating the way the board was told it would.
A unit says it cannot comply because the requirement conflicts with another obligation. What now?
Treat it as an interpretation question rather than a compliance failure, and give it a formal route. Have the unit state both obligations in writing and the specific decision it cannot make. Get a written determination from someone with authority or counsel on both sides. Then record the determination as the standing answer, so the next unit facing the same collision inherits it. The failure mode is leaving units to resolve conflicts privately, which converts a policy defect into a set of invisible individual judgments the compliance function will never see.
Should we impose funding penalties for persistent non-compliance?
Rarely, and almost never as an early step. Funding penalties sit at the heavy end of the enforcement ladder, and the money removed from a struggling unit is usually the money it needed to comply, so the consequence lands on the constituents that unit serves rather than on the officials who made the decision. Work the ladder from the top: a warning, a corrective action plan, escalation to an official with cross-agency authority, re-review of the authorization, and suspension where there is evidence of harm. If a unit is persistently non-compliant, first establish whether the cause is capacity, resources, interpretation or conduct, because only the last one is a case for sanction at all.
Our monitoring reports are clean. Does that mean our AI portfolio is safe?
It means the obligations you defined were met on the schedule you set, which is genuinely worth having and is not the same claim. A clean record says nothing about harms your metrics were never designed to detect, systems that fell out of the inventory, or requirements that were quietly weakened to make them achievable. Read a clean quarter as a reason to examine coverage rather than as a reason to relax: what is running that is not on the list, and what could go wrong that nothing currently measures?
Skill.re