←
AI for Government
Capable · M25 · lesson 25 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Minimum Risk Management Practices
📖
now learning

Minimum Risk Management Practices

15 min

Celeste Mwangi had spent four years as a compliance analyst in the Department of Housing and Urban Development's Office of the Chief Data Officer before her director handed her a new assignment: run a gap assessment against OMB M-24-10, the Office of Management and Budget memorandum issued in 2024 that set binding minimum practices for federal agency use of artificial intelligence. Celeste pulled up HUD's AI system inventory, 22 active systems, ranging from a document classification tool in the Office of General Counsel to a mortgage risk-scoring model used by the Federal Housing Administration. She built a six-column spreadsheet and walked each system owner through six required practice areas. Three systems met every requirement. Nineteen did not. This lesson explains what those six minimum practices are, what failing them looks like in practice, and how to run the same gap analysis Celeste ran.

What OMB M-24-10 Is, and Why It Is Not Optional

OMB M-24-10, formally titled "Advancing Governance, Innovation, and Risk Management for Agency Use of Artificial Intelligence," is a binding policy memorandum. Unlike voluntary AI principles or aspirational frameworks, OMB memoranda carry compliance obligations. Agencies that fail to meet them face budget consequences and Inspector General scrutiny. That is the difference between this document and, for example, a consensus risk framework an agency may choose to adopt.

One practical caution before you cite it in a memo of your own. OMB guidance is periodically revised, rescinded, and superseded, and the requirements below are stated as this curriculum states them. Confirm which memorandum is currently in force for your agency, and under what number, before you tell a program office that a specific practice is mandatory. The substance of the six practices is durable; the citation is not, and citing a superseded memorandum is the fastest way to lose a room full of skeptical system owners.

The analogy that helped Celeste explain the obligation to those program officers: think of the minimum practices as the fire code for federal AI systems. Nobody debates whether sprinklers are worth the cost on a case-by-case basis. The code sets a floor below which no building may operate. A system that has been running for five years and producing useful outputs gets no grandfather clause. The floor applies to it too.

Why a Floor Was Needed at All

Before this guidance existed, agencies approached AI governance inconsistently. Some had thoughtful, comprehensive programs. Others treated AI like any other IT system, applying frameworks designed for databases and software rather than for systems that make consequential decisions about people. The result was scattered risk management. Some agencies conducted impact assessments and some did not. Some monitored systems continuously and some did nothing until a problem surfaced. Some built human oversight and appeal mechanisms and some deployed fully automated systems with neither.

Minimum practices level that field. They are not arbitrary bureaucratic requirements, and they are not high bars. Each one is distilled from cases where AI systems caused problems that a basic safeguard would have prevented. The point of a floor is not excellence. It is that no agency is flying blind, and that when something goes wrong there is a named person, a document describing what the system was supposed to do, a test result to compare against, and a plan for what happens next.

The Six Minimum Practice Areas

The curriculum enumerates six categories of minimum safeguard: governance and accountability structures, documentation and transparency, testing for bias and performance, human review and override, ongoing monitoring and performance assessment, and incident response and escalation. Elsewhere in the same source material the count is given as seven, but no seventh practice is ever named or described; the enumerated list holds six, and six is the number used here. If your agency's own policy names a different count, follow your policy and reconcile the difference in writing rather than assuming one source is authoritative.

Every covered system must meet all six. There is no partial credit, and a system that is strong in five areas and absent in the sixth is not five-sixths compliant. It is non-compliant, and the gap is where the incident will come from.

Governance and accountability

Every AI system must have governance that establishes decision rights, assigns responsibility, and creates accountability. In practice that means four things. Clear ownership: a designated owner responsible for the system's performance, updates, and compliance, and for a rights-affecting system that owner should be a named individual, not a team and not a vendor, whose name appears in a document and who sits in an accountability chain reaching agency leadership. Decision rights: it must be clear who approves deployment, who approves changes, and who decides to escalate or retire the system. Governance review: systems pass through a review process before deployment and periodically afterward, with the depth of review scaled to the system, but never skipped entirely. Escalation paths: a staff member who notices a problem must have a route to raise it, and if that route leads nowhere there must be a second one.

What failure looks like: Celeste found a tenant complaint routing system whose entire program office leadership had turned over since deployment. The current director did not know the system existed. When it misfired, there was nobody whose job it was to act. Implementation is unglamorous and mostly clerical. Document role assignments for each system, including owner, data steward, and fairness lead. Create governance templates such as an impact assessment template and a monitoring plan template. Establish a standard review cadence, whether quarterly or annual, and write it down. Create escalation procedures and, critically, communicate them to the people who would need to use them.

Documentation and transparency

Systems must be documented adequately, and people affected by them must have transparency about how decisions are made. System documentation covers six questions: what the system does, how it works in terms of decision logic rather than source code, what data it uses including sources and known limitations, how it performs against defined metrics, what could go wrong in the form of known failure modes, and how it has been tested with results attached. A model card is the structured artifact that carries most of this, including training data provenance and where the system is deployed. Cards must be current.

Transparency to affected people is a separate obligation with four parts: notice that AI is being used in a decision affecting them, explanation of how the decision is made in plain language rather than technical terms, appeal so they can contest the decision and request human review, and redress so that harm the system caused can actually be addressed. Agencies must also be able to explain systems to the public under the Freedom of Information Act, which gives citizens the right to request government records.

What failure looks like: a natural language processing tool had a model card written at deployment in 2021. The model had been updated twice. The card still described the original version, which meant the agency's primary evidence of what the system did was wrong. Implementation: require documentation templates to be completed before deployment, publish accessible explanations of how high-impact systems work, establish appeal processes and communicate them to affected populations, create feedback mechanisms so people can report problems, and actually respond to complaints and escalate serious ones.

Testing for bias and performance

Before deployment and periodically afterward, systems must be tested for performance, accuracy, and bias. Pre-deployment testing has five components: accuracy assessment on representative data; fairness assessment of whether outcomes are equitable across demographic groups, including testing for disparate impact; edge case testing on unusual or extreme inputs; adversarial testing of whether the system can be tricked or manipulated; and data quality validation of whether the training data was appropriate and what known quality issues it carries. Ongoing testing continues with performance monitoring in production, drift detection, and fairness monitoring.

Two cautions belong with those tests. First, results must be documented by demographic group, not just as an overall accuracy number. Second, any threshold you apply to a fairness result is a threshold your agency sets for itself and records in advance. A measured disparity is a signal that triggers investigation and legal review, not a legal finding in its own right, and the direction of the finding does not change that. Passing an adversarial test is evidence that the specific attacks you tried did not work; it is not a guarantee that the system cannot be manipulated.

What failure looks like: Celeste's mortgage risk model had an 88 percent overall accuracy rate, and the program office was proud of that number. Nobody had broken it down by race, income bracket, or geography. The Office of Fair Housing and Equal Opportunity had been receiving complaints for eight months. Without demographic data, the agency could neither evaluate disparate outcomes nor defend itself if sued. It is also worth saying plainly what that 88 percent was: an estimate of how the model performed on the cases someone tested it against, not a promise about the next file it scores. Implementation: write testing plans before deployment that specify what will be tested and what the success criteria are, use fairness toolkits and model monitoring platforms, document results and use them in the deployment decision rather than filing them afterward, stand up monitoring dashboards, and create alerts for performance degradation and fairness problems.

Human review and override

For consequential decisions, humans must be involved in the decision-making process and must be able to override AI recommendations. Three requirements sit inside that sentence. Human-in-the-loop: for high-impact decisions, a human reviews and approves recommendations before they take effect. Override capability: humans can reject a recommendation and reach a different decision, and the mechanism must be documented and genuinely reachable by the people who need it, not buried in a workflow nobody uses. Meaningful review: the review is not rubber-stamping, which means reviewers need the tools, the reasoning behind the recommendation, and the training to evaluate it.

Be precise about what this practice delivers. Meaningful human review is a mitigation that reduces the rate at which system errors become final decisions. It is not a cure, it does not make an unreliable model reliable, and a design that lists it as a control and then treats the residual risk as handled has renamed the risk rather than managed it. A sprinkler system with a locked shutoff valve is not compliant, and an override mechanism no caseworker knows how to use is not compliant either. The test is whether the override is real, documented, reachable, and used.

What failure looks like: an eligibility screening tool generated denial recommendations routed to a supervisory review queue. Supervisors approved every recommendation without examining the underlying data, because no training explained what to look for or how to override. The human review step existed on paper only. Implementation: define which decision types require human review and which may be automated, design the review workflow so the case is presented with the information a reviewer needs, ensure reviewers can see why the system recommended what it did, provide training on both the system's strengths and its limitations, and track how often humans override. Read that override figure carefully. A rate at or near zero is a finding to investigate, not a success to report, and a rate near total suggests the system is not adding value.

Ongoing monitoring and performance assessment

After deployment, systems must be continuously monitored to confirm they still perform acceptably. Five things get watched: performance measurement of accuracy, fairness, and other key metrics on a regular cycle; drift detection for degradation over time; outcome tracking to see whether the system is achieving its intended purpose; population-level monitoring of whether outcomes remain equitable across groups; and feedback loops that use reports from affected people to improve the system.

Monitoring finds what you chose to watch. Drift detection surfaces degradation in the metrics you defined, on the population you sample, at the cadence you set. A failure mode nobody thought to instrument will not appear on the dashboard no matter how green the dashboard looks, which is why the documentation practice asks you to write down what could go wrong before you build the monitoring. Passive watching is not monitoring either. Someone must own the dashboard, define the thresholds, and follow a documented protocol when an alert fires.

Implementation: establish monitoring dashboards for each high-risk system, define alert thresholds so that exceeding a stated metric escalates, create mechanisms for people to report problems, hold regular review meetings on a stated cadence to discuss monitoring results, and use the monitoring data to drive updates rather than to decorate a quarterly slide. Alert thresholds are, again, numbers your agency sets and records in advance. Their value comes from having been written down before the data arrived.

Incident response and escalation

When problems are discovered there must be clear processes for responding and escalating. Six elements make up the practice: detection through monitoring, user reports, and audits; escalation along clear paths to leadership; investigation with resources to understand what went wrong; remediation to fix the problem and prevent recurrence; communication to notify affected people when problems are discovered; and learning, using incidents to improve the process itself.

A written plan must name specific individuals, not just roles. "Notify the IT Director" is a role. "Notify Keandra Willis, IT Director, at her direct line, within two hours" is a plan. Incident response and ongoing monitoring are the alarm and the evacuation procedure, and the fire code requires both. Implementation: document the incident response process and communicate it widely, establish incident response teams with defined roles, create an incident tracking system, prepare communication templates for affected populations in advance rather than drafting them during an incident, conduct retrospectives, and track incidents over time to identify patterns and systemic issues.

Two Implementation Paths: Retrofit and New Build

The practices apply in two very different situations, and the sequencing differs.

The retrofit path starts from an audit that finds significant gaps: no governance ownership for several systems, minimal documentation, no fairness testing, no monitoring, no incident response process. Over roughly twelve months the agency assigns owners to all systems, creates and completes documentation, conducts fairness audits, implements monitoring dashboards for high-risk systems, and develops incident response procedures. The result is that systems move from unmanaged to managed, meaning problems can now be detected and addressed proactively rather than discovered by a complainant. That is a genuine change in posture, and it is worth being clear that it is a change in the agency's ability to see problems, not proof that the systems no longer have any.

The new-build path uses the minimum practices as requirements for a rights-impacting system before it exists. Pre-deployment: document what the system does and how it works, test for accuracy and fairness, design the human review workflow, create the monitoring plan, and establish the appeal process. At deployment: implement with human review required, activate the monitoring dashboards, and communicate to affected populations about the system and how to appeal. Post-deployment: monitor performance weekly, review with the governance team monthly, conduct a fairness audit quarterly, and accept and respond to appeals and feedback. Completing that sequence means the system operates with appropriate safeguards and oversight in place. It is evidence of a controlled deployment, not a warranty that the system will behave, which is exactly why the monitoring and appeal steps continue after launch.

Running a Gap Analysis on Your Current Systems

Celeste's spreadsheet had one row per AI system and six columns, one per practice area. Each cell was rated green for fully meets the requirement with documentation, yellow for partial compliance or incomplete documentation, and red for does not meet the requirement. She interviewed each system owner for 45 minutes using the same five questions: Has a named individual accepted formal accountability in writing? Is there a model card reflecting the current version? Do you have demographic breakdown data from the most recent bias test? Can you show me how a caseworker would override a recommendation? Who gets called if this system fails tonight, and in how many hours?

The rating rule mattered more than the ratings. Owners had one week to provide documentation for yellow cells, and anything unresolved moved to red. Without that rule, yellow becomes a permanent resting state where a system is neither compliant nor visibly non-compliant, and the assessment quietly stops being useful. The full exercise covered 22 systems in six weeks with two analysts and one IT liaison, approximately 280 hours of staff time.

Prioritizing Remediation: Rights First

With 19 systems showing gaps, Celeste prioritized by rights exposure rather than by ease. Systems directly affecting individuals, the mortgage risk model, the benefits eligibility screener, the fraud detection tool, went to the top of the queue regardless of how close they were to compliance. A fraud detection model with no demographic breakdown of false positive rates is not just a documentation gap. It is a potential fair lending exposure. A benefits tool with no documented human override is a due process risk.

Yellow-rated systems with partial compliance were often fixable in weeks: outdated model cards, undocumented override procedures, unwritten alert thresholds. Red-rated systems sometimes required 60 to 90 days to build compliance infrastructure before the system could safely continue operating. Uneven implementation is its own hazard here. Fixing the high-risk systems first is correct sequencing, but every system should eventually come into scope, because the unmanaged remainder is where the next surprise lives and it will undermine confidence in the whole program when it surfaces.

Document Before the IG Asks

Inspectors General, the independent watchdogs inside each federal agency with authority to audit operations and refer findings to Congress, are actively building AI audit capacity. Binding minimum practices give them a clear standard to audit against. If an IG opens an inquiry before you have completed your own gap analysis, you lose control of the framing: the inquiry becomes about what the agency failed to do. A completed gap analysis with a documented remediation plan, even mid-progress, changes the conversation to an agency identifying and fixing its own gaps. It improves your position substantially. It does not determine the finding, and it is not a reason to slow the remediation itself.

Celeste documented her gap analysis in a formal internal report, obtained written acknowledgment from each system owner, and attached a remediation tracker with named owners and milestones at 30, 60, and 90 days. When the IG's office sent a preliminary inquiry four months later, the response was ready in three days.

Anti-Patterns

  • Minimum as maximum. Treating the six practices as sufficient rather than as a floor. They are necessary, not sufficient. Use them as the baseline and invest beyond them where the system's consequences justify it.
  • Compliance without understanding. Implementing practices to check boxes. Documentation gets created but is not useful; monitoring happens but drives no decisions. Teach the team why each practice exists and connect each one to an outcome it prevents.
  • Practice without support. Establishing requirements without providing time, tools, or staff. If monitoring is required, fund monitoring. If testing is required, supply testing frameworks and training. Unfunded mandates produce paperwork, not safety.
  • Uneven implementation. Applying the practices to some systems and not others. Start with the highest-risk systems, then expand, but keep every system on the list.
  • Calling a rubber stamp human oversight. Recording a review step in an inventory or impact assessment when reviewers lack the reasoning, the training, or a usable override. That misstates the agency's control posture in a document auditors will read.
  • Treating meaningful review as the answer. Listing human review as the control and closing the risk. It lowers the rate at which errors become final; it does not fix the model.
  • Reading an overall accuracy figure as a promise. An accuracy rate estimates past performance on the cases that were tested. It says nothing binding about the next case, and it conceals disparities until it is broken out by group.
  • Letting yellow become permanent. A gap analysis with no deadline for substantiating partial ratings degrades into a document that records ambiguity rather than compliance.
  • Assuming the dashboard sees everything. Monitoring surfaces the metrics you defined on the population you sample. Failure modes nobody instrumented stay invisible regardless of how healthy the dashboard looks.

Practice Prompts

  • Audit one of your agency's AI systems against the six minimum practices. Document where it meets the requirement and where gaps exist, using red, yellow, and green ratings with a stated deadline for substantiating yellow.
  • Pick one gap from that audit and design an implementation plan. What specifically would be required to meet the practice, who would do the work, and how long would it take?
  • Map the six practices onto your agency's existing governance structures. Which governance body enforces which practice, and what decision does it actually make?
  • Design an escalation and incident response process. What is the first step when someone reports a problem, where does it go if it is not resolved locally, and who is named rather than roled at each step?
  • Build a staffing plan for implementing the practices. What roles are needed, how many people, and what skills? Compare it against what your program currently has funded.
  • Write the five gap-analysis interview questions you would ask a system owner in your agency, then ask them of a system you believe is compliant. Record what you could not substantiate on the spot.

Reflection

Take three minutes on your own program. Which of the six practices is your strongest, and why is it strong? Usually the answer is that somebody owns it personally, which tells you something about how the weaker ones might be fixed. Which practice is weakest, and what specifically prevents implementation: money, tools, authority, or the absence of anyone whose job it is?

Then the harder question. If you implemented all six fully, what would change about how AI is governed in your agency, and would leadership actually support that level of governance once it produced its first inconvenient finding? A minimum practice regime is tested not when it passes a system but when it fails one that a senior official wants deployed. Deciding in advance what happens in that moment is more useful than any template.

Glossary

  • Minimum practices. The fundamental risk management practices that OMB guidance requires agencies to implement for AI systems, functioning as a floor rather than a target.
  • Model card. A structured document describing what a model does, how it was built, what data trained it, how it performs, how it fails, and where it is deployed.
  • Training data provenance. The documented record of where a model's training data came from and what its known characteristics and limitations are.
  • Human-in-the-loop. A governance model in which humans review and can override AI recommendations before those recommendations take effect.
  • Bias testing. Systematic evaluation of whether a system produces equitable outcomes across demographic groups, with results reported by group rather than in aggregate.
  • Disparate impact. A statistical difference in outcomes for protected groups that may indicate discrimination even without intent. A measured disparity is a signal for investigation and legal review, not a legal conclusion.
  • Model drift. Degradation in performance as the real-world data a system processes diverges from the data it was trained on.
  • Performance monitoring. Continuous measurement of a system's metrics in production, with alerting when a defined threshold is crossed.
  • Adversarial testing. Deliberate attempts to trick or manipulate a system before deployment, in order to find failure modes normal testing misses.
  • Incident response. The documented process for detecting, escalating, investigating, remediating, and communicating about problems with an AI system, plus the retrospective that improves the process.
  • Gap analysis. A structured assessment of each system against each required practice, rated and dated, with a deadline for substantiating partial ratings.

Closing

Minimum practices are where governance theory becomes something an agency does week in and week out. Everything upstream, governance structures, system inventories, risk classification, exists to make these six practices possible, and everything downstream, privacy impact assessments, data governance, validation protocols, is a specific instantiation of one of them. That is why a gap analysis is such a useful diagnostic: it tests the whole governance apparatus by asking whether it produced six concrete artifacts for each system.

Celeste's spreadsheet found three compliant systems out of 22. The number that mattered to her director was not the three. It was that HUD now knew which nineteen were exposed, in what way, and who owned the fix. An agency that cannot answer those three questions for every system on its inventory is not managing AI risk regardless of how many frameworks it has adopted, and an agency that can is in a position to defend its program to an Inspector General, a court, or a constituent, whichever arrives first.

Key Takeaways

  • Minimum practices are required, not advisory. They set a floor every covered AI system must meet regardless of when it was built or how well it has performed. No grandfather clauses apply.
  • Six practice areas, all mandatory. Governance and accountability, documentation and transparency, testing for bias and performance, human review and override, ongoing monitoring, and incident response and escalation. Partial compliance is non-compliance.
  • Confirm the citation before you rely on it. The substance of the practices is durable, but OMB guidance is revised and superseded. Check which memorandum currently binds your agency.
  • Named individual accountability is the foundation. A team, a vendor, or an unnamed role does not satisfy governance for a rights-affecting system. A specific human with documented ownership does, alongside written decision rights and a second escalation path.
  • Demographic breakdown in bias testing is not optional. An overall accuracy figure conceals disparate impacts, and it only ever described past performance on tested cases. Any threshold you apply is one your agency sets and records in advance, and a measured disparity triggers review rather than settling a legal question.
  • Human override must be real, not nominal. Document the procedure, train the staff, give reviewers the system's reasoning, and verify that overrides actually occur. Meaningful review is a mitigation that lowers the rate of final errors, not a fix for the model.
  • Monitoring sees only what you instrumented. Define the metrics, the population, the cadence, the thresholds, and the protocol that fires when an alert does. Write down what could go wrong before you build the dashboard.
  • Incident plans name people, not roles. Include detection, escalation, investigation, remediation, communication to affected people, and a retrospective that feeds back into the process.
  • The gap analysis is the practical tool. Rows are systems, columns are practices, cells are red, yellow, or green, with a 45-minute structured interview per owner and a one-week deadline before yellow becomes red.
  • Prioritize remediation by rights exposure, then finish the list. Benefits, housing, credit, and law enforcement systems first, but uneven implementation leaves the next incident sitting in the systems you skipped.
  • The floor is not the goal. Treating the minimum as the maximum is the most common failure. Build past it where consequences warrant, and fund the practices you mandate.

Frequently Asked Questions

How many minimum practices are there, six or seven? The enumerated list in this curriculum names six: governance and accountability, documentation and transparency, testing for bias and performance, human review and override, ongoing monitoring, and incident response and escalation. A count of seven appears in the same material without a seventh practice ever being named, so six is used here. Follow your agency's own policy where it differs, and reconcile the discrepancy in writing.

Does an old system that has worked fine for years have to comply? Yes. The floor applies to systems already in operation, not only to new deployments. Good historical performance is not a compliance record, and it is not evidence that the six artifacts exist.

Can a vendor own the system for governance purposes? No. A vendor can operate, host, and maintain a system, but accountability has to rest with a named individual inside the agency who sits in a chain reaching leadership. If the only person who can explain the system works for the contractor, the governance practice is not met.

What counts as meaningful human review? Review by someone who can see the reasoning behind the recommendation, has been trained on the system's strengths and failure modes, has an override that is documented and easy to reach, and actually uses it sometimes. Anything less is a countersignature on an automated decision.

Is a high overall accuracy figure enough to satisfy the testing practice? No. Results must be documented by demographic group. An overall figure conceals disparities, and in any case it estimates past performance on the cases that were tested rather than predicting the next one.

How long does a gap analysis take? Celeste's covered 22 systems in six weeks with two analysts and one IT liaison, at roughly 280 hours of staff time, using 45-minute structured interviews. Scale from there, and expect the interviews to be the cheap part relative to chasing documentation.

What should we do first if we fail almost everything? Assign named owners. Ownership is the practice that makes the other five fixable, because every remaining gap then has somebody whose job it is to close it. Then prioritize by rights exposure rather than by which fix is easiest.

Do the minimum practices mean our AI program is safe? They mean it is managed. Meeting the floor gives you the ability to detect and address problems rather than discover them through a complaint or an audit. That is a real improvement in posture and it is not the same as a guarantee, which is why the practices include continuous monitoring and appeal rather than ending at deployment.