NIST AI RMF: MAP, MEASURE, MANAGE
Bertrand Achille had been a program analyst at the U.S. Department of Veterans Affairs (VA) Benefits Administration for four years before anyone asked him to find out how many AI systems the regional office in Baltimore actually used. He expected the answer to be three or four: the claims-routing tool, the document classifier, and maybe the scheduling assistant. When he finished talking to every team lead and checking every active contract, the number was fourteen. Fourteen AI systems. None appeared in any official inventory. None had a documented risk assessment. Three were built by contractors whose contracts had expired. Two were running on a server nobody could locate in the current IT asset registry. Bertrand's discovery was not unusual. It was exactly the problem that the National Institute of Standards and Technology (NIST) AI Risk Management Framework (AI RMF) was built to make visible.
What the AI RMF Is and Why It Matters
The NIST AI RMF is a voluntary, non-binding framework. It carries no penalties of its own, and adopting it is not itself a legal obligation. What it provides is a shared vocabulary and a systematic method, organized into four core functions: GOVERN, MAP, MEASURE, and MANAGE. A separate lesson covers GOVERN, which establishes who decides, what the rules are, and who is accountable. This lesson covers the three operational functions that sit on that foundation. GOVERN sets the structure. MAP, MEASURE, and MANAGE are where the work actually happens.
Each function answers a plain question. MAP answers "what AI systems do we have, and what are their risk profiles?" MEASURE answers "how is each system actually performing, and is it meeting our standards?" MANAGE answers "what do we do when we find a gap or a problem?" The four functions are not a sequence you complete once. They cycle. Better governance enables better mapping, better mapping enables more sophisticated measurement, and what measurement teaches you feeds back into the policies GOVERN maintains.
Think of the three operational functions as a building inspection cycle. MAP is the walk-through, where you identify every structure on the property before anything else. MEASURE is the engineering assessment, where you apply systematic tests to understand what each structure can safely carry. MANAGE is the remediation, where you fix what needs fixing, restrict access to what is dangerous, and demolish what cannot be made safe. You cannot usefully measure a building you have not mapped. You cannot manage risks you have not measured.
What this cycle prevents, and what it does not
Before this kind of practice became normal, agencies routinely deployed AI with very little understanding of what they had built. A department implements a system to screen job applications. Years later, when someone finally investigates, they discover the system has been systematically disadvantaging certain populations, and by then thousands of decisions have already been made. Or an agency inherits a system from a previous administration, does not know how it works or what it was trained on, and keeps running it because removing it looks harder than maintaining it.
It is tempting to say the three functions prevent these outcomes. They do not. A disciplined MAP, MEASURE, MANAGE cycle does not stop a badly designed system from being built, and monitoring only ever detects what someone chose to watch. What the cycle does is far more modest and far more valuable: it shortens the distance between a problem existing and someone in the agency knowing about it, and it gives that person a defined path to act. The years-later discovery becomes a quarterly finding. That is the real benefit, and it is enough.
Where the legal obligation actually comes from
For federal agencies, the underlying duties are not optional, though they come from policy rather than from the AI RMF itself. Office of Management and Budget (OMB) Memorandum M-24-10, "Advancing Governance, Innovation, and Risk Management for Agency Use of Artificial Intelligence," issued in 2024, requires covered agencies to maintain an inventory of AI use cases, to designate a Chief AI Officer (CAIO) accountable for AI risk governance, and to apply specified minimum practices to higher-risk systems. Agencies that already run the AI RMF cycle are substantially better positioned to show M-24-10 compliance than agencies that do not, because the evidence is a byproduct of the work.
MAP: What AI Systems Do We Have?
MAP is the function Bertrand needed first. It answers a deceptively simple question, deceptively simple because answering it honestly in a large organization is extremely difficult. Some systems were developed in-house. Some were procured from vendors. Some were built internally and licensed to external partners. Some are legacy systems that have been running so long nobody remembers they exist. Some are cloud services that individual departments adopted on their own. Without MAP, risk management is fiction, because you are assessing a population you cannot see.
What counts as an AI system
For governance purposes the definition is deliberately broad. An AI system is any machine-based system that uses machine learning, deep learning, statistical models, or other computational methods to make predictions, recommendations, or decisions, or to support a human who does. That is broader than most people expect. It covers agency-built systems, vendor-operated systems, AI features embedded inside commercial software, and tools that staff adopted independently without formal procurement, which practitioners call shadow AI. It includes systems that decide entirely on their own and systems that only advise.
- Predictive models that score or rank decisions: risk assessment, resume screening, benefit eligibility.
- Classification systems that categorize inputs: spam filtering, document classification, incident categorization.
- Anomaly detection systems that flag unusual patterns: fraud detection, network intrusion detection, benefits fraud.
- Recommendation systems that suggest actions or resources.
- Natural language processing systems that extract information or generate text.
- Computer vision systems that analyze images.
The breadth matters because it is what makes an inventory honest. High-profile deep learning systems get inventoried without much prompting. Traditional statistical scoring models, embedded vendor features, and a free tool a supervisor started using last spring are the ones that go missing. Of Bertrand's fourteen systems, three were fully procured and governed, five were vendor features inside larger platforms, four were contractor-built tools that outlived their contracts, and two were shadow AI adopted by individual teams without IT or legal review. Only the first three would have appeared in a procurement-driven inventory.
Running the mapping process
Most organizations map the same way, in five steps. Send a questionnaire to every department and unit asking about systems they use to automate decisions or analyze data. Interview technical leads to surface systems that were never formally recorded. Review cloud service logs and software-as-a-service subscriptions to find independently adopted AI services. Consolidate everything into a central inventory. Then validate each entry with its system owner, because the questionnaire answers will be incomplete and the interviews will be inconsistent, and the owner is the person who knows.
For each system in the inventory, document a standard set of facts. Name and purpose. Owner and responsible team. Data sources and data characteristics. The decisions the system influences and the population affected. Risk classification. Current accuracy and performance metrics. Fairness testing status and findings. Governance approval status. Monitoring and measurement status. That last field is the one agencies skip, and it is the one that tells a reviewer whether the rest of the record is being kept current or was filled in once and abandoned.
Risk classification inside MAP
Once listed, each system needs a risk level, because the depth of assessment that follows should be proportionate. M-24-10 supplies a practical two-tier structure for federal agencies: rights-impacting AI affects individuals' access to benefits, employment, housing, or credit, and safety-impacting AI could affect human life, including clinical decision support and infrastructure controls. Of Bertrand's fourteen systems, six met the rights-impacting threshold. One, a triage tool influencing the urgency classification of incoming disability claims, met both. Those seven required priority attention.
MEASURE: Systematic Assessment, Not One-Time Testing
MEASURE evaluates AI systems against defined performance, fairness, and robustness standards using repeatable, documented methods. The key word is systematic. A test run once before deployment is not measurement in the AI RMF sense. Measurement is ongoing: a scheduled cadence, defined metrics, documented results that someone reads. A single pre-deployment test tells you how the system behaved on one day against one sample, which is exactly the information that stops being true the moment the system meets real traffic.
The dimensions worth measuring
Six dimensions cover most of what matters. Accuracy and performance: is the system producing the right outputs, measured through accuracy, precision, recall, F1 score, or a domain-specific metric, and measured both across the full population and within demographic subgroups. Fairness: is the system producing equitable outcomes. Robustness and security: how does it handle unusual inputs, typos, edge cases, and adversarial attempts. Data quality: completeness, accuracy, currency, and whether the data still represents the population being served.
The last two are the ones teams add late. Drift detection asks whether performance has changed over time, and separates data drift, where the input distribution shifts, from label drift, where the outcome distribution shifts, from concept drift, where the relationship between inputs and outputs changes. Explainability asks whether the system can account for its decisions, which for consequential government decisions is not a nicety. Feature importance, SHAP values, LIME explanations, and manual audits of individual cases are all legitimate approaches, and the right one depends on who has to understand the answer.
Performance and drift in practice
Bertrand's claims-routing tool was designed to achieve 91 percent accuracy against manual review. When his team ran the first systematic audit since deployment, accuracy had drifted to 84 percent. Performance drift is expected, not exceptional, because the real world does not hold still. The PACT Act (Sergeant First Class Heath Robinson Honoring our Promise to Address Comprehensive Toxics Act, 2022) expanded eligibility for toxic-exposure claims. The system had been trained on pre-2022 data and was never retested against the new claim population. Nothing about the model had changed. The world it described had.
Measuring fairness means breaking the data apart
Aggregate accuracy hides demographic gaps by construction. Fairness measurement therefore requires disaggregated analysis, which means computing error rates separately by group rather than reporting one overall number. There are several established ways to frame the comparison: disparate impact asks whether approval rates are similar across groups, calibration asks whether a given prediction score means the same thing for every group, equalized odds asks whether true positive and false positive rates match, and individual fairness asks whether similar individuals are treated similarly. They do not always agree, and choosing among them is a policy judgment, not a technical one.
Bertrand's team ran a disaggregated analysis on the claims-routing tool. The routing error rate for Veterans identifying as Black or African American was 11.3 percent, compared with 7.2 percent for Veterans identifying as White. The gap had gone undetected because nobody had been assigned to look for it. Finding it took a three-person effort over two weeks using data the agency already held. That is the uncomfortable lesson: the disparity was not hidden behind a technical barrier. It was hidden behind an unassigned task.
How often, and with what
Measurement cadence should follow risk and the stability of the environment. Rapidly changing programs warrant more frequent measurement than settled ones, and a major policy change is its own trigger regardless of the calendar.
| System risk tier | Measurement cadence |
|---|---|
| High risk | Continuous automated monitoring, plus weekly or monthly manual review |
| Medium risk | Monthly or quarterly measurement and review |
| Low risk | Quarterly or annual measurement |
Tooling follows the same proportionality. Open fairness toolkits such as AI Fairness 360 and Fairness Indicators handle the standard disaggregated metrics. Commercial model monitoring platforms, including offerings such as Fiddler, Evidently AI, and WhyLabs, automate drift and performance alerting at scale. General statistical tooling in Python, R, or SQL covers a great deal, and custom monitoring is sometimes the only option for a system with an unusual output. Manual approaches work fine for a handful of systems. They stop working somewhere around the point Bertrand reached.
MANAGE: Respond to What You Found
Measurement identifies problems. MANAGE is what you do about them. At the level of governance decision, the AI RMF frames four choices: avoid the risk by retiring the system, transfer it by shifting accountability to a vendor under contractual protections, mitigate it by changing the system or adding oversight, or accept it by documenting a conscious decision to continue with known limitations and compensating controls. Every response you actually implement is one of those four wearing working clothes.
Operationally, the moves available are more concrete. Mitigate by adding controls while the system stays live, such as manual review for the affected population. Adjust the system itself by retraining, changing thresholds, or redesigning the surrounding process. Escalate to human decision-makers when no technical fix is sufficient. Restrict the system to lower-risk applications only. Or retire it when the risks exceed the benefits. Which one fits depends on the severity of the problem, the feasibility of a fix, and how much value the system genuinely delivers.
Risk acceptance deserves its own note, because it is legitimate and it is routinely confused with negligence. An agency that documents "we accept this risk because mitigation would cost $2.3 million and the error rate affects fewer than 0.3 percent of cases" has exercised governance judgment, and a reviewer can examine that judgment. An agency that simply never looked has not accepted a risk. It has accumulated a liability. The difference between the two is a written record naming who decided, on what evidence, and when the decision gets revisited.
Retrain, restrict, or retire
For the claims-routing tool, MANAGE produced three decisions over four months. First, the team restricted the tool's autonomous routing for toxic-exposure claims, about 18 percent of volume, pending retraining. Second, the vendor retrained the model on 2022 and 2023 claims data, taking eleven weeks at a cost of $87,000 under an existing task order. Third, two systems with expired contracts and no active maintainer were retired, which required one alternative migration and updates to three internal standard operating procedures. Retirement is a response, not a failure.
Meaningful human oversight
M-24-10 requires meaningful human review for rights-impacting and safety-impacting AI, and the word "meaningful" carries weight. The reviewer must have enough information to evaluate the recommendation, the authority to override it, and no performance incentives that effectively punish overrides. Bertrand found reviewers processing 340 AI recommendations per hour, roughly 10 seconds each. That is nominal oversight, not meaningful review. MANAGE required redesigning the workflow: fewer cases per reviewer, a summary explanation alongside each recommendation, and performance metrics built around review quality rather than throughput.
When something goes wrong: incident response
Some findings are not routine remediation but incidents, and those need a defined path. Detection comes through monitoring or user escalation. Documentation records what happened, when, and the scope of impact. Containment takes the system offline, restricts its use, or imposes heightened review. Investigation establishes root cause rather than stopping at the symptom. Remediation fixes the immediate problem and addresses the cause. A retrospective feeds what was learned back into governance and measurement. Communication reaches affected parties and stakeholders, and in government that step is rarely optional.
The AI risk register
Every MANAGE decision has to survive personnel turnover, procurement transitions, and Inspector General (IG) review. The AI RMF recommends maintaining a living AI risk register: a record of each identified risk, the chosen response, the responsible official, the implementation date, and the next scheduled review. Bertrand created a register entry for each of the fourteen systems within thirty days of completing MANAGE and attached each entry to the corresponding record in the IT asset management system. Total documentation effort: one analyst, approximately 40 hours over three weeks.
How the Three Functions Cycle Together
The framework's power is in repetition, not in any single pass. In a first cycle, MAP identifies the systems you can find, MEASURE turns up a fairness problem in one of them, and MANAGE implements a fix and stands up monitoring. In a second cycle, GOVERN updates the fairness assessment policy based on what the first cycle taught, the improved policy surfaces additional systems that the first questionnaire missed, MEASURE assesses those, and MANAGE addresses what it finds. Nothing about the second cycle requires the first to have been perfect.
Continuous improvement is the accumulation of those passes. Models get retrained on new or refined data. Parameters get tuned as the team develops a real feel for the tradeoff between fairness and accuracy in its particular decision context. The measurement plan itself evolves, because you learn which metrics were telling you something and which were decoration. And the process improves, because managing actual systems teaches you things that designing a process on paper cannot. The principles hold constant across a three-person team and a cabinet department. Only the scale and the tooling change.
Anti-Patterns
Every one of these is a way to complete the cycle on paper while leaving the risk exactly where it was.
- Measurement theater. Sophisticated dashboards that drive no decisions. Measurement is expensive, and if findings never produce action, credibility erodes until nobody funds the next round. Tie each metric to a specific decision rule: if this exceeds X, we do Y. Then point at the times it worked.
- Incomplete measurement. Measuring accuracy because it is easy, while ignoring fairness, robustness, and data quality. The dimensions you never measured are the ones that surprise you. Define the full measurement requirement up front even if you implement it incrementally, so the gaps are visible as gaps.
- Manage without escalation. Problems handled quietly at the operational level, one at a time, so the pattern across systems is never recognized as systemic. Require documentation and escalation, expect leadership to ask for it, and review incident trends rather than individual incidents.
- Measurement without accountability. Findings identified, nobody assigned. Every finding needs a named owner, a specific timeline, and tracked progress, and remediation belongs in performance expectations rather than in goodwill.
- Treating monitoring as a safety net. Continuous monitoring only detects what you configured it to watch. A model that fails in a way nobody anticipated fails silently through a green dashboard. Periodically ask what the monitoring would miss, and test that question directly rather than trusting the absence of alerts.
- Declaring the inventory finished. An inventory is a snapshot with a short shelf life, because staff adopt new tools continuously and vendors add AI features to software you already own. Re-run the mapping process on a schedule, and treat a stale inventory as a finding in its own right.
- Human review as a rubber stamp. Assigning a reviewer satisfies the control on paper. If the reviewer's throughput target makes genuine evaluation impossible, or overriding the system is professionally costly, the oversight is nominal. Measure override rates and review time, not the existence of a reviewer.
Practice Prompts
- Map your own corner. Identify three to five AI systems in your area, including anything embedded in software you already use. For each, document name, purpose, owner, data sources, and a preliminary risk classification. This is the mapping process in miniature, and the hard part will be the systems you had to think about before remembering.
- Design a measurement plan. Pick one of those systems. What metrics would tell you it was working? How often would you measure them? What tools do you already have? Which fairness metric is the right one for the decision this system influences, and why that one rather than the others?
- Walk an incident. Assume a fairness problem is discovered in one of your systems next month. Walk the full path: how would it be detected, who investigates, who decides whether to restrict or pause the system, what gets communicated to affected people, and who signs the remediation off.
- Trace the governance connection. How do MAP, MEASURE, and MANAGE connect to your agency's governance body? What information flows up, what decisions flow down, and where in that loop would a finding currently get stuck?
- Evaluate the tooling. Look at the model monitoring platforms available to you, including whatever your existing cloud provider offers. Which would fit your context? What would implementation actually require in staff time and access approvals, not just licensing?
Reflection
Take a few minutes and assess your own organization across the three functions honestly. MAP: do you have a comprehensive AI system inventory, and if not, what is the gap between what is listed and what you believe exists? MEASURE: which metrics are actually tracked today, on what schedule, and by whom? MANAGE: what happens when a problem is identified, and can you name the last time that path was used?
Then answer one more question. Of those three, which improvement would have the highest impact in your context over the next quarter: completing the inventory, adding measurement to systems that have none, or building an incident response path that people would actually use? Most teams discover the answer is the first one, because everything downstream is limited by the population you can see.
Glossary
- AI system inventory: a comprehensive catalog of AI systems, including purpose, owner, data sources, risks, and governance status.
- Shadow AI: AI tools adopted by staff or teams without formal procurement, IT review, or legal review.
- Rights-impacting AI: under M-24-10, AI affecting individuals' access to benefits, employment, housing, or credit.
- Safety-impacting AI: under M-24-10, AI that could affect human life, including clinical decision support and infrastructure controls.
- Disaggregated analysis: computing performance separately by demographic group rather than reporting a single aggregate figure.
- Disparate impact: different outcomes for different demographic groups, even without explicit discrimination.
- Model drift: change in performance over time caused by shifts in data, population, or operating environment.
- Fairness metric: a quantitative measure of equitable outcomes across groups, such as disparate impact, calibration, or equalized odds.
- Incident response: the systematic process for detecting, containing, investigating, and remediating AI system problems.
- Root cause analysis: investigation into the underlying cause of a problem rather than its symptom.
- Automated monitoring: continuous systematic assessment through tooling and alerts, as distinct from manual periodic review.
- AI risk register: a living record of each identified risk, the chosen response, the responsible official, the implementation date, and the next review date.
Related Lessons
- NIST AI RMF: The GOVERN Function establishes the decision rights, roles, and accountability that this cycle depends on. Read it first if your agency has no governance body, because MAP findings with nowhere to land go nowhere.
- Your Agency's AI Governance Structure covers how to design and staff the bodies that receive what MEASURE produces and approve what MANAGE proposes.
- AI Use Case Inventory and Documentation (OMB M-24-10) takes the MAP output and turns it into the specific inventory record federal agencies are required to maintain.
- Risk Classification: Safety-Impacting vs. Rights-Impacting goes deeper into the two-tier distinction used here and why it shapes the assessment that follows.
- Continuous Monitoring Fundamentals develops the operational side of MEASURE, including alert design and what monitoring reliably will not catch.
- AI Incident Documentation and Response expands the incident path sketched above into a full procedure.
Closing
MAP, MEASURE, and MANAGE are where the AI RMF stops being a document and becomes a set of habits. GOVERN establishes the structures and the policies. These three execute against them, and they do it repeatedly rather than once. Agencies that treat the cycle as an annual compliance event get an annual compliance artifact. Agencies that treat it as an operating rhythm get something more useful: a shrinking gap between what is happening in their AI systems and what they know about it.
Bertrand's fourteen systems did not become safe because he found them. They became manageable. Six months earlier, a drifting model was quietly misrouting toxic-exposure claims and a demographic error gap sat undetected in data the agency already owned, and neither fact was anyone's responsibility. After one full cycle, both were written down, owned, scheduled for review, and attached to a named official. That is not a dramatic outcome. It is the whole point.
Key Takeaways
- MAP before anything else. You cannot measure or manage risks you have not inventoried. A rough spreadsheet in week one beats a perfect system in month four.
- Shadow AI belongs in the inventory. Staff-adopted tools, expired-contract systems, and AI features embedded in commercial platforms all count. A procurement-driven inventory finds only the systems that came through procurement.
- The AI RMF is voluntary; the underlying duties are not. For federal agencies, M-24-10 supplies the obligations around inventory, a Chief AI Officer, minimum practices, and meaningful human review. The RMF is the method that makes demonstrating them straightforward.
- Performance drift is expected, not exceptional. Reassess on a schedule, and immediately after any policy change that shifts the population the system serves. Bertrand's model did not change. Its world did.
- Fairness measurement requires demographic disaggregation. Aggregate accuracy conceals group-level gaps by construction, and the choice of fairness metric is a policy judgment rather than a technical default.
- Risk acceptance is legitimate when documented. Naming the decision, the evidence, the official, and the review date is governance. Never having looked is not acceptance; it is accumulated liability.
- Human oversight must be redesigned, not just declared. If reviewers cannot realistically evaluate recommendations at the volume they carry, the oversight is nominal. Fix the workflow, then measure override rates rather than reviewer headcount.
- Findings need owners. Measurement without assigned accountability produces dashboards, escalation trends nobody reviews, and problems that resurface. Every finding gets a name, a date, and a tracked status.
- Build the risk register before the auditors arrive. Forty hours of proactive documentation costs far less than months of reactive evidence-gathering under IG scrutiny, and the register is what survives turnover.
Frequently Asked Questions
Is our agency required to adopt the NIST AI RMF? No. The AI RMF is voluntary and non-binding, and nothing in it carries a penalty. What is binding for federal agencies comes from OMB policy and from law: maintaining an AI use case inventory, designating a Chief AI Officer, and applying minimum practices to rights-impacting and safety-impacting systems. The RMF is simply the most widely used method for meeting those duties in an organized way, which is why the two are so often discussed together.
Where do we start if we have no inventory at all? Start with the questionnaire and the interviews, and accept that the first pass will be incomplete. A partial inventory that exists is more useful than a complete one that is still being designed, because it lets you begin risk classification immediately and it shows leadership something concrete. Plan the second pass from the beginning, and expect it to find systems the first one missed, especially embedded vendor features.
How do we classify a system when it sits on the boundary between rights-impacting and safety-impacting? Classify it as both, as Bertrand did with the disability claims triage tool. The tiers determine the depth of assessment and oversight required, so when a system plausibly meets two thresholds, the more demanding treatment applies. Document the reasoning, because the borderline calls are the ones a reviewer will ask about.
Our team measured accuracy and it looks fine. Is that enough? No. Aggregate accuracy is one dimension of six, and it is the dimension most likely to look healthy while a subgroup gap, a robustness weakness, or a data quality problem sits underneath it. Bertrand's tool showed a single accuracy number for months while its error rate for one group ran materially higher than for another. Disaggregate, and add at least data quality and drift.
What if measurement reveals a problem we cannot afford to fix? That is what documented risk acceptance exists for. Record the finding, the cost and feasibility of each response you considered, the compensating controls you did put in place, the official who accepted the residual risk, and the date the decision will be revisited. An honest accepted risk with a review date is defensible. A finding that quietly disappears from the register is not.
How often should we re-run the whole cycle? Measurement cadence follows the risk tier, so high-risk systems are under continuous monitoring with frequent manual review while low-risk systems may only need an annual look. The mapping pass is different: re-run it on a fixed schedule regardless of findings, because new systems arrive through channels your last pass did not cover. Treat any major policy change affecting the population you serve as an immediate trigger for both.
Skill.re