NIST AI RMF: Practical Implementation Workflows
Priya Nair, a program director at a federal benefits agency, walked out of an executive briefing with a printout of the NIST AI Risk Management Framework and a directive: "Implement this." Three weeks later she had read it twice and produced nothing. The framework described what good AI risk management looks like, govern, map, measure, manage, in careful, abstract language. It did not tell her what her team should do on Monday morning. Her staff kept asking the same question: "Okay, but what's the actual process?" Priya realized the gap. The framework is a destination, not a route. "It tells you what a well-governed AI system looks like," she told her team. "Our job is to build the workflows that get an ordinary project there, every time, without me in the room."
The NIST AI Risk Management Framework, published by the National Institute of Standards and Technology, is a voluntary structure for managing the risks of AI systems. It organizes the work into four functions: Govern (set the rules and roles), Map (understand the system and its context), Measure (test and assess it), and Manage (act on what you find, then monitor). It is deliberately not a checklist. That flexibility is its strength and, for a program director told to implement it, its frustration. The work of an AI strategist is to convert these four functions into repeatable workflows, the step-by-step processes a project team follows.
Voluntary, and Consequential Anyway
Two things about the framework are true at once, and getting them straight prevents a lot of wasted argument. The framework itself is voluntary. Adopting it is not a legal obligation, no statute requires its four functions by name, and nobody can be cited for non-compliance with a document that was written to be adaptable rather than enforceable. Agencies that ignore it entirely are not breaking a law.
What makes it consequential is everything that surrounds it. Federal AI direction and OMB policy reference it, oversight bodies ask agencies how they manage AI risk, and the answers agencies give are shaped by the vocabulary the framework supplies. The obligations are real, and they come from law, appropriations, executive direction and agency policy. The framework supplies the shared language in which those obligations get discussed, satisfied and audited. Treat it as the vocabulary rather than the mandate, and both halves of the sentence stay accurate.
That distinction has a practical consequence for how you sell the work internally. An implementation pitched as "we have to comply with NIST" invites a challenge you will lose, because you do not have to. An implementation pitched as "these are the four questions our overseers will ask, and here is how we will always be able to answer them" survives the challenge, because it is true. Priya used the second framing and stopped having the argument.
Map the Framework Onto What You Already Run
The most expensive mistake in framework implementation is building a parallel universe. Most agencies already have governance processes, procurement processes, security reviews and compliance cycles. The framework is not a replacement for any of them. The work is mapping its functions onto processes that already exist, so that teams are adding AI-specific elements to a familiar path rather than learning a second one.
Take federal procurement, which most agencies run as requirements definition, vendor evaluation, contract negotiation, performance monitoring, and audit and closeout. Requirements definition is where Map belongs: what are the system's requirements and what data will it use? Vendor evaluation is a Govern question: does the vendor have appropriate AI governance of its own? Contract negotiation is also Govern: what performance standards, safety measures and escalation procedures are required in writing? Performance monitoring is Measure, asking whether the agency is getting the performance it contracted for. Audit and closeout is Manage, asking what was learned and how it applies to the next system.
Security processes map the same way. Existing threat modeling extends into Map by asking what AI-specific threats apply, model theft, data poisoning, adversarial attacks. Existing vulnerability scanning extends into Measure by asking how AI systems get tested for weaknesses. Existing incident response extends into Manage by defining what the agency does when an AI system is compromised. In neither case are you creating a new process. You are adding a small number of AI questions to a process that already has an owner, a calendar and a budget line.
From Four Functions to Four Stage Gates
Priya's breakthrough was to stop treating the framework as a document to comply with and start treating each function as a stage gate in her projects' lifecycle. Every AI system her agency builds or buys now passes through four workflows, in order, each with a defined trigger, defined steps, a defined owner, and a defined artifact that must exist before the next stage begins. The framework tells you what a well-governed AI system looks like; a workflow tells your staff what to do on Monday.
She carried one project through as her test case: a tool to help caseworkers prioritize backlogged disability claims. Watch how each function becomes a procedure.
Govern, which runs once and then continuously
Govern is the standing setup, not a project step. Priya's agency did it once and now maintains it. The workflow established a named AI governance owner; a simple inventory where every AI system is registered; a risk-tiering rule asking whether the system affects people's rights or safety; and the policy that no AI system goes live without completing the other three workflows. The artifact is the inventory entry and an assigned risk tier. The claims tool, because it influences how real people's benefits get processed, was tiered high, which set the depth of everything that followed.
Underneath those four decisions sit the ones every agency has to make explicitly rather than by default. What governance model are you running: centralized, hub-and-spoke, or federated? What decision framework applies at each risk level, so that a low-risk pilot is not routed through the same committee as a rights-affecting system? What documentation standard applies to all AI systems? And what is the escalation procedure when something goes wrong, including who has authority to stop a system? Resource allocation belongs here too, because a governance structure with no staff time attached is a diagram rather than a function.
Map, which runs at project start
Map is the workflow that forces understanding before building. Its steps: write a plain-language statement of what the system is for and who it affects; identify the data it uses and where that data came from; list the ways it could go wrong, meaning be wrong, be biased, or be misused; and name the humans who will oversee it. The artifact is a one-to-three-page system map. Operationally, the agency's job is to create the template, define exactly what information is needed about each system, build the process by which teams supply it, and set the criteria reviewers use to assess what comes back.
For the claims tool, the Map workflow surfaced a quiet danger early: the training data over-represented one type of claim, which would have biased prioritization. Caught at Map, this cost days to fix. Caught after launch, it would have meant wrongly deprioritized claimants and a remediation crisis. That asymmetry is the entire argument for the stage gate. Understanding is cheap before the build and expensive after the deployment, and nothing in the framework matters more than that ordering.
Measure, which runs before launch and then on a schedule
Measure is the testing workflow. Its steps: define what "working correctly" means in numbers; test accuracy overall and broken out by group to catch disparate performance; test how the system fails and whether humans can catch its errors; and document the results. Priya's rule is that a high-tier system cannot launch until its Measure artifact, a test report with results by group, is signed by the named owner. Operationally this means defining metrics per system type, writing testing protocols, standing up monitoring, and setting the thresholds at which a system needs attention.
The claims tool's first test showed 94% accuracy overall but only 81% for one claimant category, a gap invisible in the headline number and unacceptable for a rights-affecting system. It went back before launch. Note what the headline number did: it was not wrong, and it was not useful. Aggregate accuracy is an average across populations, and an average is exactly the statistic that conceals a group-level failure. A rights-affecting system cannot launch on the average alone.
Manage, which runs continuously after launch
Manage is the act-and-monitor workflow, the one agencies most often skip. Its steps: decide what to do about each risk Measure found, meaning fix it, add a control, or accept it with a documented reason; deploy with the human-oversight controls Map specified; and monitor in production on a set schedule, because models drift as the world changes. The artifact is a risk register entry per risk plus a recurring monitoring review. The claims tool launched only after the group-accuracy gap was narrowed and a monthly monitoring review was scheduled.
Manage also carries the parts nobody wants to write down in advance: the mitigation playbooks for the risks you expect, the incident response procedure for the ones you do not, the criteria under which a system gets shut down rather than patched, and the feedback loop that turns an incident into a change in how the next system is built. The shutdown criteria are the hardest to agree and the most valuable to have, because the moment you need them is the moment nobody is calm enough to write them.
Four Workflows That Carry the Framework
Stage gates describe a project's path. Alongside them, an agency needs standing workflows that run whenever their trigger fires, regardless of which project is involved. Four of them do most of the work, and each spans more than one function.
System onboarding covers Govern and Map, and starts when a new AI system is proposed. The team submits a system description on a standardized template. The governance body classifies the system's risk level. Documentation requirements follow from that classification rather than being uniform. The governance body reviews what was submitted, then approves, approves with conditions, or asks for more information. Finally, baseline metrics and a monitoring schedule are established before the system goes anywhere. Typical elapsed time in the source material is two to four weeks for moderate-risk systems and four to eight weeks for high-risk ones, which are targets an agency sets for itself rather than external requirements.
Ongoing monitoring is Measure in production. The system reports metrics automatically, covering accuracy, bias, latency and errors. Those metrics are compared against established thresholds. Breaches generate alerts to the governance body. Someone investigates, decides whether the issue needs mitigation, continued monitoring or escalation, and documents the finding and the action. Cadence runs continuous for critical systems and quarterly for moderate ones. The word doing the work in that workflow is "automatically": a monitoring design that depends on somebody remembering to check is not a monitoring design.
Issue escalation is Manage under pressure. A team identifies a problem and submits it. Severity gets classified as low, medium, high or critical. An initial response decides whether to monitor, mitigate or shut down. Then root cause analysis, a mitigation plan, implementation with retesting and validation, and closure with lessons documented. Response targets in the source material are one day for critical issues, one week for high-risk issues and thirty days for moderate ones. Again, these are commitments the agency makes to itself, and they are worth writing down precisely because the temperature is high when they get invoked.
System decommissioning spans Manage and Govern, and is the workflow agencies write last and need eventually. It covers the decommissioning plan and what happens to the data; an impact assessment of what else depends on this system; data handling, whether archived or deleted; stakeholder notification; lessons learned; and knowledge transfer so the next team inherits something. Four to eight weeks depending on complexity. A system that is quietly switched off without this workflow takes its institutional memory with it, and the next team rebuilds the same mistakes.
Matching Depth to Risk Tier
Every workflow in this lesson has a dial on it, and the risk tier assigned during Govern is what turns the dial. Without that mechanism, an agency has only two settings available: apply full rigor to everything, which stalls, or apply light rigor to everything, which leaves rights-affecting systems under-governed. Tiering is what lets a single process serve a low-stakes scheduling tool and a benefits eligibility system without being wrong for both.
The dial moves several things at once. It sets what documentation is required, so a moderate-risk system supplies a short system map while a high-tier system supplies the map, a group-level test report and a risk register with named acceptances. It sets who approves, so a low-risk pilot clears at the program level while a rights-affecting system reaches the governance body itself. It sets monitoring cadence, running continuous for critical systems and quarterly for moderate ones. And it sets the clock: the source material puts onboarding at two to four weeks for moderate-risk systems and four to eight weeks for high-risk ones, the extra time being the cost of the deeper review rather than a queue.
Two failure modes bracket the practice. Under-tiering is the familiar one: a system gets classified as moderate because the high-tier path looks slow, and a system that shapes people's benefits ends up governed like an internal dashboard. Over-tiering is quieter and just as damaging, because when everything is high-risk nothing is, reviewers stop reading carefully, and the governance body becomes a queue. The defence against both is that the tier is a written determination with a named signer and a stated reason, made before the project has momentum and revisited when the system's use changes.
Priya's agency writes the tier into the inventory entry and treats a change of use as a re-tiering trigger. The claims tool was tiered high at registration, which is why its Measure artifact needed a signature and why the group breakdown was mandatory rather than optional. Had the same tool been built to help staff sort an internal reading queue, the identical model would have carried a lower tier and a lighter path, and both answers would have been correct. The tier is not a judgment about the technology. It is a judgment about what happens to a person when the technology is wrong.
Documentation as Memory, Not Bureaucracy
Priya's staff initially saw the artifacts, system map, test report, risk register, as paperwork. She reframed them as institutional memory. When a new caseworker asks why the claims tool works the way it does, the system map answers. When an auditor or the Government Accountability Office asks how the agency knew the tool was fair, the test report answers. When the tool drifts in a year, the risk register tells the next team what was already known.
The documentation standard is simple: each workflow produces one living artifact, written in plain language, owned by a named person, and updated rather than recreated as the system changes. The point is not the document. It is that the agency can always answer the question "how do you know?" A document that exists but was never revisited answers that question badly, which is why "living" and "owned by a named person" are the two words in the standard that actually do work.
Knowing Whether the Implementation Is Working
An implementation can look complete and be hollow. The way to tell is to measure the implementation itself, not just the systems it governs. The source material offers a set of indicators across all four functions, each paired with a target. Every one of these targets is a number the agency chooses and records in advance; none of them is a legal test, and none separates lawful from unlawful practice.
| Function | Indicator | Target from the source material |
|---|---|---|
| Govern | AI systems documented | 100% within 6 months |
| Govern | Systems in inventory | Baseline: how many do you actually have? |
| Govern | Time from proposal to approval | 4 weeks for moderate systems |
| Govern | Escalations triggered per quarter | Zero critical escalations |
| Map | Completeness of system documentation | 100% of required fields |
| Map | Coverage of risk assessment criteria | All major risks identified in mapping |
| Map | Alignment of mapped risks to actual incidents | Are the risks we identified the ones causing problems? |
| Measure | Systems with defined metrics | 100% of operational systems |
| Measure | Monitoring frequency | Real-time for critical, quarterly for moderate |
| Measure | Metrics meeting thresholds | 90%+ of systems meeting performance targets |
| Manage | Time to respond to critical issues | 24 hours |
| Manage | Issues with documented mitigation | 100% |
| Manage | Repeat issues | Zero, indicating the agency is actually learning |
Two of those targets need a warning label. A target of zero critical escalations measures an outcome that staff can improve by not escalating, and a workforce that learns escalations are counted against the programme will quietly stop raising them. Read it as a goal for the systems, never as a performance measure for the people who report problems. The alignment indicator in the Map row is the most informative of the set and the least used: comparing the risks you predicted against the incidents you actually had is the only honest test of whether your mapping is any good.
The Four-Workflow Implementation Checklist
Priya distilled her system into a one-page checklist every project lead now runs. No AI system reaches the public until every box is checked and every artifact exists.
- Govern, registered. Is the system in the AI inventory with an assigned risk tier and a named accountable owner?
- Govern, routed. Does the risk tier determine which review path, documentation set and approval authority apply, rather than everything going to one committee?
- Map, purpose stated. Is there a plain-language statement of what it does, who it affects, and its data sources?
- Map, risks listed. Are the ways it could be wrong, biased, or misused written down, with named human overseers?
- Measure, defined success. Is "working correctly" defined in measurable terms before testing begins?
- Measure, tested by group. Has accuracy been tested overall and broken out by affected group, with results documented and signed?
- Manage, risks resolved. Does every identified risk have a decision, fix, control, or documented acceptance, recorded in the risk register?
- Manage, oversight live. Are the human-in-the-loop controls in place and a monitoring schedule set for after launch?
- Manage, exit planned. Is there a shutdown criterion and a decommissioning path, agreed while everyone is calm?
- Integrated, not parallel. Does each step attach to an existing agency process rather than creating a second one teams must learn?
- Tier-matched depth. Does the rigor of each workflow match the system's risk tier, deeper for rights- or safety-affecting systems?
Three Agency Shapes, Three Implementations
The framework is meant to be adapted to context, and the right implementation looks very different at different scales. Three shapes recur.
The small agency. A regional agency with ten staff and five AI systems can run a complete implementation with almost no apparatus: a single five-person governance board meeting monthly, a simple spreadsheet inventory with a template, the framework mapped onto the IT security processes it already runs, a quarterly review cadence covering all systems, and a one-page lessons-learned note after each review. Within six months it has a full inventory documented, risk levels assigned and monitoring baselines established. That is lightweight and complete, and it is aligned with the framework without heavy bureaucracy.
The large agency. A federal agency with fifty or more AI systems across multiple departments needs a hub-and-spoke model: a central office producing implementation guidance, a local governance committee in each department applying that guidance, central review of high-risk systems and cross-department coordination, monthly metrics reporting up to the centre, and a quarterly all-hands to share lessons. Departments keep autonomy over how they implement, while every system is assessed against the same criteria. The consistency lives in the criteria, not in the org chart.
The cross-agency initiative. A multi-agency effort, for instance around AI in healthcare, works best federated: each agency implements according to its own governance structure, with quarterly coordination meetings, shared documentation templates and tools, an escalation process for issues that cross agency lines, and a joint annual assessment of the initiative overall. This respects agency autonomy while enabling coordination. Agencies do not adopt identical structures; they adopt common standards and common tools, which is the smallest thing that makes their results comparable.
Start With One Project, Not the Whole Agency
Priya's last piece of advice to peers is about sequencing. Trying to implement the framework agency-wide at once produces a binder no one uses. She ran the four workflows on a single project, the claims tool, learned where they were too heavy or too light, tuned them, and only then rolled the refined version out.
The framework rewards this. Its functions are meant to be adapted to context, and the fastest way to learn the right depth for your agency is to run one real project end to end. Complexity is the enemy of adoption: a simple process that teams actually use beats a comprehensive one nobody follows, every time and by a wide margin. The destination is the same for everyone; the route is something you have to walk before you can map it for others.
Anti-Patterns
- The framework as a checkbox. Documentation gets created, compliance gets declared, and everyone moves on. It happens because leadership wants to say the agency has adopted the framework, and producing documents once is easier than running processes forever. Then a system starts exhibiting bias, nobody escalates it through the formal process, and an external audit finds it. The agency now has to explain why a documented process did not catch what it was designed to catch, which is a worse position than having had no process. Implement the workflows, the monitoring and the escalation; let the documentation be the byproduct.
- Implementing everything at once. The framework is comprehensive, so thorough people conclude they must implement it comprehensively. Implementation stalls under its own weight, buy-in falters, and valuable systems cannot get through an approval process nobody can navigate. One agency built a forty-page system documentation template; teams found it so burdensome they stopped submitting systems for review, and systems went live with no oversight at all. Start with the essential elements and add depth as the organization matures.
- Building a parallel process. The framework is new and existing processes are entrenched, so bolting on a separate track feels easier than integrating. The result is duplicated work, conflicts between processes, and teams unsure which one applies. An agency running both a traditional IT procurement process and a separate AI procurement process left a team unable to determine which governed their system, and the procurement was delayed while it got sorted out. Map onto what exists.
- Treating the framework as a mandate. Overstating its legal status wins the first meeting and loses the second, when someone checks. The framework is voluntary, and an implementation justified on a false premise collapses the moment the premise is tested. Ground the obligation where it actually lives, in law, policy and oversight expectations, and use the framework for what it genuinely provides, which is a shared vocabulary.
- Launching on aggregate accuracy. A strong headline number is the most comfortable place to stop testing and the most dangerous. An average across populations is precisely the statistic that hides a group-level failure, and for a rights-affecting system the group breakdown is the test that matters. A system that has only been measured in aggregate has not been measured.
- Counting escalations as failures. A target of zero critical escalations, applied to people rather than systems, teaches staff that raising a problem is a black mark. The metric then improves while the underlying risk grows, and the first anyone hears of a serious issue is from outside the agency. Measure escalations for what they reveal about systems; never let them score the people who file them.
- Skipping Manage. This is the function agencies most often drop, because Govern feels like leadership and Map and Measure feel like progress, while Manage is unglamorous maintenance forever. It is also where the real risk lives, since models drift, populations change, and a system validated once is only known to have been correct once.
Practice Prompts
- Inventory your existing processes first. List the governance, procurement, security and compliance processes your organization already runs, then write next to each the framework function it most naturally carries. The gaps in that column are your actual implementation backlog.
- Identify the integration points. For each existing process, name the specific step where an AI-specific question would attach, and name the person who owns that step today. If you cannot name the owner, you have not found the integration point.
- Design your core workflows. Draft onboarding, monitoring, escalation and decommissioning for your agency, each with a trigger, steps, an owner and a required artifact. Keep the first draft short enough that a project lead reads it in one sitting.
- Define your success metrics. Pick a small number of indicators across the four functions and set targets your agency can defend, then write down explicitly that they are self-set targets rather than legal tests.
- Run the stage gates on one real system end to end. Choose something live and consequential rather than a comfortable pilot, and record every point where the process was too heavy or too light before you scale it.
- Take your highest-tier system and ask what its shutdown criterion is. If nobody can state it, draft it now, while nothing is going wrong.
- Compare the risks you mapped for one system against the incidents it has actually produced. Where they diverge, your mapping process needs work more urgently than your monitoring does.
Reflection
Pick the AI system in your agency with the most exposure to the public. Walk it backwards through the four functions and ask, at each one, not whether a document exists but whether the work happened: is it registered with a tier and an owner, was its purpose and data understood before it was built, was it tested by group and not just in aggregate, and is anyone monitoring it on a schedule today. Then ask the question that separates a real implementation from a documented one: if that system started drifting this month, how would you find out, and how long would it take? If the answer is that someone would eventually notice, you have a Govern function and a Map function and no Manage function at all.
Glossary
- GOVERN: Establishing the structures, policies and processes for managing AI risks, including who decides, what documentation is required, and how issues escalate.
- MAP: Understanding the AI system, the data it uses, and the ways it could go wrong, before it is built rather than after it is deployed.
- MEASURE: Quantifying system performance and risk, including testing broken out by affected group rather than reported only in aggregate.
- MANAGE: Responding to identified risks, deploying the oversight controls, monitoring in production, and learning from incidents.
- Risk tier: The documented classification of a system by whether it affects people's rights or safety, which then sets the depth required of every subsequent workflow.
- Workflow: A defined sequence of steps for accomplishing a governance function, with a trigger, an owner and a required artifact.
- Stage gate: A point in a project's lifecycle that cannot be passed until a named artifact exists and has been signed by the accountable owner.
- System map: The short plain-language artifact produced by the Map workflow, covering purpose, affected people, data sources, failure modes and named overseers.
- Risk register: The record of each identified risk and the decision taken about it, whether fixed, controlled, or accepted with a documented reason.
- Escalation: The process for raising issues from operational teams to the governance body, with severity classification and defined response commitments.
- Hub-and-spoke model: Governance in which a central office sets standards and reviews high-risk systems while local offices implement within their own departments.
- Federated governance: Multiple independent governance bodies coordinating on shared standards and tools without adopting identical structures.
- Model drift: The degradation of a system's performance over time as the world it was trained on changes, which is why monitoring is a schedule and not an event.
Related Lessons
- NIST AI RMF: The GOVERN Function goes deeper on the standing setup this lesson runs once.
- NIST AI RMF: MAP, MEASURE, MANAGE covers the three project-facing functions in their own right.
- Establishing an AI Governance Board details the body that classifies risk and receives escalations here.
- OMB M-24-10 Deep Dive: Full Implementation sets out the federal obligations this framework gives you language for.
- OMB M-24-18 and AI Procurement Governance is where the procurement mapping in this lesson becomes contract clauses.
- Risk Classification: Safety-Impacting vs. Rights-Impacting is the tiering rule the Govern workflow depends on.
- Continuous Monitoring Fundamentals builds out the Measure and Manage schedules in operational detail.
- Testing and Validating AI Systems is the methodology behind the group-level test report.
- AI Incident Response Planning extends the escalation workflow into a full incident capability.
Closing
The framework is powerful because it covers the full lifecycle of an AI system, and power of that kind is inert until somebody operationalizes it. Your job is taking four functions and making them real: workflows teams can follow, metrics that tell you how you are doing, and integration with existing processes that makes adoption cheap rather than heroic. The specific implementation details matter less than the core principles, understanding your systems, maintaining oversight, responding to problems, and governing the whole.
Priya's agency did not become well-governed by finishing a document. It became well-governed when a project lead she had never met ran the four gates on a system she never saw, found a data problem at Map, fixed it in days, and never told her about it. That is what an implementation is for: it moves the judgment out of one person's head and into a process the organization owns. Start now, start simple, learn, iterate. That is how an organization moves from aligned on paper to aligned in practice.
Key Takeaways
- The framework is voluntary; the obligations around it are not. Nothing requires its four functions by name. Federal policy references it, oversight bodies ask the questions it organizes, and the real duties come from law and agency direction. Pitch it as the vocabulary your overseers use, not as a mandate that will not survive scrutiny.
- The framework is a destination, not a route. It describes what good looks like; your job is to build the workflows that get an ordinary project there every time, without you in the room.
- Turn each function into a stage gate. Govern, Map, Measure and Manage become workflows with a trigger, steps, an owner and a required artifact that must exist before the next stage begins.
- Map onto existing processes rather than building parallel ones. Procurement and security reviews already have owners, calendars and budgets. Adding AI questions to them costs a fraction of what a second track costs, and teams actually follow it.
- Map catches the cheap-to-fix problems. Forcing understanding before building surfaces data and bias risks while they cost days, rather than after launch when they cost a remediation crisis.
- Measure must break results out by group. Aggregate accuracy is an average across populations, and an average is the statistic that hides a group-level failure. A rights-affecting system cannot launch on the headline number alone.
- Manage is the step agencies skip. Acting on findings and monitoring after launch is where real risk lives, because models drift as the world changes and a system validated once is only known to have been correct once.
- Four standing workflows carry the framework. Onboarding, ongoing monitoring, issue escalation and decommissioning each span multiple functions and each need a trigger, an owner and an artifact.
- Measure the implementation, not just the systems. Coverage, cycle times, monitoring frequency and repeat-issue counts show whether the process is real, provided you remember these are targets you set for yourself, not legal tests.
- Documentation is memory, not bureaucracy. One living artifact per workflow, owned by a named person and updated rather than recreated, lets the agency always answer "how do you know?"
- Start with one real project. Run the workflows end to end on a single system, tune them, then scale. Complexity is the enemy of adoption, and a simple process that gets used beats a comprehensive one that does not.
Frequently Asked Questions
Is our agency legally required to implement the NIST AI RMF?
The framework itself is voluntary and no statute requires its four functions by name. What is not voluntary is the underlying duty to manage AI risk, which arrives through law, executive direction, OMB policy and the expectations of oversight bodies. Those obligations exist whether or not you use the framework's vocabulary; the framework simply gives you a structured way to answer them and a shared language for discussing them with auditors. Check with your own policy office for what applies to your agency specifically.
Where do we start if we have no governance at all?
Start with the inventory. You cannot tier what you have not listed, and most agencies discover during that exercise that they have more AI systems than anyone believed, several of them acquired inside larger software purchases. Once you have a list, assign each entry a risk tier and a named owner. Those two fields alone let you route the next decision correctly, and they cost days rather than months.
Our system tested at high accuracy overall. Is that enough to launch?
Not for anything that affects people's rights or safety. An overall figure is an average across every population the system touches, and averages conceal exactly the pattern you most need to see, which is a group for whom the system performs materially worse. Break the result out by affected group before launch, document what you find, and treat a gap as a finding to resolve rather than a rounding detail. Passing the group test is evidence the system works for those groups on the data you tested, not a guarantee it will keep doing so.
How do we avoid the framework becoming a paperwork exercise?
Attach every artifact to a decision that cannot be made without it. A system map that nobody reads is paperwork; a system map that a reviewer must sign before the build is funded is a control. The same document does different work depending on whether skipping it stops anything. If an artifact can go missing without any consequence, it will, and no amount of exhortation changes that.
How much of this applies to a small agency with a handful of systems?
All of the functions, almost none of the apparatus. A five-person board meeting monthly, a spreadsheet inventory, the framework mapped onto the security review you already run, and a quarterly cadence is a complete implementation at that scale. Depth should track risk, not headcount, so a small agency running one rights-affecting system needs full rigor on that system and very little on the rest.
Should we set a target of zero critical escalations?
As a goal for your systems, yes. As a measure applied to the people who report problems, no, and the distinction is not academic. A team that learns escalations count against the programme stops filing them, the metric improves, the risk grows, and the first anyone hears of a serious issue comes from outside the agency. Track escalation counts for what they reveal about system health, and make it explicit in writing that raising an issue is never held against the person who raised it.
Skill.re