←
AI for Government
Strategic · M23 · lesson 23 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Drafting Agency AI Policies
📖
now learning

Drafting Agency AI Policies

15 min

Priya Nair, deputy CIO of a state transportation department, was handed a four-word directive by her secretary after a news story broke about a neighboring state's AI tool wrongly suspending drivers' licenses: "Get us a policy." She had two weeks before a legislative hearing. What she did not want to produce was the thing most agencies produce under deadline pressure: a glossy two-page statement of principles that said the department would use AI "responsibly and ethically" and answered none of the questions a frontline employee, a vendor, or an auditor would actually ask. A policy that cannot be enforced is not a policy. It is a press release.

This lesson is about drafting an agency AI policy that actually governs behavior: one that tells a caseworker what they may and may not do, tells a vendor what they must deliver, and gives an auditor something to check against. We will build it the way Priya built hers, section by section, and then step back to the wider federal picture: how policy is layered across an agency, what document elements make it binding, how it gets drafted and reviewed, and how it gets enforced once published.

What a real policy does that a principles statement does not

Principles are necessary but they are not sufficient. "Be fair" is a principle. "Any AI system that affects a resident's eligibility for a benefit must undergo a documented bias evaluation before launch, and a human must review every denial" is a policy. The difference is that the second one creates obligations a specific person owns, with consequences if they are skipped. That second sentence is deliberately demanding, and note that it is written to err toward protection rather than convenience: it commits the agency to human review of every adverse outcome, not a sample.

Agency AI policy is not a glossy external document. It is the load-bearing operating system of the AI program. It is what allows a Chief AI Officer to say yes or no to a use case and have the decision stick. It is what gives a contracting officer language to cite when a vendor is not performing. It is what an auditor from the Government Accountability Office asks to see first.

Above all it is what converts abstract obligations into concrete actions staff can take tomorrow morning. Those obligations arrive from several directions at once: OMB Memorandum M-24-10, Executive Order 14110, the Federal Information Security Modernization Act, and the NIST AI Risk Management Framework, which is voluntary rather than binding but widely expected. Because executive actions are revised and revoked over time, confirm the current status of any one of them before citing it as authority in a document you sign. Where policy is well drafted, program managers operate inside a known envelope. Where it is missing, every AI decision becomes a one-off negotiation: slow, inconsistent, and legally fragile.

A workable policy answers four questions for any AI use in the agency. What is allowed and what is forbidden? Who decides and who is accountable? What must happen before, during, and after deployment? And what occurs when something goes wrong? If your draft does not answer all four, it will not survive contact with a real situation, and certainly not a legislative hearing. Write it so the person reading it at four o'clock on a Friday knows exactly what they are allowed to do.

Three ways agency AI policy fails

Policy failure in government AI falls into three recognizable buckets, and each has a different remedy. The first is the phantom policy: a document that exists on the agency website but does not map to any actual decision process. The classic presentation is a policy that assigns authority to an AI governance board which has not met in months, where staff who are interviewed cannot name the board's decision authority. The remedy is structural rather than editorial: re-charter the board, tie each policy to a named decision authority, and publish decision logs so the mapping between document and practice is visible.

The second is the inherited policy, where an agency copies language from another agency or a commercial framework without adapting it to local authorities. The result references controls the agency does not own or use. It fails audit and it confuses staff, who cannot find the systems the policy talks about. The third is the frozen policy, still describing machine learning as a single technology with no reference to generative AI, agentic systems, or the rights-impacting and safety-impacting taxonomy that OMB Memorandum M-24-10 introduced in 2024. A frozen policy is worse than no policy, because it signals to staff that the old categories still apply when they do not.

The contrast is a drafting discipline rather than a talent difference. Good policy starts with clarity about authority: what statute, regulation, or executive action is this document implementing? The set an agency under scrutiny may be asked to produce on demand runs from M-24-10 and the NIST AI RMF through FISMA and, increasingly, the audit expectations forming around ISO/IEC 42001. It moves through explicit stakeholder engagement that gets the Office of General Counsel, the Senior Agency Official for Privacy, the Chief Information Security Officer, the civil rights office, and affected mission owners on record before publication. It produces plain language a new employee can follow without a glossary. And it establishes a living enforcement and revision process, so that the policy keeps pace with both the technology and the guidance.

The five-level policy taxonomy

Federal AI policy works best organized into five levels, each with a distinct purpose, audience, and signatory. The reason for the layering is that a single document cannot simultaneously establish agency posture and tell a frontline benefits adjudicator what to do on Tuesday. Trying to make one document do both produces something too vague to follow and too specific to endorse. Priya's state department needed only three of the five, but knowing the full structure told her which of her draft's sections belonged in different documents entirely.

LevelDocumentWhat it doesWho signsRevision
OneSecretary's or Administrator's DirectiveShort document establishing agency AI posture and naming the Chief AI Officer as senior accountable official under OMB M-24-10; cites the executive action, the memorandum, and the agency's statutory mission it implementsSecretary or AdministratorInfrequent; once every two years is typical
TwoChief AI Officer policy memoOperationalizes the directive: use-case inventory process, rights-impacting and safety-impacting classifications, training requirements, incident reportingChief AI OfficerTypically annual
ThreeAI governance board charter and operating proceduresBoard membership, decision authority, quorum, escalation; how use cases are intaken and reviewed, what artifacts approval requires, how decisions are documentedChief AI Officer signs the charter; the board approves its own proceduresAs the board's remit changes
FourProgram standard operating procedureTranslates levels one through three into actions a frontline team can execute in a specific program areaProgram's Assistant Secretary or deputyWith program change
FiveSystem concept of operationsFor one system: technical design, operational plan, monitoring plan, human-in-the-loop design, appeal path, training, and incident responseProgram manager, with concurrence from the Chief AI Officer's officeWith system change

Failing to maintain all five levels produces specific and diagnosable pathologies, which is what makes the taxonomy useful as a diagnostic rather than just a filing scheme. If the directive is missing or vague, the agency has no unified posture and program offices drift apart. If the Chief AI Officer memo is missing, staff cannot determine what counts as rights-impacting or how to report an incident. If the board charter is weak, decisions feel arbitrary and get relitigated. If procedures are missing, frontline staff improvise. If system-level concepts of operations are missing, the agency cannot explain to an inspector general why a given system is operating the way it is. An auditor asked to evaluate AI governance walks this taxonomy from top to bottom, and the gaps become the findings.

The seven components, built around one running case

Within any one policy document, seven subject areas have to be governed. We will carry one realistic use through every section: the transportation department wants to deploy an AI tool that flags license applications for possible fraud. It is a useful example because it touches eligibility, due process, bias, and public trust all at once, exactly the kind of high-stakes use a policy exists to govern. These seven are what the policy must govern; the section on document elements later covers how the document itself is structured, which is a different cut of the same problem.

1. Scope

Scope says what the policy covers and what it does not. Be concrete. Priya's scope covered "any system that uses machine learning or generative models to inform, recommend, or make a decision affecting a member of the public or a department employee." It explicitly excluded ordinary spreadsheet formulas and rule-based software, so the policy did not collapse under its own breadth. A scope that tries to cover everything governs nothing, and a scope that is silent about exclusions invites every program to argue itself out of coverage.

2. Governance

Governance names who is in charge. This section establishes the roles: who can approve a new AI use, who must be consulted, and who holds final accountability. For the fraud-flagging tool, the policy named the deputy CIO as approver, required sign-off from privacy and legal, and made the program director accountable for outcomes. Critically, it placed final accountability with a named role, not "the department," because a policy where everyone is responsible means no one is. Avoid the phrase "as appropriate" entirely; it is where accountability goes to disappear.

3. Risk

Risk is where the policy connects to a recognized method instead of improvising. Priya anchored hers to the National Institute of Standards and Technology AI Risk Management Framework, which is voluntary rather than legally binding but is widely expected as the reference method, and which organizes the work into four plain activities: govern, map, measure, and manage. In policy terms that became a requirement to classify every AI use by impact level, and to require heavier review as impact rises. The fraud tool, affecting people's licenses, landed in the highest tier and triggered the full review path.

4. Procurement

Most government AI is bought, not built, so the policy must reach into purchasing. This section sets what the agency demands from vendors: the right to test for bias, access to documentation explaining how the system works, data-handling terms that keep resident data under agency control, and an exit clause so the agency is not trapped. These requirements ride on top of standard acquisition rules like the Federal Acquisition Regulation, translating "be careful with vendors" into specific contract clauses a contracting officer can insert and, when performance slips, actually cite.

5. Ethics

Ethics in a government policy is not vague aspiration; it is enforceable commitment. Priya's policy required three things for any public-facing AI: disclosure that AI is being used, a meaningful human appeal of any adverse decision, and accessibility compliance under Section 508 so the tool works for people with disabilities. For the fraud tool, that meant an applicant flagged by the system would be told, and could request human review before any action. Where a use is rights-impacting, OMB M-24-10 makes civil rights office engagement non-optional, and the policy should say so rather than leave it to judgment.

6. Monitoring

A policy that ends at launch is half a policy. This section requires ongoing checks: tracking accuracy and bias over time, logging decisions, and re-reviewing high-impact systems on a set schedule. The fraud tool was set for quarterly bias audits, because a model that is fair on launch day can drift as the world changes around it. Be specific about what is monitored, since monitoring only surfaces the failure modes you chose to watch for, and a dashboard with no fairness metric will report a healthy system all the way through a fairness failure.

7. Enforcement

Enforcement is the component agencies most often omit, and its absence is what turns a policy into a press release. This section states what happens when the policy is not followed: who can pause or shut down a non-compliant system, what the consequences are for deploying without approval, and how exceptions are granted and documented. Without teeth, the other six sections are suggestions. The strongest available pattern for exceptions is the one M-24-10 uses, where a rights-impacting use case can be exempted only through a documented risk acceptance signed by the Chief AI Officer.

The nine elements every policy document must contain

The seven components above are what the policy governs. The nine elements below are how the document itself is built, in roughly this order, and a document missing any of them fails in a predictable way. First, Authority, citing the statutes, regulations, executive actions, and prior agency documents the policy implements; this section is what makes the policy binding rather than advisory. Second, Scope, identifying the systems, use cases, programs, and personnel covered, explicit about exclusions as well as inclusions. Third, Definitions, reusing federal guidance and framework terminology rather than reinventing it, with agency-specific terms clearly marked as such.

Fourth, Roles and Responsibilities, naming the Chief AI Officer, Chief Information Officer, Chief Information Security Officer, Senior Agency Official for Privacy, Chief Data Officer, civil rights office, general counsel, and mission owners, each with specific responsibilities. Fifth, Requirements, the substantive rules: what staff must do, what is prohibited, what must be documented, what must be reviewed. Use "shall" for mandatory items and "should" for recommended practice, and never mix the two ambiguously in one sentence. Sixth, Exceptions and Waivers, describing the limited circumstances in which requirements may be relaxed and precisely who may grant relief.

Seventh, Enforcement, covering compliance monitoring, training requirements, attestation, and consequences for non-compliance; enforcement language should hook into the agency's existing accountability systems rather than inventing parallel ones nobody will run. Eighth, Revision, including a cadence, a trigger list, and the named authority who may revise. Ninth, Effective Date and Signatures, with the appropriate official signing and unambiguous dates. Findings against federal AI policy gaps map, with striking regularity, onto one of these nine elements being absent or hollow.

The seven-step drafting process

Durable policy comes out of a defined process rather than a heroic weekend. Step one is scoping, which names the policy's purpose, the authority it implements, its audience, and its level in the taxonomy. Scoping produces a one-page memo approved before substantive drafting begins, which prevents weeks of work aimed at the wrong level or audience. Step two is the first draft, produced by the policy lead with input from at least one mission owner, the general counsel's office, and the security office. This draft is deliberately rough. Its purpose is to give stakeholders something concrete to react to, not to be good.

Step three is legal review, coordinated by the Office of General Counsel, which checks alignment with statute, prior agency policy, executive branch guidance, and litigation posture, and flags any language that would create enforceable rights in third parties, something internal policy typically avoids. Step four is privacy, records, and civil rights review. The Senior Agency Official for Privacy reviews for Privacy Act implications and for the privacy impact obligations under Section 208 of the E-Government Act. The agency records officer reviews records management obligations under 44 U.S.C. chapter 31. The civil rights office reviews for Title VI, Section 504, and any agency-specific civil rights obligations, and for rights-impacting uses M-24-10 requires that engagement, which makes this review non-optional rather than discretionary.

Step five is executive review, bringing the draft to the AI governance board and, for high-visibility policies, to the deputy secretary or administrator. The board earns its place here as an integration point that catches cross-mission conflicts nobody in a single program would see. Step six is public engagement where applicable: for rights-impacting policies, agencies increasingly post a draft for a limited comment window, gathering outside expertise and demonstrating procedural legitimacy. Step seven is publication and socialization, which means posting in the policy system of record, notifying affected staff, training program leads on new obligations, and scheduling the first compliance check.

Each skipped step produces its own predictable pathology. Skip legal review and you publish something unenforceable. Skip privacy review and you carry an unassessed Privacy Act risk into production. Skip civil rights review on a rights-impacting use and you have removed the check most likely to catch the harm the policy exists to prevent. Skip socialization and you have a technically valid policy that nobody follows, which in an audit is indistinguishable from having no policy at all.

Plain-language drafting and readability

The Plain Writing Act of 2010 requires federal agencies to use plain language in communications with the public, and plain-language expectations have since been extended to internal policy as well. For AI policy this is not optional polish. It is the difference between a policy staff can follow and a shelf ornament. The workable targets are modest and testable: sentences averaging under twenty words, active voice as the strong majority of constructions, technical terms defined at first use in one plain sentence rather than a technical footnote, and an overall readability level around eighth to tenth grade for internal policy.

Four specific practices separate the best-drafted policies from the rest. They define "AI" and "AI system" once, at the start, borrowing the definitions from federal guidance and the risk framework rather than inventing new ones. They use consistent subject-verb patterns for requirements, so that every obligation reads as "the Chief AI Officer shall," "program managers shall," or "users may." They include worked examples at the end of key sections, showing how a requirement applies to a use case the agency actually runs. And they include a decision flow or checklist for the most common cases, so a new program manager can orient in minutes instead of reading the whole document.

The consistent subject-verb pattern is worth one caution. It makes it easy for an auditor to check whether a requirement was written and who it binds. It does not make it easy to check whether the requirement was met, which is a separate exercise requiring evidence rather than reading. The anti-patterns to avoid are equally concrete: undefined acronyms, passive voice that hides the actor as in "it shall be ensured that," nested conditionals more than two layers deep, cross-references to external documents without a date and version, and vendor-specific terminology that will be obsolete by the next product cycle.

A usable artifact: the agency AI policy skeleton

Below is a fill-in skeleton you can adapt directly. Each row has its governing question and a sample clause drawn from the fraud-tool case. Replace the bracketed items with your agency's specifics, route it past legal, privacy, records, and civil rights, and you have a draft that governs rather than decorates.

SectionQuestion it answersSample clause to adapt
ScopeWhat is covered?"This policy applies to any system using machine learning or generative AI to inform or make decisions affecting the public or staff. It excludes rule-based and standard office software."
GovernanceWho decides and is accountable?"New AI uses require approval by [the CIO], consultation with privacy and legal, and a named accountable program owner."
RiskHow is risk assessed?"Every AI use is classified by impact (low, moderate, high) following the NIST AI RMF. High-impact uses require full review and an approved risk register entry."
ProcurementWhat do vendors owe us?"AI contracts must grant rights to bias testing, system documentation, agency-controlled data handling, and a no-penalty exit clause."
EthicsWhat do we owe the public?"Public-facing AI must disclose its use, provide a human appeal of adverse decisions, and meet Section 508 accessibility."
MonitoringHow do we watch it over time?"High-impact systems undergo quarterly bias and accuracy audits and full re-review annually. All decisions are logged."
EnforcementWhat if the policy is broken?"[The CIO] may pause or decommission any non-compliant system. Deployment without approval is a documented policy violation. Exceptions require written, reviewed authorization."

Enforcement, evidence, and revision cadence

A policy that is not enforced is a suggestion. Enforcement in practice rests on five mechanisms working together, and each one alone is weak. Training: every covered employee completes training on the policy's requirements, with completion tracked. Attestation: for higher-sensitivity obligations, an annual affirmation that the employee understands and will comply. Compliance monitoring: the Chief AI Officer's office or internal audit periodically samples use cases and verifies adherence, reporting findings to the governance board. Incident-driven review: any AI incident triggers a compliance check. And an evidence package you can hand an auditor on request.

Be clear-eyed about what each mechanism proves. Tracked training completion proves attendance, not comprehension. An attestation log proves that someone affirmed compliance, not that they complied, though it does establish that they were on notice, which matters. Periodic sampling finds problems in the sample, and only the kinds of problems the sample was designed to surface. These are real controls and they are worth building, but describing them to leadership as assurance that the policy is being followed overstates them. They bound your exposure and they generate evidence. They do not guarantee compliance.

Mature agencies maintain a perpetual evidence folder per policy: the signed document, training logs, attestation logs, compliance reports, incident cross-references, and revision history. Revision is the other half of enforcement. Annual review is the common cadence, and revision triggers are additive rather than alternative: a new statute, executive action, or guidance memorandum; a significant incident; a new audit finding; or a major technology shift, as when the arrival of widely deployed generative AI in 2023 drove a wave of agency policy revision the following year. Revision should incorporate new obligations while preserving the stable core rather than rewriting everything.

A revision log at the end of the policy captures the history and answers the question auditors actually ask, which is what the policy said at the time of a specific decision rather than what it says now. Policies that revise chaotically without logs are nearly as audit-fragile as frozen ones. The final discipline is consequence management: when a program office fails a compliance check, the Chief AI Officer's office documents the failure, sets a remediation plan, and tracks it to closure, with repeated non-compliance escalating to the deputy secretary. Agencies that treat compliance as a living process produce policy that shapes behavior. Agencies that treat it as a formality produce policy that shapes paper.

How to get it adopted, not just written

Priya did one more thing that separated her policy from the shelf-ware versions. She tested the draft against three real scenarios before the hearing: the fraud tool, a back-office document summarizer, and a public chatbot. She walked each one through all seven components and asked her team, "does the policy tell us clearly what to do here?" Where it did not, she fixed it. A policy stress-tested against concrete cases is one you can defend under questioning, because you have already answered the hard questions in private rather than discovering the gaps in front of legislators.

The stress test also does something a legal review cannot. Legal review asks whether the policy is defensible. Scenario walkthrough asks whether it is followable, which is a different question with a different failure mode. A clause can be perfectly lawful and completely unactionable, and the person who discovers that is either you in a conference room this month or a caseworker at four o'clock on a Friday next quarter. Pick which. Agencies that adopt plain-language drafting and scenario testing as standing disciplines find that staff follow policy they can understand and apply, which is a low bar that a surprising amount of published policy fails to clear.

Anti-Patterns

  • The principles statement wearing a policy's clothes. A document full of commitments to fairness and responsibility with no named owner, no required action, and no consequence. It survives the press cycle and dies on contact with the first real question. Test every paragraph by asking who must do what by when, and cut or rewrite any that cannot answer.
  • The inherited policy. Copying another agency's language wholesale, so the document references boards you do not have, controls you do not run, and authorities you do not hold. It fails audit and confuses staff simultaneously. Borrow structure freely; borrow authorities never, because your authorities are what make your policy binding.
  • The frozen policy. Still treating machine learning as one technology, silent on generative and agentic systems and on the rights-impacting and safety-impacting distinction. It is worse than no policy, because it tells staff the old categories still hold. Set a revision cadence with named triggers and a revision log, and treat a missed review as an incident.
  • Treating a completed checklist as compliance. Training completion records attendance, attestation records that someone clicked, and a sampled audit records what happened in the sample. Each is genuine evidence and none is proof the policy was followed. Say this explicitly to leadership before someone else says it to an auditor.
  • Enforcement with no named enforcer. A policy that says non-compliant systems "may be paused" without saying who may pause them, on what evidence, and over whose objection. Name the role, give it the authority in writing, and route the exception path through the same role so refusals are visible.
  • Mixing shall and should in one breath. "Programs should document their evaluation and shall report incidents as appropriate" leaves an auditor and a caseworker with different readings of the same sentence. Keep mandatory and recommended language structurally separate, and strike "as appropriate" wherever it modifies an obligation.

Practice Prompts

  • Audit your taxonomy. For your agency, identify which of the five levels actually exist in signed form: directive, senior official memo, board charter and procedures, program procedures, and system concepts of operations. For each gap, write the one sentence describing the pathology it is currently producing.
  • Rewrite one principle into a requirement. Take a sentence from your agency's current AI statement, something like "we will use AI fairly," and rewrite it as an obligation naming who must do what, before which event, documented how, and reviewed by whom. Note how much longer it gets. That length is the point.
  • Walk a real case through all seven components. Pick one AI use your agency runs or is considering. Walk it through scope, governance, risk, procurement, ethics, monitoring, and enforcement, and record every point where your current policy does not tell you clearly what to do.
  • Test the nine elements. Take your existing policy document and check for authority, scope, definitions, roles and responsibilities, requirements, exceptions and waivers, enforcement, revision, and effective date with signatures. Which are missing, and which are present in name but hollow in content?
  • Build the evidence folder. For one existing policy, assemble what you would hand an auditor tomorrow: signed document, training logs, attestation logs, compliance reports, incident cross-references, revision history. Write down what you could not find, and how long it would take to reconstruct.

Reflection

Set this down in writing rather than turning it over mentally. If your agency's AI policy were tested tomorrow by a caseworker facing an ambiguous case rather than by a lawyer reading for defensibility, what would break first? Which of the three failure modes does your current document most resemble: phantom, inherited, or frozen? And if it is more than one, which would you fix first given that you cannot fix all of them at once?

Then consider the harder political question. Enforcement is the component agencies most often omit, and the reason is rarely oversight. It is that naming who may pause a program's system creates a conflict someone would rather not have in writing. Who in your agency would have to accept that authority for your policy to have teeth, what would they want in exchange, and what would you be willing to give? A policy nobody was uncomfortable signing is usually a policy that asks nothing of anyone.

Glossary

  • Policy taxonomy. The layered set of documents through which an agency issues AI policy, from a short directive at the top to a single system's concept of operations at the bottom, each with its own signatory.
  • Concept of operations. The system-level document combining technical design, operational plan, monitoring plan, human-review design, appeal path, training requirements, and incident response for one AI system.
  • Rights-impacting use. An AI use affecting an individual's rights, opportunities, or access to benefits, carrying heavier minimum practices and mandatory civil rights review under federal guidance.
  • Authority section. The opening element citing the statutes, regulations, executive actions, and prior agency documents a policy implements. It is what makes the document binding rather than advisory.
  • Attestation. A periodic affirmation by a covered employee that they understand and will comply with a policy. Evidence that they were on notice, not evidence that they complied.
  • Revision trigger. A defined event, such as new guidance, a significant incident, or an audit finding, that requires the policy to be reopened regardless of where the scheduled review cycle stands.
  • Waiver. A documented, signed relaxation of a specific requirement for a specific use, granted by a named authority and recorded so it can be reviewed rather than quietly renewed.
  • Evidence package. The assembled record for one policy that an agency can hand an auditor: signed document, training and attestation logs, compliance reports, incident cross-references, and revision history.

Closing

Priya made her hearing. What got her through it was not eloquence about responsible AI. It was that when a legislator asked "who decides whether one of these tools can go live," she could name a role, and when asked "what happens if someone skips that," she could name a consequence, and when asked "how would you know," she could describe the log. None of those answers require a large staff or a big budget. They require a document that was written to be followed rather than to be published.

The recurring temptation under deadline is to produce the two-page statement, because it can be written in an afternoon and nobody objects to it. The recurring cost is that six months later, every AI decision is still a one-off negotiation and the statement is not helping. Layer your policy properly, build in all nine document elements, run the seven-step process even when compressed, write it in plain language, and give it enforcement with a named enforcer.

Then keep it alive. The single most reliable predictor of whether an agency's AI policy still works two years from now is whether anyone owns revising it, with triggers written down and a log that shows the history. Policy is not a document you finish. It is a process you run, and the document is just its current output.

Key Takeaways

  • A principles statement is not a policy. Real policy creates specific obligations that a named person owns, with consequences if they are skipped.
  • Answer four questions. What is allowed, who decides, what must happen across the lifecycle, and what occurs when something goes wrong.
  • Layer policy across five levels. Directive, senior official memo, board charter and procedures, program procedures, and system concept of operations, each with its own signatory and revision rhythm.
  • Govern seven components and build nine document elements. Scope, governance, risk, procurement, ethics, monitoring, and enforcement are what the policy governs; authority through signatures is how the document is constructed.
  • Anchor risk to a recognized framework. Classifying uses by impact level using the NIST AI RMF, which is voluntary rather than binding, replaces improvisation with a defensible method.
  • Reach into procurement. Since most government AI is bought, turn caution into specific contract clauses: bias testing, documentation, data control, and an exit ramp.
  • Run all seven drafting steps. Scoping, rough draft, legal, then privacy, records and civil rights review, executive review, public engagement where applicable, and socialization. Each skipped step has its own predictable failure.
  • Enforcement is the part most agencies omit. Name who can pause a system, the consequences of skipping approval, and how exceptions are documented and reviewed.
  • Do not oversell your controls. Training records attendance, attestation records notice, sampling records the sample. They bound exposure and generate evidence; they are not proof of compliance.
  • Stress-test before you publish. Walk concrete use cases through every component so you find the gaps in private, not at a legislative hearing.

Frequently Asked Questions

How long should an agency AI policy be?

Length follows level rather than ambition. A directive establishing posture and naming the accountable official runs one to three pages, because its job is to be signed by someone very senior and cited by everything below it. A senior official memo operationalizing that directive runs considerably longer, because it has to define classifications, inventory processes, training requirements, and incident reporting. If your single document is trying to do both jobs, that is usually the reason it feels simultaneously too long and too vague.

Can we adopt another agency's policy to save time?

Borrow its structure, its section order, and its plain-language patterns freely. Do not borrow its authority section, its role names, or its control references, because those are what tie the policy to your agency's actual legal footing and organization. An inherited policy that cites boards you do not have and controls you do not run will fail its first audit and, more immediately, will confuse the staff who try to follow it and cannot find the things it describes.

What if our agency has no Chief AI Officer?

The taxonomy still works; the signatures move. What matters is that each level has a named accountable role with real authority, not that the role carries a particular title. Many state and local agencies place these responsibilities with a deputy CIO, a chief data officer, or a designated program executive. Write the actual title into the policy rather than a generic one, and update it when the organization changes, since a policy assigning authority to a vacant or renamed role is a phantom policy waiting to be found.

How often should the policy be revised?

Annual review is the common cadence, but the cadence is the floor, not the mechanism. What keeps policy current is the trigger list: new guidance or statute, a significant incident, a new audit finding, or a substantial technology shift. Any of those reopens the document regardless of where the calendar stands. Keep a revision log at the end, because the question an auditor asks is what the policy said at the time of a specific decision, and only the log can answer that.

Does a strong policy protect the agency if an AI system causes harm?

It helps, and it is worth being precise about how. A well-drafted, followed, and evidenced policy demonstrates that the agency identified the risk, assigned accountability, required review, and acted on what it found. That is materially better than the alternative in any oversight or legal setting. What it does not do is make a harmful outcome acceptable or transfer responsibility away from the agency. A completed process has never made a wrong decision right, and the policy's value comes from the decisions it changed before deployment, not from its existence in a binder afterward.