←
AI for Government
Visionary · M32 · lesson 32 of 46 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Multi-Level Government AI Governance
📖
now learning

Multi-Level Government AI Governance

15 min

Maria Delgado, deputy secretary of a state health and human services agency, thought she had solved AI governance. Her department had a board, a policy, a risk register. Then a fraud-detection tool, built by the state, deployed by a county social-services office, fed by federal benefits data, wrongly flagged 1,200 families for benefits review in six weeks. Each level pointed at the others. The state said the county misconfigured it. The county said the state never trained them. The federal program office said its data-sharing agreement was silent on AI. No one was lying. No one was accountable. The families waited, some losing benefits for 30 days while the finger-pointing ran its course. "We didn't have a governance problem inside any one agency," Maria told a national working group. "We had a governance gap between them." This lesson is about governing AI across the seams of government.

Multi-level government AI governance is the work of coordinating how AI is built, bought, deployed, and held accountable across federal, state, local, and tribal governments, levels that share missions and data but answer to different authorities, budgets, and laws. For an agency head or national AI leader, this is the hardest governance problem there is, because no one is in charge of all of it, and the failures fall through exactly the cracks that no single level owns. The mechanics of aligning federal, state, and local programs are developed in Federal/State/Local AI Alignment; this lesson is about designing accountability that survives the moment everyone starts pointing.

Why the Seams Fail

Maria's fraud tool failed not at a level but between levels. That distinction matters, because every remedy her team first proposed was an internal one: a better risk register, a stricter review board, more thorough documentation. None of them would have prevented anything, because the failure happened in the space between three organizations, each of which had done its own job adequately. Three structural features of the system produce these gaps, and a leader has to design around all three at once.

  • Split authority. Federal agencies set program rules and fund them; states administer them; counties and cities deliver services to actual people; tribal nations are sovereign governments rather than subordinate units. Each level has a genuine mandate and none of them was designed to supervise an algorithm running at another level.
  • Tangled funding. Money flows down with conditions attached. A federal grant funds a state program that a county runs, and each transfer carries its own terms. When AI enters that chain, the question of who is responsible for governing it is rarely written down anywhere, because the agreements predate the technology.
  • Capability gaps. A federal agency may have AI expertise that a rural county will never have, and expertise does not travel with software. Handing a county a powerful tool without the capacity to configure, monitor, and challenge it is how Maria's 1,200 families got flagged.

Being Honest About Who Can Tell Whom What

Before choosing a governance model, get the authority question right, because most cross-level designs fail on a wrong assumption about it in one direction or the other. The convenient story is that no level can tell another level anything, which excuses inaction. The equally convenient opposite story is that a federal policy automatically governs everyone who touches a federal program, which produces a paper baseline nobody is actually bound by.

The accurate picture is messier and more useful. Federal bodies rarely command states directly, but they routinely attach conditions to the money and the data they provide, and those conditions bind whoever accepted them. Many states do exercise authority over their local governments, though how far that authority reaches depends on state law and on how the locality is chartered, which is why a rule that holds in one county may be unenforceable in the next one over. Tribal nations occupy a different category entirely and are engaged government to government. So the practical question is never whether you can order another level. It is which instrument you actually hold: a funding condition, a data-sharing agreement, a contract, a certification, or nothing but persuasion.

Maria's working group made this its first exercise. For each level involved in the fraud tool, they wrote down the specific instrument through which any obligation could travel, and left the box empty where there was none. The empty boxes were the governance gap, drawn on one page, and they were not where anyone had expected them to be.

Three Governance Models, and When Each Fits

There is no single right structure for governing across levels. Maria's working group settled on three recurring patterns and a rule for choosing among them, which is to pick the model your available instruments can actually support rather than the one that sounds most rigorous.

The first is centralized, in which one level, usually federal or state, sets binding rules that everyone follows. It produces consistency and comparable evidence, and it strains against the limits of one level's authority over another while ignoring genuine local variation. The second is federated, in which a shared baseline applies to everyone with freedom to go further locally. It is the most common workable answer, because it concentrates enforcement effort on a short list of non-negotiables. The third is coordinated, in which levels keep their own authority but agree to align through shared standards, data agreements, and joint bodies. It respects sovereignty and independence, and it moves slowly and depends on goodwill that evaporates under political pressure.

Maria's choice for the fraud tool, in hindsight, was a federated model: a short, non-negotiable baseline that every deploying office must meet, covering pre-deployment testing, human review of adverse determinations, and a working appeal path, with room for local adaptation above that floor. Tribal nations were engaged as the sovereign governments they are, through government-to-government agreements, rather than handed a state mandate with a comment period attached.

Govern Where the Data and the Harm Live

The cleanest way to assign accountability across levels is to follow two things: the data and the harm. In Maria's case, federal benefits data flowed to a state-built model deployed by a county office, and the harm landed on families in that county. Those are the two ends of the chain, and they are where the instruments actually exist. Data moves under agreements that someone signs. Harm lands in a jurisdiction whose officials answer to the people it landed on.

A workable governance design therefore names, for each of those, a single accountable owner, and writes the AI obligations directly into the agreements that move the data and the money. Everything in the middle, the model build, the vendor, the integration work, hangs off those two anchors. This also resolves the argument that consumed six weeks after the failure. If the agreement governing the federal data had said what testing was required before that data could feed an automated determination, and the county's own service commitments had said who fixes a wrong determination and how fast, there would have been nothing left to argue about except whether each party had done what it signed.

How a Federal Standard Actually Travels Down the Chain

This is where federal frameworks become practical tools rather than paperwork, and where leaders most often overestimate what they have. The federal AI policy direction issued in 2024, OMB Memorandum M-24-10, Advancing Governance, Innovation, and Risk Management for Agencies' Use of Artificial Intelligence, directs federal agencies to identify AI that affects people's rights or safety and to apply concrete protections before deploying it, including testing, human oversight, and a path to contest a decision. A benefits fraud tool plainly falls in that territory.

What it does not do is reach a state or a county by itself. It is guidance to federal agencies, and it does not automatically bind a grantee, a contractor, or another level of government, however sensible its requirements are. That gap is precisely the trap: a state can adopt the language of a federal memorandum in its own policy documents and believe it has imported an obligation, when it has imported a vocabulary. The obligation exists only where someone wrote it into a specific instrument that a specific party accepted.

So the fix Maria pushed was mechanical rather than rhetorical. Rather than announcing that the federal standard applied, she worked to have the substance of the rights-protecting baseline written into the actual data-sharing agreement, so that any downstream party using that federal benefits data in an automated determination had accepted the baseline as a term. Whether a program office can impose such a term, and how far it can reach, varies by program and by the authority behind it, which is a question to answer with counsel before you promise the protection publicly. Done properly, the standard travels with the data. Assumed rather than written, it travels nowhere.

Sovereignty Is Not a Consultation Step

Tribal nations are sovereign governments, and the most common failure in cross-level AI work is to treat that sentence as a courtesy rather than a design constraint. A state that sends a tribal government its new AI policy for comment has not engaged a sovereign; it has cc'd one. Engagement runs government to government, on terms both parties negotiate, and it covers the questions that matter most in AI work: whose data this is, where it may be stored, who may build models on it, and who decides when a model touching tribal citizens is withdrawn.

Practically, this means the sovereignty question belongs at the start of a project rather than at the compliance review, because the answers change the architecture. If data cannot leave a jurisdiction, that is a design input. If a model may be used for one purpose and not another, that is a contract term. Discovering either after a system is built produces the worst outcome, which is a service that works and cannot lawfully or legitimately be used.

Capacity Is a Governance Control, Not a Training Line Item

Maria's hardest lesson was that governance obligations placed on a level that cannot meet them are not governance. The county office running her fraud tool had no one whose job included understanding what the model was doing, no threshold to tune it against, and no basis on which to challenge an output that looked wrong. It was accountable on paper and helpless in practice, which is how 1,200 households became a caseload nobody had planned for.

Treat the capacity check as a gate rather than as a support activity. Before a tool goes live at another level, establish whether the deploying office has someone who can configure it correctly, someone who can recognize and escalate a wrong output, and a route to expert help when something ambiguous happens late in the day. If those are absent, the answer is not more training slides. It is either supplying the capability as a shared service, narrowing the tool so that it cannot do the damaging thing unsupervised, or not deploying at that site yet. Each of those is a legitimate choice. Deploying anyway and calling the local staff accountable is not.

Nobody Was Watching the Whole Thing

The detail in Maria's case that troubled her longest was not the misconfiguration. It was that the pattern had been visible for weeks and no one had been positioned to see it. The county saw its own caseload rising and assumed a busy period. The state saw a tool performing within the ranges it had tested. The federal program office saw nothing at all, because its relationship was with the data and not with the outcomes. The aggregate picture, which was alarming, existed in no one's dashboard because the aggregate crossed three organizations.

Cross-level monitoring is therefore a design problem rather than a technical one. Someone has to be assigned the whole view, told what would count as alarming, and given a route to act on it. Decide three things explicitly before go-live: which signals each level reports upward or sideways, who assembles them into one picture, and what that person is authorized to do when the picture looks wrong. Without the third, you have built a reporting obligation rather than a control, and the reports will accumulate quietly while everyone assumes someone else is reading them.

Keep the signal list short and outcome-oriented, because a long list guarantees that nothing is examined. What usually matters is the rate at which the system produces adverse outcomes, how that rate is distributed across communities and offices, how many of those outcomes are overturned on appeal, and how long the overturning takes. A high overturn rate is the single most useful cross-level signal available, because it is measured after a human has looked, it is hard to argue with, and it is generated by the people the system affected rather than by the people who built it.

A Multi-Level AI Governance Responsibility Charter

Maria's lasting contribution was a one-page charter that any cross-level AI initiative must complete and sign before deployment. It forces the accountability questions into the open while there is still time to answer them, and it is deliberately short enough that a program manager will actually fill it in. Each row names a real official, not an office.

  1. System and scope. What does this AI do, who does it affect, and could it influence anyone's rights, benefits, or safety? If yes, the rights-protecting baseline applies.
  2. Data owner. Which level owns the data feeding this system, and what does its sharing agreement say about automated use? If the agreement is silent, fix the agreement first.
  3. Build owner. Who built or bought the model, and who is accountable for its testing, its measured accuracy, and its documented limits?
  4. Deployment owner. Which office actually runs it on real people, and who there is accountable for correct configuration and trained staff?
  5. Baseline compliance. Has the deploying office met the shared minimum, meaning pre-deployment testing, human review of adverse outcomes, monitoring, and a working appeal path?
  6. Instrument. Through which specific agreement, grant term, or contract does each obligation above actually bind the party it names? An obligation with no instrument is an intention.
  7. Capacity check. Does the deploying level have the skills to govern this tool? If not, what support, shared service, or narrowed scope fills the gap before go-live?
  8. Sovereignty. If tribal nations or independently chartered jurisdictions are involved, is there a negotiated government-to-government agreement rather than a mandate?
  9. Appeal and redress. When the system is wrong about a person, exactly who fixes it, how fast, and who tells the person?
  10. Escalation owner. When something fails across levels, who convenes the parties and makes the call? Name the person and the timeline.

Standard-Setting as the Quiet Lever

An agency head rarely commands other levels, but can shape them through standards, and this is the instrument most available to leaders who hold none of the others. Shared standards are how a federated or coordinated system stays coherent without anyone surrendering authority. The most durable ones are voluntary frameworks that levels adopt because they are useful rather than because they were forced, since a framework adopted willingly survives a change of administration in a way that an imposed rule usually does not.

The NIST AI Risk Management Framework is the common example, and it is voluntary and non-binding by design. Its usefulness here is linguistic before it is technical. When every level uses the same vocabulary of govern, map, measure, and manage to describe how it handles AI risk, the seams between them become inspectable, because a county and a federal program office can at least establish whether they mean the same thing by the sentence we tested it. Maria's working group adopted it as the shared spine of every cross-level agreement for exactly that reason. Note what it does not do: adopting a common framework does not verify anyone's claims, and it does not create an obligation. It makes claims comparable, which is the precondition for someone doing the verifying.

Anti-Patterns to Avoid

  • Announcing that a federal standard applies. Restating federal guidance in a state or local policy document and treating the obligation as imported. Guidance to federal agencies does not by itself bind a grantee or another level of government; the duty exists only where it was written into an instrument someone accepted.
  • Governing the level instead of the seam. Responding to a cross-level failure with a stronger internal review board. Every party may have done its own job adequately and the harm still lands, because nobody owned the space between them.
  • The consultation that mistakes itself for sovereignty. Circulating a finished policy to tribal governments for comment and recording it as engagement. Sovereign relationships are negotiated at the start, when the answers can still change the architecture.
  • Assuming a rule that holds in one locality holds in all of them. States vary in the authority they have over their political subdivisions, and localities vary in how they are chartered, so a uniform statement about what local governments must do is usually wrong somewhere in your own jurisdiction.
  • Accountability without capacity. Naming a deploying office as responsible for a tool nobody there can interpret, tune, or challenge. This converts a governance design into a plan for blaming the smallest party after the fact.
  • The unowned handoff. Shipping a model from the level that built it to the level that runs it with documentation and no named counterpart. The build owner disappears at the exact moment the deployment questions start.
  • Treating a shared framework as verification. Reading another level's statement of alignment with a voluntary framework as evidence that its system was tested. A common vocabulary makes claims comparable; someone still has to check them.
  • Leaving the escalation owner blank. Designing careful accountability for normal operation and none for the failure that crosses levels. That is the only scenario where the design is load-bearing, and it is the one everyone skips.

Practice Prompts

  • Draw the instrument map. Take one live cross-level AI system and, for each obligation you believe exists, name the specific agreement, grant term, or contract that carries it. Leave the box empty where there is none, and treat the empty boxes as your work plan.
  • Read the data agreement for silence. Pull the sharing agreement behind your most consequential data flow and find what it says about automated decision-making. Draft the clause you would need if it says nothing, before you need it.
  • Run the capacity gate. For a tool about to deploy at another level, establish by name who there can configure it, who can recognize a wrong output, and who they call when something ambiguous happens. If any of the three is missing, decide which of narrowing, supporting, or delaying you will do.
  • Write the failure script. Take the cross-level system you most rely on and write down, in advance, exactly who convenes the parties when it fails, on what timeline, and who is authorized to suspend it. Circulate it to those named and see whether they agree.
  • Complete the charter for a system already running. Fill in all ten rows for something in production rather than something proposed. The rows you cannot answer describe the exposure you already carry.

Reflection

Think about the AI systems in your own environment that cross a boundary between organizations, and ask what would happen tomorrow if one of them were badly wrong about the people it touches. Trace the conversation: who would be called first, who would have the authority to switch it off, who would have to tell the affected people, and how long each of those steps would take while the parties established whose problem it was. Then ask which of those answers is written down somewhere that a party has actually signed, and which of them is your assumption about how reasonable colleagues would behave under pressure. The distance between those two lists is the real state of your cross-level governance.

Glossary

  • Multi-level governance. Coordinating how a capability is built, bought, deployed, and held accountable across governments that share missions and data but answer to different authorities, budgets, and laws.
  • Centralized model. A structure in which one level sets binding rules for the others, producing consistency and comparable evidence at the cost of local fit and of straining the limits of authority.
  • Federated model. A structure with a shared mandatory baseline and freedom to exceed it locally, which concentrates enforcement effort on a short list of non-negotiables.
  • Coordinated model. A structure in which levels retain full authority and align voluntarily through shared standards and joint bodies, respecting independence at the cost of speed and durability.
  • Instrument. The specific agreement, grant term, contract, or certification through which an obligation actually binds a party. An obligation without an instrument is a stated intention.
  • Data-sharing agreement. The document governing transfer and use of data between organizations, and the most common place where automated use is unaddressed in agreements written before the technology arrived.
  • Government-to-government engagement. Negotiation between sovereigns on terms both parties set, as distinct from consulting a stakeholder on a decision already made.
  • Rights-impacting use. An application of AI that can affect a person's rights, benefits, or safety, and therefore attracts heightened requirements such as testing, human oversight, and a route to contest the outcome.
  • Escalation owner. The named person with authority to convene the parties and make a decision when a system fails across organizational boundaries.

Closing

The six weeks Maria's agency spent establishing whose fault it was were not a failure of goodwill. Every official in that argument was acting reasonably given what they had signed, and what they had signed did not contemplate a model. Cross-level governance is unglamorous for exactly this reason: it consists of writing down, in advance and in specific instruments, the answers to questions that only become urgent on the worst day. You will not get credit for it, because its success looks like an incident that resolved quietly in an afternoon. Start with the one system whose failure would be hardest to explain, and fill in the empty boxes.

Key Takeaways

  • The failures fall in the seams. Cross-level AI breaks down between governments rather than inside them, and internal remedies do not touch it.
  • Ask which instrument you hold, not whether you have authority. Federal bodies attach conditions to money and data; states often have authority over localities, though how far it reaches depends on state law and how the locality is chartered.
  • Choose the model your instruments can support. Centralized brings consistency, federated sets a short mandatory floor, coordinated respects independence and moves slowly on goodwill.
  • Federal guidance does not travel by itself. A memorandum directing federal agencies binds another level only where its substance was written into an agreement that party accepted, and how far that can reach varies by program.
  • Anchor accountability on the data and the harm. Those are the two ends of the chain where real instruments and real answerability already exist.
  • Name a real owner for each stage. Data, build, deployment, appeal, and escalation each need a named official before go-live rather than after the failure.
  • Capacity is a gate. Placing an obligation on an office that cannot interpret, tune, or challenge the tool is a plan for blaming the smallest party.
  • Sovereignty is negotiated at the start. Tribal nations are engaged government to government on terms both parties set, not consulted on a finished policy.
  • Shared standards make claims comparable, not true. A voluntary framework such as the NIST AI RMF gives independent levels one vocabulary; verification remains someone's job.

Frequently Asked Questions

We are a state agency. Can we require counties to follow our AI baseline? Sometimes, and the answer is a legal one specific to your state and to how the county is chartered, so establish it with counsel before you announce anything. What is more reliably available to you is the set of instruments you already control: the funding you pass through, the data you share, the systems you host, and the contracts you sign on their behalf. A baseline written into those reaches everyone who accepts them and does not depend on winning an argument about supervisory authority. Start there and treat broader mandates as a separate, slower project.

How much variation should a federated baseline tolerate? Keep the mandatory floor short enough that you can actually verify it and be genuinely permissive above it. The failure mode is a sprawling baseline that nobody checks, which produces the appearance of a shared standard and none of the effect. Pick the small number of things whose absence caused, or would have caused, real harm, usually pre-deployment testing, human review of adverse outcomes, monitoring, and a working appeal path, and enforce those seriously while letting local offices exceed them however they see fit.

What if the deploying level simply does not have the skills? Then you have three honest options and one dishonest one. You can supply the capability centrally as a shared service so the expertise sits where it exists. You can narrow the tool so that the actions requiring judgment are not available without support. You can delay deployment at that site until the gap is closed. The dishonest option is to deploy anyway with a training session and a policy naming the local office as accountable, which is a decision to place the risk on the party least able to carry it.

Nobody will convene the cross-level group. How do we get started without authority? Convene something narrower than governance. Pick one shared system and one concrete question, usually who does what when it is wrong about someone, and invite the specific people who would have to act. Cross-level bodies formed around a real operational problem tend to survive; those formed around an ambition to coordinate tend to hold two meetings. Once the group has answered a question that mattered, widening its scope becomes a proposal rather than a request for permission.