←
AI for Government
Strategic · M26 · lesson 26 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Enterprise Risk Frameworks for AI
📖
now learning

Enterprise Risk Frameworks for AI

15 min

Declan Hargrove is the Chief Risk Officer at a federal department with 22,000 employees, a $14 billion annual budget, and an AI portfolio that had grown to 31 active systems in less than two years. His office had always maintained the department's enterprise risk register, the structured inventory of significant risks to mission, operations, finance, legal compliance, and reputation that senior leadership and the department Secretary reviewed quarterly. For the first 18 months of AI deployment, AI risks lived in individual program offices' registers rather than the enterprise one, and each program office assessed its own system in isolation. Nobody was looking at the portfolio in aggregate. When the department's senior risk committee asked Declan to produce the first enterprise-level AI risk report, he found something his program colleagues had not seen: three separate AI systems, each assessed as medium risk on its own, were all making recommendations that fed into the same downstream human decision process. If all three were wrong in the same direction at the same time, a correlated failure mode, the combined effect on that decision was not medium risk. The aggregate was not the sum of the individual assessments. It was something different.

Why AI Risk Requires Enterprise-Level Thinking

Traditional enterprise risk management, the structured practice of identifying, assessing, and managing risk across an organization at the leadership level, was designed for risks that are visible and discrete. A major contract failure, a finding of regulatory non-compliance, a cybersecurity breach: each announces itself, each has an owner, and each can be described in a register entry that a senior committee understands on first reading. AI risk has characteristics that make it harder to manage with those same tools, and the difficulty is not that AI is technically exotic. It is that the failure modes do not arrive in the shape the register was built to hold.

AI risks can stay invisible until they become operational. A bias in a benefits eligibility model may produce systematically incorrect recommendations for months before the pattern becomes apparent in outcome data. A model whose accuracy is degrading because its training data no longer reflects current conditions will keep producing confident-sounding outputs while becoming progressively less reliable, and the confidence of the output is exactly what stops anyone from questioning it. Risk identification processes that depend on a human noticing an operational problem will not catch either failure until it is large enough to notice, which in a high-volume program means large enough to have affected a great many people.

AI risks can also be correlated across systems in ways that no individual system's risk assessment reveals. Declan's three-system finding is the classic case: each system's own risk is manageable, but the combination creates an amplified exposure that none of the three assessments captured, because none of them was looking outside its own boundary. Correlation arises when multiple systems share training data, share a foundation model, make recommendations that feed the same downstream process, or operate in the same shifting environment. Enterprise-level aggregation is the mechanism that reveals these correlations, and a program office working alone has no vantage point from which to see them.

Think of enterprise AI risk like a flight crew's combined instrument readings. Any single instrument can be managed against its own alert threshold, and each one on its own may read within tolerance. But a particular combination, low altitude with low airspeed and a high bank angle at the same moment, describes a situation that no single reading conveys. The crew needs to see the panel at once to understand what is actually happening. Enterprise AI risk management is the instrument panel for the AI portfolio, and the point of building it is that no individual gauge was ever going to tell you.

Integrating AI Into the Enterprise Risk Register

The first practical step is putting AI systems and their associated risks into the enterprise risk register itself, the same register that carries the department's major operational, financial, legal, and reputational risks. This sounds administrative and is not. A risk that lives only in a program office register is a risk that senior leadership cannot weigh against anything else, cannot fund against anything else, and will not hear about until it has already produced an incident. Placement determines audience, and audience determines whether the risk is ever traded off against the other calls on the department's attention.

Each AI system's register entry should carry the system name and description, the program it serves, the decisions or determinations it assists with, the population it affects, the risk classification the department has assigned it, the date of the last governance review, and the current status of each major risk dimension: performance, bias, security, and compliance. The classification and the review date are the two fields most often left stale, and they are the two that a senior committee will rely on most heavily, because they are the fields that let a non-specialist reader tell an urgent entry from a routine one.

The entry should also capture the system's dependency relationships: what data sources it relies on, what other systems feed it or consume its outputs, and what human decision process sits at the end of the chain. These dependency records are not documentation for its own sake. They are the direct inputs to aggregation analysis, and an entry that omits them makes the system invisible to the exact analysis that found Declan's correlated exposure. Wherever a dependency is unknown, record it as unknown rather than leaving the field blank, because a blank field reads as an absence of dependencies rather than an absence of knowledge.

Risk Aggregation Across the AI Portfolio

Aggregation is the analytic step that turns a list of register entries into a portfolio view. It starts from the dependency records and asks a small number of questions of the whole set: which systems draw on the same data source, which are built on the same underlying model, which feed the same downstream decision, and which serve the same population. Each shared element is a channel through which an error in one place can appear in several places at once, which is what makes several individually tolerable risks intolerable in combination.

The output of aggregation is not a single portfolio risk score. A single number would hide precisely the structure that makes the analysis worth doing, and it would invite a committee to track the number rather than the exposure underneath it. The useful output is a short list of identified correlation clusters, each described in terms of the shared element, the systems involved, the decision process affected, and what a simultaneous failure across the cluster would mean for the people on the other side of that decision. Declan's three systems were one such cluster, and describing them as a cluster was what made the exposure legible to leadership.

Aggregation should be repeated on a schedule rather than performed once, because the portfolio changes underneath it. A new system added to a program office register creates new shared dependencies the moment it goes live, and a model swap inside an existing system can create a shared dependency where none existed before without any register entry visibly changing. The cadence matters more than the sophistication of the method. A simple analysis run every quarter will find more real correlation than an elaborate one run once at the start of a fiscal year and never repeated.

Setting Risk Appetite for AI

Risk appetite is the level of risk an organization is willing to accept in pursuit of its objectives, and for AI it must be defined before systems are deployed rather than negotiated case by case afterward. The reason is not tidiness. Without a defined appetite, each program office makes its own tolerance decision under its own delivery pressure, every decision looks defensible in isolation, and the cumulative risk of the portfolio drifts upward without anyone having explicitly authorized the level it reaches. Nobody chose the resulting exposure. It simply accumulated, one reasonable local judgment at a time.

Appetite for a government department should be defined at several levels at once. At the portfolio level: what is the maximum acceptable number of high-risk AI systems in active production at any time? At the system level: what is the maximum acceptable error rate for a system making recommendations that affect individual benefit eligibility, and what is the maximum acceptable disparity in error rates across protected class categories? At the incident level: what constitutes a reportable AI incident, and within what timeframe must it reach department leadership and external oversight bodies? The incident question is the one most often deferred and the one most costly to answer under pressure.

Appetite statements have to be concrete and measurable rather than aspirational. "We will not tolerate bias" is not an appetite statement. "No AI system in production will have a demographic disparity in error rates that exceeds 0.10 across any protected class category defined in the department's Equity Action Plan" is one. The difference is that the first cannot be monitored or enforced and the second can. That threshold is the department's own choice, set against its own equity commitments, and not an external standard; the discipline being taught is that a department must choose a number it can measure, not that any particular number is correct.

Board-Level and Senior Leadership Risk Reporting

In federal agencies, board-level reporting has no board. Its analogs are the Secretary's senior risk committee, the chief management officer, and, for programs subject to sustained congressional interest, the relevant authorizing and appropriations committees. These audiences differ in a way that matters: none of them will read a technical risk artifact, all of them can act on what they do read, and each of them will be asked about the program in a setting where a wrong answer is expensive. What they need is AI risk information that is accurate, aggregated, and actionable, in that order.

Effective senior-level reporting has three components. The first is an executive summary that translates technical indicators into mission terms. Not "bias score 0.14 in category 3" but this: the benefits processing system is producing denial recommendations at a rate 14 percentage points higher for applicants in rural counties than in urban counties, the same county type where the program has documented access equity issues in the past, and we are investigating whether the disparity reflects a real eligibility difference or a model performance gap. The translated version is the same measurement. It is the only version a committee can act on.

The second component is a trend line showing whether the aggregate risk profile of the portfolio is improving, stable, or worsening over the reporting period. A profile worsening across multiple systems at once is a signal of something systemic: a shared data quality problem, a governance staffing gap, a change in the operational environment that every system is exposed to. That signal is invisible in any single point-in-time score, and it is the earliest warning the enterprise level is capable of producing. A stable profile deserves the same scrutiny when the portfolio grew during the period, because stability across a larger portfolio is a change in the underlying rate.

The third component is a prioritized list of the risks that require a leadership decision or a resource allocation to address, each stated as a request. Reporting that presents problems without asking for anything is a missed opportunity dressed as diligence: the committee reads it, notes it, and moves on, and the risk manager records that leadership was informed. Leadership needs to know what action is being requested of them, by when, and what happens if the answer is no. That last part is what converts a briefing into a decision.

What Enterprise Risk Management Is Not For

Enterprise AI risk management is not about eliminating AI risk, and a program that measures itself by how few risks appear on the register has misunderstood its own function. The purpose is to ensure that the risks the organization is accepting are the ones it has deliberately chosen to accept, at levels it has explicitly authorized, with monitoring in place to know when those levels are being exceeded. A portfolio with fifteen well-characterized risks under active watch is in better condition than one with three, where the difference is that the second department has not looked.

It follows that the register and the review calendar are instruments, not achievements. A completed governance review tells you that someone examined the system on a particular date against a particular set of questions. It does not tell you the system is safe now, and it does not transfer responsibility for the system's behavior to the reviewer. The same caution applies to a risk classification: a system classified as standard is a system nobody has yet found a rights or safety impact in, which is a statement about the search rather than about the system. Treating either artifact as a conclusion is how a department ends up surprised by a risk that its own paperwork said was managed.

The enterprise function also has a boundary worth naming. It is not the operational control layer, and it does not replace the testing, monitoring, and human oversight that individual systems require. Those obligations sit with the program office and are covered in the department's own risk management practice; the enterprise layer exists to see what no program office can see from where it stands. A department that pulls operational controls upward into the risk office usually ends up with slower controls and a risk office that is too busy to aggregate anything.

Anti-Patterns

  • Treating a completed governance review as proof the system is currently safe. A review records that someone looked, on one date, at the questions on one form. Between reviews the data shifts, the population changes, and the model may be updated. Report the review date next to the risk status so a reader can see how old the evidence is, and treat a long gap as a risk indicator in itself.
  • Rolling the portfolio up into a single enterprise AI risk score. One number briefs easily and hides the correlation structure that makes aggregation worth doing. Committees start managing the number, and three medium-risk systems feeding one determination vanish into an average. Report correlation clusters by name, with the shared element and the affected decision.
  • Leaving AI risks in program registers because that is where the expertise sits. Expertise and visibility are different problems. The program office knows the system best and structurally cannot see the portfolio, so keeping the risk local guarantees nobody runs the analysis only the enterprise level can run. Put the entry in the enterprise register and keep the program office as its owner.
  • Negotiating risk appetite system by system after deployment. Each negotiation is reasonable, delivery pressure pushes all of them the same way, and the aggregate is a risk level nobody authorized. Set appetite in advance at portfolio, system, and incident levels, and make any request to exceed it an escalation with a named approver.

Practice Prompts

  • Draft a full enterprise register entry for one AI system you support: system, program, decisions assisted, population affected, risk classification, date of last governance review, and status across performance, bias, security, and compliance. Note every field you could not fill from existing documentation and who would have to be asked.
  • Map the dependencies of three AI systems: data sources, shared models, upstream and downstream systems, and the human decision each connects to. Identify any element appearing in more than one map, and describe what a simultaneous failure across those systems would mean for the people affected.
  • Write three risk appetite statements, one each at portfolio, system, and incident level. Test each against one question: could someone with access to your monitoring data determine today whether it is being met? Rewrite any that fails.
  • Take a recent technical risk finding and rewrite it twice, once for the program manager and once for a senior committee with no technical background. The second version must state the mission consequence, the affected population, and the decision or resource you are requesting.
  • Draft the portfolio trend section your senior committee would receive next quarter, with the evidence behind the trend direction. If you have only one period of data, write down what you would need to collect now so the next report can show a trend.

Reflection

Ask Declan's question of the systems your organization runs today: if two or three of them were wrong in the same direction at the same time, whose decision would be affected, and would anyone notice? Then ask where that question is supposed to be answered in your governance structure, and who has the standing and the information to answer it. If no forum owns the question, you have found the gap enterprise AI risk management exists to close, and naming it is the first useful step.

Glossary

  • Enterprise risk management (ERM). The structured practice of identifying, assessing, and managing risk across an organization at leadership level rather than inside individual programs.
  • Enterprise risk register. The central inventory of significant risks to mission, operations, finance, legal compliance, and reputation that senior leadership reviews on a set cadence. A system absent from it is invisible to portfolio analysis and to the resourcing decisions that follow.
  • Risk appetite. The level of risk an organization is willing to accept in pursuit of its objectives, expressed for AI at portfolio, system, and incident level. Usable only if someone can tell from monitoring data whether it is currently being met.
  • Risk aggregation. Analysis across the portfolio that finds where individually acceptable risks combine into an unacceptable one, through a shared data source, model, downstream decision, or operating environment. Its input is the dependency information in each register entry.
  • Correlated failure mode. Several systems failing in the same direction at the same time because of something they have in common, so the combined effect exceeds what any individual assessment predicted.

Closing

Declan's finding did not come from a better evaluation method or a new testing tool. It came from asking a question nobody in the department was positioned to ask: what do these systems have in common, and what happens if they are all wrong at once? Every part of the enterprise function follows from making that question routine. The register makes systems visible, dependency records make shared elements visible, aggregation turns those into named clusters, appetite decides in advance what the department will tolerate, and reporting puts the result where a decision can be made. None of it removes risk. It means that when something goes wrong, the department can say what it knew and who authorized the exposure.

Key Takeaways

  • AI risks can be correlated across systems in ways individual assessments do not reveal. Multiple medium-risk systems feeding one decision process can create a high aggregate exposure that only enterprise-level aggregation will detect.
  • AI systems must appear in the enterprise risk register alongside other major organizational risks. Keeping AI risk isolated in program registers prevents leadership from seeing the portfolio picture or weighing it against other calls on departmental resources.
  • Dependency mapping is the input to aggregation analysis. Recording which systems share data sources, models, or downstream decisions is what lets correlated risk be found before it becomes correlated failure. Record unknown dependencies as unknown, never as blank.
  • Risk appetite for AI must be defined before deployment, not negotiated after. Measurable statements at portfolio, system, and incident level prevent the upward drift that results when each program office sets its own tolerance under its own delivery pressure.
  • Senior leadership reporting must translate technical indicators into mission terms. A disparity metric stated statistically will not prompt a decision; the same metric stated as a program equity consequence will.
  • Trend lines matter as much as point-in-time scores. A profile worsening across several systems at once signals a systemic cause, and a flat profile across a growing portfolio is itself a change worth explaining.
  • Risk reporting should carry a specific request. Leadership needs to know what it is being asked to decide, by when, and what follows from declining.
  • A completed review is evidence, not a guarantee. Reviews and classifications describe what was examined and when. Neither establishes that a system is safe today, and neither moves responsibility away from the program that operates it.

Frequently Asked Questions

Does adding AI systems to the enterprise register mean the risk office now owns them? No. Ownership stays with the program office that operates the system. The register determines who can see the risk, not who is accountable for managing it. Keeping those questions separate is what lets you raise an entry to enterprise level without diluting program accountability.

How many systems does a portfolio need before aggregation is worth doing? The trigger is a shared element, not a count. Two systems drawing on the same data source and feeding the same determination are already a cluster; two with nothing in common may never produce one. Start when the first shared dependency appears in your register.

What if we cannot agree on numeric thresholds for appetite? Write the statements you can agree on and record the rest as open items with a named decision owner and a date. A partial appetite that is genuinely enforced beats a complete one nobody signed. What does not work is deferring the exercise until deployment, when every threshold discussion carries a live system behind it.

Our committee already receives a cybersecurity risk report. Should AI risk fold into it? Security is one dimension of AI risk and rarely the largest. Performance degradation, bias, and compliance failures do not surface in a security report and are not assessed by security controls. Whether the two arrive as one document or two matters less than making sure the non-security dimensions get their own indicators and trend line.

How do we report a risk we cannot yet quantify? Report it as identified and unquantified, with what you do know: systems involved, decision affected, population exposed, and what you would need to measure it. Withholding a known risk until it has a number is how significant exposures stay off the register for a year. The request attached to such an entry is usually for the measurement capability itself.