←
AI for Government
Proficient · M7 · lesson 7 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Portfolio Management
📖
now learning

AI Portfolio Management

15 min

Diane Okafor took the job of Chief Data Officer at a county government with one inherited surprise: 23 separate AI projects, scattered across nine departments, with a combined annual spend of $6.1 million and no single person who could say which ones worked. Some were thriving. Some had been "almost done" for two years. Two were paying cloud bills for models nobody had used since the pilot ended. When the county executive asked Diane for a one page answer to "is our AI money well spent?", she realized she did not have a portfolio. She had a pile.

Managing a pile means firefighting the loudest project. Managing a portfolio means seeing all your AI investments together, comparing them on the same terms, moving money toward what works, and deliberately ending what does not. That shift, from project firefighter to portfolio manager, is the whole job at this level.

From a pile to a portfolio

A financial advisor does not fall in love with one stock. They hold a mix, rebalance as conditions change, and sell losers without shame. AI portfolio management borrows exactly that discipline. You hold a set of AI initiatives, you score them on the same scale, you fund the strong ones more and the weak ones less, and you retire the ones that no longer earn their keep. The skill that is hardest for government leaders is not starting projects. It is stopping them.

The federal framing is worth stealing even if you work for a county. An AI portfolio is not a list of projects. It is the set of AI enabled commitments an agency is making to its mission, its residents, its workforce and its budget. Without portfolio management, an organization accumulates disconnected pilots, vendor dependencies and unmanaged risks that nobody has ever seen assembled in one place. Assessments cited by federal acquisition and budget bodies suggest that up to eighty percent of federal AI pilots never reach sustained production, and the absence of portfolio governance is named as one root cause. Another is the failure to align individual initiatives to strategic plans, budget cycles and statutory obligations.

This matters more in government than in a company because your return is not only dollars. The Government Accountability Office framework for overseeing federal AI asks agencies to track governance, data, performance and monitoring, not cost savings alone. So your portfolio scoring has to value mission impact and public trust alongside money, or you will quietly optimize for the cheapest projects rather than the ones that serve people best. Agencies that build a portfolio with documented governance, monitoring and the authority to pause a system are in a position to lead. Agencies that operate a scatter of disconnected pilots are in a position to produce the next fraud detection scandal.

The inventory is the source of truth

Portfolio governance starts with an inventory, and in federal practice that inventory already exists as an obligation. Executive Order 14110, issued in 2023, directed agencies to maintain an AI use case inventory, designate Chief AI Officers and integrate AI governance with existing governance bodies. OMB Memorandum M-24-10, issued in 2024, establishes minimum practices for rights impacting and safety impacting AI, including impact assessments, pre deployment testing, ongoing monitoring, human oversight, public notice, decommissioning procedures and workforce consultation. The inventory is not a compliance artifact sitting beside the portfolio. It is the portfolio, in its most auditable form.

Each entry should carry enough to make comparison possible: purpose, current stage, whether the use case is rights impacting or safety impacting, data sources, vendor identification, cloud authorization status, performance metrics, incident history, workforce and equity considerations, and dependencies on other systems or agencies. Diane's county has no federal inventory obligation, and she built the same table anyway, because the first thing she discovered was that two departments had bought overlapping capabilities from different vendors without either knowing about the other. An inventory is how that stops being a surprise.

Stage gates from ideation to retirement

A portfolio moves initiatives through stages rather than holding them in a permanent present tense. The standard gates run ideation, pilot, scale, sustain and retirement, with explicit go or no go criteria at each boundary. The value of a gate is not the ceremony; it is that somebody has to say the words "we are proceeding" and own them, on a date, against criteria that were written before the results came in. A project that has never passed a gate and never failed one is not being managed, it is being tolerated.

Retirement is the gate everybody forgets to build. Outdated AI systems have to be decommissioned deliberately, with documented reasons, a data disposition decision and a successor plan for the work the system was doing. In federal practice, decommissioning procedures are among the minimum practices M-24-10 sets out, which means retirement is not an administrative afterthought but part of the governance record. It is also where Diane found real money, because a system nobody has retired is a system somebody is still paying for.

The portfolio scorecard: scoring Diane's 23 projects

Diane's first move is to put every initiative on one scorecard with the same five dimensions, each scored 1 to 5. This is what lets her compare a fraud detection model against a permitting chatbot without comparing apples to bridges. The dimensions are hers, chosen for a county and recorded before any project was scored, which is the part that makes the results defensible when a department head disputes them.

DimensionQuestionScore of 1Score of 5
Mission impactDoes it measurably improve a public outcome?No measured outcomeClear, sustained outcome gain
AdoptionAre people actually using it?ShelfwareEmbedded in daily work
Cost efficiencyIs value worth the run cost?Costs exceed valueStrong value per dollar
Risk and complianceGoverned, fair, secure, lawful?Unmanaged riskFully governed and documented
Strategic fitDoes it serve a county priority?Orphan projectCore to the strategy

Five dimensions scored 1 to 5 give a maximum of 25 points, and that maximum is the number every later decision rule refers back to. When Diane scores the 23 projects, a pattern jumps out. Eight score high across the board and are her producers. Six are promising but stuck, with high strategic fit and low adoption, usually because of a missing integration or absent training. Five are zombies, low on everything and still drawing budget. Four are surprises: quietly excellent projects in small departments that nobody was funding properly. Those four categories account for all 23 projects, which is the first useful thing the scorecard produced.

The federal balanced scorecard, and why weights are dangerous

Federal portfolio practice uses a wider balanced scorecard than Diane's five dimensions, and the extra ones are worth adding as your portfolio matures. Mission value asks whether the initiative advances the mission in measurable ways. Risk posture asks whether its risks are manageable within current controls. Equity asks whether it serves or harms specific populations, an idea drawn from the algorithmic discrimination protections principle in the Blueprint for an AI Bill of Rights, which is a non binding policy document rather than a statute. Workforce asks whether the initiative respects worker voice and collective bargaining obligations. Financial sustainability asks whether recurring costs are covered beyond the initial funding. Technical maturity asks whether the infrastructure, skills and authorization boundaries exist to deliver it.

The rule attached to that scorecard is strict and worth repeating exactly: no initiative should scale without acceptable scores across all dimensions. A high mission value score does not buy forgiveness on risk posture, and a strong technical maturity score does not compensate for a project that has never been costed past its pilot funding.

Two arithmetic warnings apply the moment you start combining dimensions into a single number. First, if you add dimensions to a scorecard, the maximum changes, and every threshold expressed against the old maximum becomes meaningless until you restate it. Diane's five dimensions cap at 25; a six dimension version does not, and a rule written against 25 cannot be carried across unchanged. Second, if you weight dimensions rather than treating them equally, write the weights down, confirm they sum to whatever total you claim, and publish that total with the scores. Weighted scoring is where portfolio arithmetic most often quietly breaks, usually because a weight was changed for one project and not for the others.

Rebalancing: moving money toward what works

A scorecard is useless if it does not move money. Diane sorts every project into one of four actions, and ties a budget decision to each. This is the part that requires nerve.

  • Grow, for the producers. Increase funding and remove blockers. Diane shifts $900,000 toward the grow set, which is the eight high scorers plus the four underfunded surprises.
  • Fix, for the stuck. A 90 day improvement plan naming one specific blocker to clear. If a permitting tool has 9 percent adoption because it does not connect to the records system, the fix is the integration, not more features.
  • Hold, for the genuinely unclear. Keep funding flat, gather data, decide next cycle, no new money until the picture improves. In Diane's first review the other three actions accounted for all 23 projects, so nothing landed here; the category still belongs in the framework, because a portfolio without a hold option pressures reviewers into premature verdicts.
  • Sunset, for the zombies. End them on a schedule. The five zombies were costing $740,000 a year combined, which pays for most of the $900,000 the grow decisions required.

The courage to sunset a failing project is what pays for the success of a thriving one. That is not a motto, it is the cash flow: the sunset money and the grow money are the same money, and a portfolio manager who will not do the second half of that sentence cannot do the first half either.

Sunsetting without blowing things up

In government, killing a project can feel politically radioactive. Someone championed it, a vendor depends on it, a press release once celebrated it. So you sunset by criteria, not by mood. Define the trigger in advance, in writing, before you know which project it will catch. Diane's rule is that any initiative scoring under 10 of the available 25 points for two consecutive review cycles, with no improvement plan in place, is retired. That threshold is a convention her governance board chose and recorded, not a measurement of anything, and its whole value comes from having been written down before the scores existed. Because the rule was set ahead of time, retiring a project is policy rather than an insult to its champion.

A clean sunset has steps. Notify stakeholders early. Archive the data and the model documentation for the record, with an explicit decision about data disposition rather than a default. Decommission the cloud resources so the bills actually stop. Capture the lessons learned in writing while the team still exists. Plan for the work the system was doing, because retirement without a successor plan just moves the problem. Then reassign the freed budget on the record. The decommission step alone surprises leaders regularly: Diane found one retired model still costing $4,100 a month in idle accelerator capacity nine months after anyone last used it.

Capital planning, procurement and vendor concentration

AI investments live inside the budget machinery that already exists, and portfolio managers who ignore it find their decisions overturned by a calendar. Federal budget formulation runs under OMB Circular A-11, information resource management under OMB Circular A-130, and larger investments run through the Capital Planning and Investment Control cycle with Major IT Business Case requirements. AI components have to be accounted for inside those artifacts rather than tracked in a parallel spreadsheet, and AI governance milestones belong on the same investment timeline as everything else. A portfolio decision that is not reflected in the budget submission is a preference, not a decision.

Vendor considerations are substantial and they are portfolio level rather than project level. Cloud services and AI interfaces handling federal data require appropriate authorization, and that status belongs in the inventory entry. Concentration risk means leaning so heavily on a single supplier that an outage, an acquisition or a pricing change disrupts several initiatives at once, which is a risk no individual project owner can see from inside their own project. Governmentwide and agency specific contract vehicles exist to give you alternatives, and every vendor relationship should be documented with service levels, exit rights and audit provisions. Contract modifications run under the Federal Acquisition Regulation, and contractor obligations under federal equal opportunity enforcement apply as they would anywhere else.

Cross agency dependencies and the workforce

Portfolios rarely stop at the boundary of one organization. An agency using a shared service inherits that service's performance and security posture, which means a dependency you did not procure can still take down an initiative you own. Large departments handle this with an enterprise AI function that coordinates across components, and cross component programs such as biometric processing at the border show why explicit governance is needed when one capability spans several operating units with different privacy and civil liberties obligations. Map the dependencies as part of scoring, because an initiative whose success depends on someone else's roadmap is carrying a risk its scorecard will otherwise miss.

Workforce belongs in the portfolio for the same reason. Change management across many initiatives is a portfolio problem, not a project problem, because the same staff absorb all of them at once. In federal practice, collective bargaining obligations under the Federal Service Labor-Management Relations Statute and consultation with the unions representing affected employees are part of the deployment path, and equal employment guidance on AI workplace tools applies to systems that touch hiring, assignment or evaluation. Workforce consultation is also named among the M-24-10 minimum practices, so a portfolio that plans deployments without it has planned deployments that will stop.

Risk governance across the portfolio

Portfolio risk governance maps onto frameworks the agency is already using. The NIST AI Risk Management Framework, which is voluntary rather than binding, contributes its GOVERN function for roles, responsibilities and policies, and its MANAGE function for tracking ongoing risks and responses. The GAO framework organizes accountability into governance, data, performance and monitoring. M-24-10 adds a specific and often overlooked instrument: the authority to pause a rights impacting system and fix it when monitoring reveals a problem. A portfolio review that can only add and remove funding is missing that middle option, and pause is frequently the correct answer for a system that is valuable and currently misbehaving.

Case history is the cheapest form of experience here. A federal identity verification rollout became a portfolio lesson when the department reversed course under public pressure, which is what vendor concentration and inadequate subcontractor review look like at the moment they become news. Multi year claims modernization efforts show what sustained portfolio commitment with union coordination actually costs and delivers. Disability triage systems show the balance between caseworker workflow and statutory adjudication responsibilities that cannot be automated away. Each of those is worth mapping onto your own portfolio and asking which of your initiatives is the local version.

Three numbers for the county executive

Diane's boss does not want 23 status reports. He wants a heartbeat. Track three portfolio level health metrics and report them every quarter.

  1. Portfolio value ratio. Total measured mission value plus realized savings, divided by total run cost. Trending up means the portfolio is getting healthier rather than merely bigger.
  2. Adoption rate. The share of deployed initiatives actually in regular use. A portfolio that is mostly shelfware is a portfolio in trouble no matter how many projects it contains.
  3. Stuck ratio. The share of initiatives that have been in progress longer than 18 months. This is the early warning that pilot purgatory is creeping back.

These three turn Diane's pile into the one page answer the executive asked for. After the rebalance she can say that the portfolio value ratio rose from 1.3 to 2.1, adoption climbed from 55 to 78 percent, and the stuck ratio fell from 26 percent to 9 percent. That last one is checkable against her own numbers: six stuck projects out of 23 is where the 26 percent came from. Whenever you report a portfolio ratio, report the two numbers underneath it as well, so the reader can do the division themselves.

Reporting and a cadence that keeps it alive

A portfolio is not a one time spring cleaning. Set a quarterly review where every initiative is re scored, the four actions are reassigned, and money is rebalanced. Make it a standing meeting with finance, the affected departments and your governance lead in the room. The discipline of the calendar prevents the pile from rebuilding itself, because projects drift and a quarterly re score catches both the producer that started slipping and the zombie that crept back onto the budget.

Reporting has more than one audience and they need the same numbers presented differently. Leadership needs concise, comparable views across initiatives. Budget and oversight bodies may request portfolio summaries. Auditors conduct audits and legislatures hold hearings. A portfolio dashboard with clear status, risk and value for each initiative serves all of them, and public transparency through the use case inventory reinforces the whole structure by making the portfolio legible to people outside it. Build one dashboard and cut different views from it rather than maintaining a separate story for each audience, because the separate stories are what fall apart under questioning.

Anti-patterns

  • The pile that calls itself a portfolio. A list of projects in a spreadsheet is an inventory. Without common scoring and budget consequences attached, nothing about it is portfolio management.
  • Scoring without moving money. The most common failure. The scorecard is completed, everyone agrees with it, and every project keeps exactly the budget it had.
  • Sunset by mood. Retiring the project whose champion left rather than the project that scored lowest. Pre set criteria exist precisely to remove that discretion.
  • Stopping without decommissioning. The team disbands, the announcement goes out, and the cloud resources keep billing because nobody owned the shutdown.
  • Threshold drift. Adding a dimension to the scorecard while leaving a cutoff written against the old maximum, or changing a weight for one project and not the rest.
  • Portfolio without dependencies. Scoring each initiative as though it stood alone, so a shared service outage or a stalled upstream program takes out three projects that all scored green.
  • Skipping the workforce. Planning deployment sequencing with no consultation, then discovering the obligation at the point where it stops the deployment.
  • Concentration blindness. Every project independently choosing the same supplier, so a single pricing change or outage becomes a portfolio event nobody modeled.
  • Only two verbs. A review that can fund or cancel but cannot pause and fix, which forces a binary answer for systems whose real answer is to stop, correct and resume.
  • Ratios without their inputs. Reporting a value ratio or an adoption percentage with no numerator and denominator attached, which makes the number impossible to check and easy to drift.

Practice prompts

  1. List every AI initiative in your organization, including the ones in other departments you only half know about. Count them. Most people find the count is higher than they expected.
  2. Score a handful of them on Diane's five dimensions and write down the total out of 25. Then write down what budget decision each score implies, and notice whether you are willing to make it.
  3. Draft your sunset trigger before your next review cycle, and get it recorded by whoever governs the portfolio, so the first project it catches cannot argue the rule was written for them.
  4. For your largest initiative, list every external dependency: shared services, other agencies, single suppliers. Which of them could stop your project without asking you?
  5. Pick the initiative your organization is proudest of and find its recurring cost past the current funding line. If nobody can answer, financial sustainability is your lowest score.
  6. Identify one portfolio improvement you can implement within 60 days, name the three biggest risks in the portfolio as you see it today, and put both in writing where your governance body will see them.

Reflection

Diane's hardest conversation was not about the zombies. It was about a project with a real champion, real effort behind it, and a score that would not move. Portfolio management works by making that conversation routine and unpersonal, which is only possible if the rules were agreed before anyone knew who they would catch. Ask yourself which initiative in your organization would be retired today if your scoring rule already existed, and then ask why it has not been. If the honest answer involves a person rather than a number, you have found the reason the pile keeps rebuilding itself, and the fix is a rule written in advance rather than a braver conversation later.

Glossary

  • AI portfolio. The full set of AI enabled commitments an organization has made to its mission, residents, workforce and budget, managed as one object rather than as separate procurements.
  • Stage gate. A decision point between lifecycle stages where someone must explicitly approve continuation against criteria written in advance.
  • Sunset trigger. A pre recorded rule that identifies which initiatives are retired, so retirement is policy rather than judgment applied after the fact.
  • Decommissioning. The controlled shutdown of a retired system, including data disposition, documentation archiving, resource shutdown and successor planning.
  • Concentration risk. Exposure created when several initiatives depend on the same supplier, so one outage, acquisition or pricing change affects all of them.
  • Capital planning and investment control. The federal cycle that selects, controls and evaluates major investments, and the cycle AI initiatives must fit inside.
  • Pause and fix. The authority to suspend a rights impacting system and remediate it, rather than choosing only between continuing and canceling.
  • Portfolio value ratio. Measured mission value plus realized savings divided by total run cost, reported with both inputs visible.
  • Stuck ratio. The share of initiatives that have been in progress longer than a stated duration, used as the early warning for pilot purgatory.

Closing

The county executive's question was not really about money. It was about whether anyone was in charge of the whole thing. Diane could not answer it while she had 23 projects and no shared basis for comparing them, and she could answer it the moment she did, even though several of the answers were unflattering. That is the trade a portfolio makes: you give up the comfort of judging each project on its own terms, and you get the ability to say what the whole investment is doing and what you are going to change about it. The pile always feels safer, right up to the moment somebody asks.

Key takeaways

  • Manage the set, not the loudest project. A portfolio view lets you compare every AI initiative on the same terms and move money deliberately.
  • Score on one scorecard. Mission impact, adoption, cost efficiency, risk and compliance, and strategic fit, scored 1 to 5 for a maximum of 25, make a chatbot and a fraud model genuinely comparable.
  • Restate thresholds when the scorecard changes. Adding a dimension changes the maximum, and weighted scores must sum to a stated total that is published with the scores.
  • The inventory is the portfolio. Purpose, stage, impact designation, vendor, authorization, metrics, incidents and dependencies in one place is what makes comparison possible.
  • Sort into grow, fix, hold and sunset. Every initiative gets one action with a budget decision attached, or the scorecard changes nothing.
  • Sunset by pre set criteria. A retirement rule recorded in advance turns ending a weak project from a political insult into routine policy.
  • Always decommission, never merely stop. Idle models keep billing, and freeing that budget is often exactly what funds your strongest initiatives.
  • Keep pause and fix on the table. A review that can only fund or cancel gives the wrong answer for a valuable system that is currently misbehaving.
  • Report three health numbers with their inputs. Portfolio value ratio, adoption rate and stuck ratio give leadership a heartbeat instead of 23 status reports, and every ratio should show the numbers underneath it.
  • Re score quarterly. A standing review with finance and the departments keeps the portfolio honest and stops the pile from rebuilding.

Frequently Asked Questions

How many dimensions should our scorecard have?

Enough to prevent a single strength from hiding a fatal weakness, and few enough that people will actually score them. Diane's five work for a county. Federal practice adds equity, workforce, financial sustainability and technical maturity as separate dimensions, which is the right direction as a portfolio grows. Whatever you choose, fix the set before scoring anything, and if you change it later, restate every threshold that was written against the old maximum.

Should we weight the dimensions?

Only if you are prepared to write the weights down, apply them identically to every initiative, and publish the total alongside the scores. Unweighted scoring is transparent and hard to manipulate. Weighted scoring is more expressive and is where the arithmetic most often breaks, usually because someone adjusted a weight for one project mid cycle. If you cannot state your weights and confirm what they sum to, you do not have a weighted model, you have a preference with decimals attached.

What if a low scoring project has a powerful champion?

That is the situation the pre recorded sunset trigger exists for. The rule has to be agreed before anyone knows which project it will catch, ideally by the body that governs the portfolio rather than by you alone. Then the conversation is about whether the rule was applied correctly, which is a factual argument, instead of about whether you respect someone's work, which is not an argument you can win.

Is a pause the same as a failure?

No, and treating it that way is what makes people avoid it. A system that is delivering value and has developed a monitoring problem should stop, be corrected and resume, which federal guidance for rights impacting systems addresses directly as a pause and fix expectation. A portfolio review that offers only continue or cancel pushes teams toward continuing, because cancellation is too expensive an admission, and continuing is exactly the wrong answer for a system that is currently getting decisions wrong.

How do we handle a project that spans several departments or agencies?

Score it once, but record every dependency explicitly and name a single accountable owner. Shared services and cross component programs inherit performance and security posture from parties you do not control, and that inherited risk belongs on the scorecard rather than in a footnote. If no single person can be named as accountable across the boundary, that is your finding, and it usually needs governance attention before the project needs more funding.

How often should the portfolio be reviewed?

Quarterly is the cadence that catches drift without exhausting the participants, with the review aligned to the budget calendar so its decisions can actually be executed rather than admired. The standing attendance matters as much as the frequency: finance, the affected departments and the governance lead in one room. A review whose conclusions arrive after the budget submission has closed is a report, not a review.