←
AI for Government
Proficient · M13 · lesson 13 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

AI Use Case Inventory Management

15 min

Marcus Bell, deputy CIO at a mid-sized state Department of Human Services, thought he knew every AI system his agency ran. Then a reporter filed a public-records request asking for the agency's complete list of automated decision tools. Marcus pulled together what he had: three systems he could name from memory. Two weeks later, a junior caseworker mentioned the chatbot the call center had been quietly piloting, the document-summarizer the eligibility unit bought on a credit card, and the fraud-scoring model a contractor had installed eighteen months earlier. The reporter's request had asked for a list. Marcus did not have one. He had a guess.

That gap between what leadership thinks is running and what is actually running is the single most common failure in government AI governance. You cannot govern, audit, defend, or improve a system you do not know exists. An AI use case inventory closes that gap. It is a living register of every place your agency uses artificial intelligence, who owns it, what it touches, and how risky it is. This lesson shows you how to build one, keep it current, and turn it from a compliance chore into a management tool.

Why an Inventory Is the Foundation of Everything Else

Every serious AI governance framework starts with the same step: know what you have. The federal Office of Management and Budget memorandum on agency AI use, known as M-24-10, directs federal agencies to publish an annual inventory of AI use cases and to flag which ones are "rights-impacting" or "safety-impacting." The National Institute of Standards and Technology AI Risk Management Framework, a voluntary standard many state and local governments adopt, builds its entire "Map" function around identifying and contextualizing each system. The Government Accountability Office accountability framework asks auditors to start by confirming the agency can produce a complete list.

Notice the pattern. Inventory is not one task among many. It is the spine. Your governance board cannot prioritize reviews without a list. Your procurement team cannot prevent duplicate purchases without a list. Your privacy officer cannot assess data exposure without a list. When Marcus could only name three systems, he had not failed at AI policy. He had failed at AI bookkeeping, and everything downstream collapsed because of it.

A good inventory earns its keep by answering questions that are otherwise unanswerable. What AI systems does our agency operate? What risks do they pose? Who is accountable for each one? When was each last reviewed? Are they performing as expected? And the question that keeps chief information officers awake: are there systems we do not know about? Every one of those is a routine query against a maintained register and an expensive research project without one.

Be clear about the limit, though. The inventory is a precondition for governance, not governance itself. A complete, beautifully maintained list of systems that nobody reviews is still an ungoverned portfolio, and an agency can be fully compliant with a publication requirement while never once asking whether one of the listed systems is harming anyone. The list makes the work possible. It does not do the work.

What the Inventory Actually Feeds

The case for the register is easier to fund when you can name the four functions that break without it, because each one fails in a different and recognizable way. Marcus could point at all four inside his own agency within a month of building the list.

The governance board is the most direct consumer. Without a list, the board reviews whatever gets escalated to it, which means it reviews the systems whose owners are conscientious and never sees the ones whose owners are not. With a list sorted by risk class and review date, the board's agenda writes itself and the systems nobody is worried about stop being systems nobody is watching. That inversion is the single biggest change a register produces.

Procurement needs the register to prevent duplicate purchases and to catch new systems at the point of entry rather than at the next annual sweep. Two program offices separately buying overlapping document-summarization tools is a routine outcome in an agency without a list, and it is not only a waste of money: it doubles the governance surface for one capability. The register also gives the contracting officer a place to check whether a proposed tool duplicates something already under contract, which is a question nobody can reliably answer from memory once the portfolio grows.

The privacy officer needs it to assess data exposure. The question "which of our systems send personal data to an external service?" is unanswerable without a data-used field maintained per entry, and it is the question that arrives urgently, usually after something has happened elsewhere. Auditors need it because a complete list is the first thing they ask for, and the speed and confidence of that answer sets the tone for everything that follows. An agency that can produce the register on request is treated differently from one that has to assemble a guess.

The Four Jobs: Track, Classify, Update, Report

A working inventory does four things. Get these right and the rest is detail.

Track: capture every use case, not just the official ones

The hard part is discovery. The fraud model the contractor installed did not announce itself. Marcus found his missing systems by combining four discovery methods, and you should run all four at least once a year:

  • Procurement scan. Pull every contract, purchase order, and software invoice over the past three years. Search for terms like "machine learning," "predictive," "automated," "model," "analytics," and the names of common vendors. Credit-card and micro-purchases are where shadow systems hide.
  • Survey the workforce. Ask every program manager a plain-language question: "Does any tool you use make a recommendation, prediction, score, or decision automatically?" Avoid the word "AI" because staff often do not realize a vendor feature counts.
  • Data-flow review. Work with IT to see which systems call external AI services or feed data to a model. Network logs reveal tools nobody declared.
  • Vendor disclosure. Ask existing vendors directly whether their product uses AI, including features added through updates. A document system that was rules-based last year may have a generative summarizer today.

Run all four rather than picking the one that fits your access, because each method is blind exactly where another one sees. The procurement scan finds anything that was paid for and misses anything that arrived as a free feature or inside a contractor's deliverable. The workforce survey finds the tools people actually use daily and misses the batch process running unattended in a back office where nobody would think to mention it. The data-flow review finds systems by their network behavior, catching tools nobody declared, and misses anything running entirely inside a vendor's cloud where your logs show only one outbound connection. Vendor disclosure finds capabilities added by update, which is the only route none of the other three can see at all, because nothing was purchased, nobody changed their workflow, and no new integration appeared.

Expect the yield to shift year to year. A method that surfaces nothing in one cycle has not been disproved; it has told you that this year's new systems arrived by other routes. Drop a method and you will find out which one it was the hard way, usually when a records request asks for a list.

Classify: sort by risk, not by technology

Once you have the list, classify each entry by how much it could harm a person, not by how sophisticated the math is. A simple model that decides who gets food assistance matters more than a fancy one that routes internal email. Borrow the federal distinction:

  • Rights-impacting: the system affects a person's legal rights, benefits, access, or opportunities. Eligibility scoring, fraud flags, hiring screens, and parole recommendations all qualify.
  • Safety-impacting: the system could affect physical safety or critical infrastructure. Think traffic signal optimization or emergency dispatch triage.
  • Operational: internal efficiency tools with no direct effect on the public, like meeting transcription or code completion.

Underneath that headline distinction, a consistent classification scheme lets you sort and report. A four-level risk scale works for most agencies, and the value of writing the definitions down is that two reviewers looking at the same system land on the same level.

Risk levelWhat it means
LowLimited impact if it fails, with no rights or safety implications
ModerateSome impact if it fails, affecting the user experience
HighSignificant impact if it fails, affecting important decisions
CriticalSevere impact if it fails, affecting rights or safety

Record the level as a field rather than a note, so the register can be sorted by it, and record the key risks you identified alongside it. The level tells the board how much attention an entry deserves; the named risks tell the reviewer what to actually look at when the entry reaches the top of the queue. An entry rated critical with no risks written down sends a reviewer in with nothing but the rating, which is how a review turns into a conversation about whether the rating was right.

Three further dimensions are worth capturing because they answer different questions. System type records what the model does: classification, prediction and regression, clustering, generation, ranking, or recommendation. That field tells your technical reviewers which failure modes to look for. Mission area records where it sits, across benefits, human resources and recruitment, law enforcement, operations, and public-facing service. That field tells your civil rights office which entries deserve their attention first. Development status records how far along it is: development, testing, pilot, production, or retired. That field stops the register from quietly filling up with tools nobody has used in two years.

Update: make it a living register, not a snapshot

An inventory taken once and filed away is worse than none, because it creates false confidence. Tie updates to events that already happen: every new contract signing, every system change request, every vendor renewal triggers an inventory check. Assign a named owner per entry who attests once a quarter that the record is still accurate.

Three update rhythms cover most of what changes. Quarterly, the named owner verifies the system still exists and the record is still accurate, updates status and performance metrics, and notes what has changed since the last review. Annually, run a comprehensive review of every system, update risk classifications where warranted, confirm that safeguards were verified rather than merely documented, and sweep for new systems that entered by routes your triggers missed. In real time, log the things that cannot wait: major incidents reported immediately, system shutdowns noted when they happen, and significant changes documented as they are made rather than reconstructed later.

Report: produce the annual public account

Federal agencies publish their inventory yearly, and the trend toward transparency is reaching states and cities. Build your internal record so the public version is a simple filtered export, not a fresh project. If reporting feels like a fire drill, your tracking is broken.

Designing for the export changes how you build the register. Decide in advance which fields are publishable and which are internal only, and hold that distinction as a property of the field rather than a judgment someone makes each year under deadline. If producing the public version requires rewriting a dozen entries by hand because the plain-language name was never filled in, or redacting a free-text notes field that owners have been using for everything, those are design problems in the internal register showing up as an annual scramble. Fix them once, in the field structure, and publication becomes a filter rather than a project.

Internal reporting should be more frequent than the public account and more pointed. A quarterly report to the governance board covers the total system count; the count broken down by risk level; new systems added; systems retired; major incidents; overdue reviews; and the recommendations that follow. That last pair is what turns a report into a decision. A board that sees three critical-risk systems with reviews overdue by two quarters has been handed an agenda, not a status update.

The Inventory Record: A Field-by-Field Template

Here is a record structure you can put into a spreadsheet on Monday. Each row is one use case. Keep the fields lean enough that owners will actually fill them in.

  • ID - a stable internal identifier (e.g., HHS-2026-014)
  • Use case name - plain-language title ("Call center benefits chatbot")
  • Owning program - the office accountable for outcomes
  • Business owner - named person, with email
  • Purpose - one sentence on what decision or task it supports
  • Vendor / build - commercial product, contractor-built, or in-house
  • Data used - categories of data in and out, including personal data
  • Risk class - rights-impacting / safety-impacting / operational
  • Human role - does a person review before action, or is it automatic?
  • Status - pilot / production / retired
  • Date added & last verified - for the quarterly attestation
  • Governance review date - when the board last assessed it

That last column is where the inventory stops being a list and becomes a management dashboard. Sort by risk class, then by review date, and your governance board instantly sees which high-stakes systems are overdue for scrutiny.

As the register matures, entries grow additional groups of fields. Technical details record the data sources, the model type, the version, the date of last update, and the accuracy or performance metrics claimed for it. Governance records the last board review date, the next scheduled review, the approvals obtained, and any outstanding issues, so the review history lives with the system rather than in a separate minutes file. Monitoring records current performance, the status of safeguards, recent incidents, and audit status. Add these in the order your governance board actually asks for them. A field nobody uses is a field owners will stop filling in accurately, and an inaccurate field is worse than an absent one.

Start Simple and Graduate

There are three levels of inventory infrastructure, and agencies routinely fail by starting at the wrong one. The simple approach is a spreadsheet holding the fields above. The medium approach is a database with reporting built on top, which becomes worth the effort once sorting and cross-tabulating by hand is consuming real time. The comprehensive approach is a dedicated inventory system with API access, which lets other systems read and write to the register directly rather than relying on a person to transcribe.

Start simple and graduate as your inventory grows. The failure mode at the top end is an inventory so elaborate that people avoid using it, at which point the register drifts out of alignment with reality and you are back to Marcus with his three remembered systems, except now with a database to give the guess an air of authority. The failure mode at the bottom end is staying on a spreadsheet long past the point where nobody can answer a board question without an afternoon of filtering. Move when the pain is real, not when the tooling is fashionable.

A Worked Walkthrough: Marcus Rebuilds His List

Marcus ran the four discovery methods over six weeks. The procurement scan surfaced 11 tools; the workforce survey added 7 more that staff had not connected to "AI"; the data-flow review caught 2 hidden integrations. Final count: 20 use cases, not the 3 he started with. The vendor disclosure round turned up nothing new that cycle, which is itself worth recording, because a method that finds nothing this year is not a method to drop next year.

Classification changed his priorities completely. Of the 20, four were rights-impacting: the fraud-scoring model, the eligibility document summarizer, an automated appeals-routing tool, and a new resume-screening system in HR. Those four moved to the top of the governance board's queue. The other 16 were logged but flagged as lower urgency. Marcus did not try to review everything at once. He let risk class set the order.

The fraud model alone justified the whole effort. Once it was on the list with a named owner, the board asked a question nobody had asked in eighteen months: had anyone checked whether it flagged some neighborhoods at higher rates than others? No one had. That question, impossible before the inventory existed, is the entire point.

Notice what the exercise revealed about the shape of the gap. The systems Marcus already knew about were the ones bought through normal procurement with a named program sponsor. The ones he did not know about arrived by the routes that bypass sponsorship: a micro-purchase on a credit card, a contractor deliverable, a feature switched on inside a product the agency already owned. Design your discovery around those routes specifically, because the official channel is the one that was never going to hide anything from you.

Anti-Patterns

  • The "we'll do it once" trap. Without event-based triggers and quarterly attestation, the list rots within a year. A stale inventory is not a neutral artifact; it actively misleads the people relying on it.
  • Counting tools, not decisions. One vendor platform may run three distinct rights-impacting use cases. Inventory the use cases, not the logos, because governance attaches to the decision a system makes about a person.
  • Letting IT own it alone. IT can find systems, but only program owners know what decisions those systems drive. Make ownership shared, with the named business owner accountable for the accuracy of the record.
  • Hiding behind "it's just a pilot." Pilots make real decisions about real people. Inventory them from day one, and record the development status so the board can see how much of the portfolio is running on provisional footing.
  • Building an inventory so complex that people avoid it. Every field you add is a field an owner has to maintain quarterly. Owners under pressure do not refuse; they fill in something plausible. An elaborate register full of stale entries is harder to fix than a lean one that is accurate.
  • Duplicate entries under different names. The same system logged twice, once by the program that bought it and once by the office that inherited it, inflates counts and splits review history in half. Reconcile against stable identifiers at the annual review, not by memory.
  • An inventory disconnected from governance. A register that no board reads, that triggers no review, and that feeds no procurement check is a compliance artifact rather than a control. The review-date field exists precisely to close this loop, and it only works if someone acts on it.
  • Treating the complete list as the finish line. The inventory tells you what exists and what it touches. It cannot tell you whether any of those systems is treating people fairly, and an agency that can produce a perfect list while never reviewing a single high-risk entry has bought itself documentation rather than governance.
  • Reconciling only on paper. Discrepancies between the inventory and operational reality are themselves a governance finding. When the annual sweep turns up a production system nobody registered, the fix is not only to add the row but to work out which control let it through.

Practice Prompts

  • Run the procurement scan. Pull three years of contracts, purchase orders and software invoices for one program office. Search the terms from this lesson. Count how many results were already known to you before you started, and how many were not. The ratio is your current visibility.
  • Write the workforce survey question. Draft the plain-language question you would send to program managers, avoiding the word "AI." Test it on colleagues who do not work in technology and revise until none of them asks what you mean.
  • Classify a sample. Take a set of tools your agency runs and assign each a risk class of rights-impacting, safety-impacting or operational, a risk level from low through critical, a system type, a mission area and a development status. Note where two colleagues classify the same system differently, because that disagreement is the classification guidance you still need to write.
  • Design the update triggers. List the events already happening in your agency that should fire an inventory check: contract signings, change requests, vendor renewals, and any others specific to you. For each, name who would notice and how the trigger would reach the register.
  • Draft the quarterly report. Build the report your governance board would receive: total count, counts by risk level, additions, retirements, incidents, overdue reviews, and recommendations. Then ask what decision the board could make from it. If the answer is none, the report needs different fields.

Reflection

If a reporter filed the same request tomorrow, how long would it take your agency to produce a defensible list, and how confident would you be that it was complete? Think specifically about the routes that bypass procurement: micro-purchases, contractor deliverables, and features switched on inside products you already own. Which of those three would you find, and which would you miss?

Then ask the question Marcus's board eventually asked. Of the systems you do know about, which highest-risk one has gone longest without anyone examining how it treats the people it decides about? If the register cannot answer that in a single sort, the gap is not in your discovery. It is in the fields you chose and who reads them.

Glossary

  • AI use case inventory. A maintained register of every place an agency uses artificial intelligence, recording ownership, purpose, data, risk classification and review history.
  • Rights-impacting use. A use whose decisions affect a person's legal rights, benefits, access or opportunities, and which therefore carries heightened review and disclosure obligations.
  • Safety-impacting use. A use whose decisions could affect physical safety or critical infrastructure.
  • Shadow system. An AI tool running inside the agency without appearing in any register, typically arriving through a micro-purchase, a contractor deliverable, or a feature added to an existing product.
  • Attestation. A periodic confirmation by the named owner that an inventory record is still accurate, which converts the register from a document into a maintained control.
  • Use case. One distinct decision or task a system performs. A single vendor platform can host several, and each is inventoried separately.
  • Development status. Where a system sits on the path from development through testing, pilot and production to retired.

Closing

A comprehensive inventory is foundational to governance, and it is only foundational: everything valuable happens in what the register lets you do next. Start simple, on a spreadsheet if that is what you can maintain, and graduate to more sophisticated infrastructure as the portfolio grows. Maintain it actively through event triggers and quarterly attestation. Connect it to the governance decisions it is supposed to drive, and audit it regularly against reality, because a discrepancy between the register and what is actually running is not a bookkeeping error. It is a governance finding about how systems enter your agency unnoticed.

Marcus went from three remembered systems to a maintained register of twenty in six weeks. What changed was not his diligence. It was that the question "which of our high-risk systems has nobody looked at?" became answerable, and once it was answerable somebody finally asked it.

Key Takeaways

  • Inventory is the spine, not a side task. Every governance, procurement, privacy, and audit activity depends on a complete, current list of where your agency uses AI.
  • Discovery takes four methods. Procurement scans, plain-language workforce surveys, data-flow reviews, and direct vendor disclosure each catch systems the others miss, especially shadow tools that arrive by micro-purchase, contractor deliverable, or a feature switched on in a product you already own.
  • Classify by harm, not by sophistication. Sort entries into rights-impacting, safety-impacting, and operational, and record a risk level from low through critical, so the highest-stakes systems rise to the top of the review queue.
  • Capture type, mission area and status too. System type tells reviewers which failure modes to look for, mission area tells the civil rights office where to look first, and development status keeps retired tools from cluttering the register.
  • Make it living. Tie updates to contracts and change requests, run a quarterly owner attestation, sweep comprehensively once a year, and log incidents and shutdowns in real time rather than reconstructing them later.
  • Design reporting in from the start. If the annual public inventory is a fresh fire drill, your tracking is broken; the public version should be a filtered export, and the internal quarterly report should end in recommendations a board can act on.
  • Start simple and graduate. Spreadsheet, then database with reporting, then a dedicated system with API access. An inventory so complex that people avoid it produces confident-looking entries nobody has verified.
  • The list is a management tool. Adding a review-date field turns a static register into a dashboard that shows which high-risk systems are overdue for governance.
  • Count decisions, not logos. One platform can host several distinct rights-impacting use cases; inventory each decision a system makes about the public.
  • A complete list is not governance. The register makes review possible; it does not perform it, and it says nothing about whether any listed system is treating people fairly. Discrepancies between the register and reality are governance findings in their own right.

Frequently Asked Questions

How do we find systems nobody told us about?

Aim your discovery at the routes that bypass normal procurement, because the official channel was never hiding anything. Scan credit-card and micro-purchase records, not just contracts. Read contractor deliverables for tools installed as part of a project. Ask vendors directly whether features added in updates use AI, since a rules-based product last year may carry a generative summarizer today. And ask staff about behavior rather than technology: does any tool you use make a recommendation, prediction, score or decision on its own?

What counts as an AI system for inventory purposes?

Inventory the decision, not the technology. If a tool produces a recommendation, prediction, score or automatic decision that affects how a person is treated, it belongs on the list regardless of whether anyone involved calls it AI. Drawing the boundary by technical sophistication is how a simple scoring rule that decides who gets food assistance stays off the register while a complex tool that routes internal email goes on it.

How often should entries be reviewed?

Run three rhythms rather than one. Quarterly, the named owner attests the record is still accurate and updates status and performance. Annually, review everything comprehensively, revisit risk classifications, confirm safeguards were actually verified, and sweep for systems that entered by an unexpected route. In real time, log incidents, shutdowns and significant changes as they happen. Separately, tie a check to the events that already occur: contract signing, change request, vendor renewal.

Should pilots be in the inventory?

Yes, from day one. A pilot makes real decisions about real people, and the pilot label describes the agency's confidence rather than the resident's exposure. Record the development status so your board can see how much of the portfolio is running provisionally, which is a governance question in itself when the answer turns out to be most of it.

Our inventory is complete. Are we compliant and safe?

Those are two different questions and the list only speaks to the first. A published, accurate inventory can satisfy a reporting obligation while nobody has yet examined whether the fraud model flags some neighborhoods more than others. Compliance is the register; safety is what the governance board does with it. Sort by risk class and review date, and if the highest-risk entries have never been assessed, you have documentation rather than governance.