←
AI for Government
Capable · M8 · lesson 8 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Use Case Inventory and Documentation (OMB M-24-10)
📖
now learning

AI Use Case Inventory and Documentation (OMB M-24-10)

15 min

Daniel Okwu was three weeks into a detail supporting his agency's newly designated Chief AI Officer (CAIO) when the scope of the assignment finally landed on him. The agency had to publish its first AI use-case inventory under OMB Memorandum M-24-10, and the working list he had inherited contained exactly six entries. Daniel was nearly certain the real number was several times that. Somewhere in the agency a program office was running a chatbot it had bought on a credit card, a data team was using a machine-learning model to prioritize cases, and a contractor was delivering "analytics" that nobody had ever called AI out loud. His job was to find all of it, classify it correctly, document it so it would survive an Inspector General audit, and decide what the public would see. This lesson is the playbook Daniel built.

What OMB M-24-10 Actually Requires

OMB Memorandum M-24-10, "Advancing Governance, Innovation, and Risk Management for Agency Use of Artificial Intelligence," was issued in 2024 as implementing guidance for federal agency use of AI under Executive Order 14110. Federal AI guidance moves, and the memo that binds your agency this fiscal year may not be the one that bound it last year, so confirm the current instrument with your CAIO or counsel before you build a compliance process on any summary, including this one. What follows is the substance of the requirement set, which has proved more durable than any particular document number.

  • Designate a Chief AI Officer. A senior official accountable for the agency's AI governance, coordination and risk management, supported by governance structures with actual authority to oversee AI systems. Daniel's CAIO was that person, and the inventory was their first major deliverable.
  • Maintain and publish an AI use-case inventory. A catalog of the agency's AI use cases, covering systems in development as well as systems in production, updated annually, with a public version. This is the artifact Daniel owned.
  • Apply minimum practices to rights-impacting and safety-impacting AI. For the higher-risk categories the memo requires concrete steps before deployment, among them an AI impact assessment, real-world performance testing, ongoing monitoring and meaningful human oversight.
  • Report and disclose publicly. Much of the inventory is published so the public can see how the agency uses AI, with carve-outs for genuinely sensitive cases, and reporting follows specified standards rather than each office's preferred format.
  • Use waivers only deliberately. A minimum practice can be waived where applying it would increase risk to safety or rights overall, or create an unacceptable impediment to critical operations, but the waiver must be made by the CAIO, documented, and revisited. A waiver is a governed exception, not a loophole.

For the inventory specifically, the required content is more prescriptive than most agencies expect on first reading. Each entry documents the system name and a description of what it does; the agency component and office that operates it; its status, meaning in development, deployed or retired; the AI techniques it uses; its data sources; its intended use; the performance metrics by which it is measured; its risk classification as safety-impacting, rights-impacting or neither; the responsible official accountable for it; and its approval and monitoring status. For the safety-impacting and rights-impacting entries, the impact assessments themselves must also be documented.

Why the Inventory Is Not Paperwork

Before this requirement existed, many agencies genuinely did not know how many AI systems they were operating. Systems were scattered across departments and divisions. Different teams used different technologies for similar purposes, duplicating effort and producing inconsistent results for similar decisions. When something went wrong, there was often no central record of what existed or how it worked, which meant the first hours of any incident were spent establishing facts that should have been written down years earlier.

A maintained inventory changes what governance can do. It demonstrates compliance, which is the visible part, but the useful part is that it enables measurement and monitoring, makes it far harder for a harmful system to run under the radar, and lets one division learn from another instead of repeating its mistakes. An inventory does not by itself stop a bad system; nothing on a spreadsheet stops anything. What it does is remove the excuse of not knowing, and move an agency from reacting to problems as they surface toward checking systems against standards before they cause any.

Running the Inventory Drive: Finding the Shadow AI

The hardest part of the inventory is not the template. It is finding the systems nobody volunteers. Daniel ran the drive on three parallel tracks, because no single method finds everything, and he sequenced the whole effort through a series of steps that other agencies have found roughly reproducible.

Surveys, with the right definition

Before anything else, settle what counts. The government-wide definition is deliberately broad: a system that uses machine learning, statistical models or other computational methods to make or support decisions is in scope. Announce the effort to all components, explain what you are doing and why, and ask for voluntary reporting first, because a drive that opens with an audit posture produces defensive answers. Then send a short structured survey. Daniel learned fast that asking "do you use artificial intelligence?" returns almost nothing, because people do not think of their tool as AI. So he asked behavioral questions instead: does any tool you use make a prediction, a recommendation, a classification or a score, or generate text or images? Does anything you rely on automate a decision or analyze data for you? That phrasing surfaced eleven systems the first round had missed.

Contract review and cloud audit

A great deal of government AI arrives through vendors, so Daniel pulled active contracts and task orders and searched them for the tells: "machine learning," "predictive," "natural language," "automated decision," "model." Two contractor-delivered systems turned up here that no program office had reported, because the staff using the output did not know what was inside the deliverable. The parallel move is auditing what the agency is already paying for in the cloud. Review the organization's cloud accounts and SaaS subscriptions and identify the AI services running in them, because a managed language or vision service consumed through an existing account leaves a billing trail long before it leaves a governance trail.

Hunting shadow AI

Shadow AI is any AI tool adopted without going through formal procurement or IT review, often a free or low-cost web tool a team started using on its own. Daniel checked expense reports for AI subscriptions, asked IT about traffic to known AI services, and held no-blame conversations with program staff. The point was never to punish the team that found a useful tool. It was to get the tool into the inventory so it could be governed. By the end, Daniel's list had grown from six entries to twenty-three.

Consolidating what you found

A raw list is not an inventory. Follow up with detailed interviews with the owners of each reported system, both to fill in the fields and because one interview routinely surfaces a second system the interviewee forgot to report. Then consolidate and deduplicate, since multiple divisions often run the same system, or one system gets reported under three different names, and establish a single canonical name for each. Validate every record with the person responsible for it, so that the entry is something they have agreed to rather than something written about them. Only then classify and prioritize, and decide who maintains the inventory afterwards, how often, and what events force an update.

The Determination: Safety-Impacting vs Rights-Impacting

Every use case has to be classified, because the classification decides how much governance attaches. The memo defines two high-risk categories, and Daniel treated the determination as the most consequential judgment in the whole exercise.

  • Safety-impacting AI is AI whose output could significantly affect the safety of human life or well-being, the climate or environment, or critical infrastructure. Think AI influencing the control of physical systems, emergency response, or the safety of food, drugs or buildings.
  • Rights-impacting AI is AI whose output serves as a principal basis for a decision affecting an individual's civil rights, civil liberties, privacy, equal opportunity, or access to government benefits and services. Think AI influencing eligibility for a benefit, an enforcement action or a hiring decision.

The memo provides presumed categories for both, and an agency can rebut the presumption with documentation. Daniel's rule for his analysts: when in doubt, classify up, and write down the reasoning either way. An undocumented "not rights-impacting" call is exactly what an Inspector General will question. A documented one, even if later revised, shows governance was exercised. Of Daniel's twenty-three systems, nine landed in one of the two high-risk categories, and those nine pulled in the full set of minimum practices.

The determination is not a label applied for tidiness. It is a switch. Flip it to rights-impacting or safety-impacting and you have committed the agency to an impact assessment, real-world testing, ongoing monitoring and meaningful human oversight before that system may keep operating. That is why the reasoning behind each determination has to live in the record, and why the temptation to classify downward is the single most predictable governance failure in this whole area.

Edge cases in the determination

Daniel kept a short list of the calls his analysts found hardest, because these were the ones an auditor would probe. An AI that drafts a document a human then edits and signs is usually not rights-impacting, since the human is the decision-maker, but only if the human review is genuine and not a rubber stamp. An AI used purely for internal efficiency, such as routing mail, is typically neither category, until the routing starts affecting how fast someone's benefit claim is handled, at which point it can edge toward rights-impacting. A commercial tool with an embedded AI feature still counts even though the agency did not build it. For each edge case, Daniel's standing instruction was the same: document the reasoning, name who made the call, and revisit it if the use changes. The edge cases are exactly where undocumented judgment becomes an audit finding.

What the High-Risk Documentation Has to Answer

For the systems that land in either high-risk category, the inventory entry is only the cover sheet. The impact assessment underneath it has to answer six sets of questions, and an assessment that skips one of them will be sent back. Daniel found it easier to hand owners the questions than to hand them a form, because the questions make it obvious when an answer is missing.

Assessment areaQuestions the record must answer
Technical specificationsWhat are the inputs and outputs? What model or algorithm is used? What are its performance characteristics? Has it been tested, and with what results? How often is it retrained or updated?
Data documentationWhat data trains the model? How representative is it of the real-world population? What data quality issues are known? How is the data refreshed?
Risk assessmentWhat are the failure modes? What is the impact if the system fails, and how likely is that? What harm could it cause, and who would be affected?
Fairness and equityHas the system been tested for disparate impact? Do outcomes differ across demographic groups? If they do, are those differences justified, and what is the plan if they are not?
Transparency and explanationCan the system explain its outputs? Do the people affected understand how decisions are made? Is there an appeal or override mechanism?
Governance and oversightWho is responsible? What monitoring is in place? What is the incident response process? How often is the system reviewed, and what would trigger escalation or shutdown?

What Gets Disclosed and What Gets Withheld

The inventory has a public face, but not everything is published. Daniel worked the disclosure question case by case with his agency's legal and security teams. The default is publication, because transparency about government AI is the point. The carve-outs are narrow: use cases whose disclosure would compromise sensitive law-enforcement or national-security functions, or reveal information protected from release, may be withheld or aggregated. The discipline Daniel imposed was that "withhold" required a stated reason tied to a real exemption, not mere discomfort. An agency that withholds everything has defeated the purpose; an agency that publishes a system it should have protected has created a different problem. Each use case therefore carried an explicit public-disclosure flag with a documented basis.

Documentation That Survives IG and GAO Audit

Daniel wrote every record as if a Government Accountability Office (GAO) auditor or his own Inspector General would read it in two years, because one of them would. Audit-surviving documentation has a few properties: it names a responsible owner, it dates every determination, it states the basis for each judgment rather than just the conclusion, and it records the risk-acceptance decision and who made it. A record that says "rights-impacting: yes, basis: affects benefit eligibility, determined by CAIO on 2024-09-12, impact assessment completed, residual risk accepted by CAIO with quarterly review" is defensible. A record that says only "rights-impacting: yes" is an open question waiting to be asked.

Where the inventory should live

Most agencies start in a spreadsheet, and for a small inventory that is a reasonable place to start. Its limits arrive quickly: a spreadsheet works for roughly under fifty systems, cannot keep linked documents such as impact assessments and monitoring reports synchronized with their entries, and offers almost no query or analysis capability. A database or purpose-built inventory system handles larger footprints, links each entry to its supporting documents, and lets you ask questions across the whole portfolio. Dedicated AI governance platforms go further, folding inventory, impact assessment, monitoring and incident response into one place. The progression is normal, and the trigger for moving is usually not the system count but the first time somebody asks a question the spreadsheet cannot answer.

The link to NIST AI RMF MAP

Daniel did not invent his questions from scratch. The MAP function of the NIST AI Risk Management Framework, a voluntary framework rather than a binding requirement, asks exactly what an inventory record needs: what is the system, in what context does it operate, who are the affected parties, and what could go wrong. Running MAP and building the inventory record are very nearly the same activity. Daniel treated the inventory as the place where his agency's MAP work was written down, which meant the two efforts reinforced each other instead of duplicating. The wider frame is worth holding onto: GOVERN establishes who decides, MAP establishes what exists, MEASURE establishes how it performs, MANAGE establishes what to do about problems, and classification establishes how intensive all of that needs to be for a given system.

Keeping the Inventory Alive

The mistake Daniel was determined to avoid was treating the inventory as a one-time project that produces a frozen document. The memo requires annual updating, and the real world moves faster than that: new tools get adopted, vendors push model updates, and a use case that was a low-risk pilot last year becomes a rights-impacting production system this year. He built three habits into the agency's process so the inventory would not rot. First, intake: any new AI procurement or tool adoption had to register a draft inventory record before it went live, so shadow AI could not re-accumulate. Second, an annual re-attestation, where each owner confirms their record is still accurate or flags what changed, including whether a model version update altered the system's behavior enough to revisit its determination. Third, a trigger-based review, so that a material change, such as expanding a tool from a pilot office to agency-wide use, forced a fresh look at the determination rather than waiting for the annual cycle.

Underneath those habits sits an ordinary review cadence. Quarterly, sweep for new systems, systems that have moved from development into production, and systems that have been retired. Annually, review everything: refresh risk classifications where the use has changed and update monitoring status. Handle change management explicitly, so that deployment adds an entry, scope expansion re-opens the risk classification, and retirement marks an entry archived rather than deleting it, because the historical record is what an auditor will want when asking what you were running in a given year. And define the escalation path in advance: a changed risk classification goes to governance leadership, a system showing performance problems goes up with its documentation updated, and repeated problems across systems become a governance change rather than a series of individual fixes.

Then actually use the thing. Governance boards should work from the inventory to pick systems for impact assessment. Measurement teams should use it to decide what to monitor. Leaders should use it to see where governance investment is needed. Auditors and researchers will use it to understand the agency's AI footprint whether you intended that or not. A six-month-old inventory nobody maintains is worse than no inventory, because it gives false confidence to the CAIO who certifies it.

The Artifact: A Complete Use-Case Inventory Record

Daniel's core deliverable was a single record template, applied to all twenty-three systems. Every field exists to answer a question an auditor, the CAIO or the public will eventually ask. Below is the template, filled in for one of his rights-impacting cases, an AI tool that scores benefit applications for likely eligibility. The dates and version numbers are the example's own; the point is that every one of them is present rather than implied.

FieldEntry (example: benefit-eligibility scoring tool)
Use-case nameBenefit Application Eligibility Scoring Assistant
OwnerDirector, Benefits Processing Division (named individual)
PurposePrioritize incoming applications by likely eligibility to speed review; does not deny benefits.
StatusDeployed, agency-wide, following a single-office pilot.
Rights-impacting / safety-impacting determinationRights-impacting. Basis: score influences a decision affecting access to a government benefit. Determined by CAIO, 2024-09-12.
Data sourcesHistorical application records and eligibility outcomes; no third-party data; covered by published system-of-records notice.
Model / vendorVendor-built model under Task Order 04 on existing GSA contract vehicle; model version 2.1.
Human oversightEvery score reviewed by a caseworker with authority to override; no autonomous denial; override rate monitored.
Testing statusAI impact assessment completed 2024-09-30; real-world performance and disaggregated fairness testing completed; quarterly monitoring scheduled.
Public-disclosure flagPublic. No law-enforcement or national-security exemption applies.
Risk-acceptance decisionResidual risk accepted by CAIO on 2024-10-15 with conditions: quarterly fairness re-measurement, and a retraining trigger if any group accuracy gap exceeds 4 points.

That last condition is worth pausing on. The four-point gap is a threshold this agency set for itself as a monitoring trigger, not a legal test and not a line that separates lawful from unlawful. Its value is that it was written down in advance, so nobody argues after the fact about whether an observed gap is big enough to act on. Your agency's number will be your own, chosen with your legal team, and the discipline is the same: decide before you measure, and record who decided.

What the First Inventory Looks Like, and What Audits Find

Daniel's agency was small enough that twenty-three entries covered it. A large federal agency running the same process reported a different shape of result: an initial survey uncovered 47 systems, and further investigation through cloud account audits and owner interviews identified 23 more, for a total of 70. Classification put 8 in the rights-impacting category and 4 in the safety-impacting category, leaving 58 in neither. The 12 high-risk systems needed impact assessments on a six to twelve month timeline, monitoring plans followed on a one to two month timeline, and the governance board took on quarterly review of the whole portfolio. The agency's gain was not the document. It was that governance could finally aim its resources at something specific.

The other view of an inventory is the one an auditor takes. In an Inspector General audit of an agency's AI governance, the inventory is where the audit starts, and the findings tend to fall into the same five buckets: systems operating that are not in the inventory at all, entries for systems no longer deployed, risk classifications that look wrong on their face, systems with no documented monitoring, and systems with no identified responsible official. One audit of this kind reported 15 missing systems, 6 stale entries, 8 questionable classifications, 12 systems without documented monitoring, and 3 without a named official. Every one of those findings is cheap to prevent and expensive to receive, and every one of them is a direct product of an inventory that was built once and then left alone.

Anti-Patterns

  • Inventory without action. The inventory is complete, documented, and used for nothing. It becomes an inert artifact, checked off for compliance and never opened again. Avoid it by giving the inventory a job: governance board review on a schedule, selection of systems for measurement, and tracking of impact assessment progress against it.
  • Gaming the classification. A rights-impacting system is labeled "other" to avoid the documentation burden that the higher classification triggers. This is the most consequential failure in the whole process, because it switches off an entire safeguard set silently. Avoid it with clear classification criteria, judgments supported by written evidence, governance oversight of the calls, and spot checks.
  • Inventory sprawl. The definition of AI system is stretched until every predictive calculation is an entry, and the inventory becomes too unwieldy to govern. Avoid it by holding to a clear definition that requires a machine learning or statistical model rather than ordinary arithmetic, and by applying that definition consistently rather than case by case.
  • Incomplete documentation on high-risk entries. The system is inventoried, the classification is right, and the impact assessment was never finished. The inventory then documents gaps rather than controls. Avoid it by making documentation a condition of deployment, so a system cannot go into production until its required record exists.
  • Treating the inventory as proof of safety. A complete inventory tells you what you are running and what you have decided about it. It does not tell you that those systems are safe, fair or accurate, and it cannot surface a risk nobody wrote down. Read it as a map of your known footprint, not as a clearance.
  • Asking people whether they use AI. The single most reliable way to produce a short inventory is to ask the question in the vocabulary of the policy rather than the vocabulary of the work. Ask what predicts, scores, classifies, routes or generates, and ask it about tools people would describe as software.
  • The one-time drive. A heroic inventory effort with no intake process behind it re-accumulates shadow AI immediately, and the second drive is as expensive as the first. Register new tools at adoption instead.
  • Records written for the file rather than the reader. An entry that states conclusions without their basis is an invitation to be asked. Write the basis, the date and the decision-maker into every determination the first time.

Practice Prompts

  • Inventory your own area. List every AI system you are aware of in your part of the agency. For each one, capture name, purpose, owner, data sources and a provisional risk classification. Expect the list to grow once you start asking colleagues behavioral questions.
  • Fill in the required fields. Take one system and complete the full field set: name and description, component and office, status, techniques, data sources, intended use, performance metrics, risk classification, responsible official, and approval and monitoring status. Note which fields nobody can answer.
  • Draft an impact assessment outline. For a high-risk system, sketch what you would include under each of the six assessment areas. Be specific about what fairness information you would need and where it would come from.
  • Design the maintenance process. Write how systems get identified, how records get updated, what events trigger a re-determination, and who owns the inventory when the person who built it moves on.
  • Write the communication plan. You are rolling out an inventory drive. Draft what you would say to technical teams, to leadership and to compliance, and be honest about what each audience is actually worried about.
  • Run the disclosure call. Take one system that feels sensitive and write the public-disclosure flag with its stated basis. If you cannot name an exemption, the answer is publication.

Reflection

Assess your own organization honestly against four questions. Do you have a centralized inventory of AI systems, and if you do, is it current? If you do not, what is actually preventing it, and is that a resource problem or an ownership problem? For your high-risk systems, are the impact assessments complete, and where are the gaps? Who owns the inventory, how is it maintained, and is it a living resource or an artifact somebody produced once for a deadline?

Then the question underneath all of them. If your inventory were complete and accurate tomorrow morning, what would you learn about your agency's AI footprint that you do not know today, and would that knowledge change how you govern? Most people who work through this exercise discover that the number of systems is not the surprise. The surprise is how many of them have no named owner, and how many were classified by someone who never wrote down why.

Glossary

  • AI use-case inventory. The comprehensive catalog of AI systems an organization operates or is developing, documenting the key information about each and maintained on a defined cadence.
  • Chief AI Officer (CAIO). The senior official accountable for an agency's AI governance, coordination and risk management, including waiver decisions and certification of the inventory.
  • Impact assessment. Formal documentation of a system's purpose, data, risks, fairness implications and management approach, required for safety-impacting and rights-impacting systems.
  • Risk classification. The determination of whether a system is safety-impacting, rights-impacting or lower-risk. It sets how intensive the governance around that system must be.
  • Rights-impacting AI. AI whose output is a principal basis for a decision affecting an individual's civil rights, civil liberties, privacy, equal opportunity, or access to benefits and services.
  • Safety-impacting AI. AI whose output could significantly affect human safety or well-being, the climate or environment, or critical infrastructure.
  • Responsible official or system owner. The named person accountable for a specific system's operation, compliance and performance. A record without one is an audit finding waiting to happen.
  • Shadow AI. Any AI tool adopted outside formal procurement or IT review, typically a low-cost web tool a team started using on its own initiative.

Closing

Daniel's CAIO ended up with something real to certify: a complete inventory, a defensible determination behind every high-risk flag, a public version ready to publish, and a documented risk-acceptance decision on file for each rights-impacting system. The six-entry list he inherited would have failed the first audit question. The twenty-three-entry inventory, each record built to answer the questions before they were asked, was the difference between compliance on paper and compliance that holds.

A bigger inventory is a better inventory, and that is the least intuitive thing about this work. Finding seventeen systems nobody had reported is not evidence that the agency was out of control; it is the exercise working. The failure mode is the short list, produced quickly, that lets everyone believe the footprint is small. Creating a good inventory is usually the first concrete step an agency takes toward mature AI governance, because it forces the hard questions in a form nobody can defer: what are we actually operating, who owns it, is it working as intended, and does it meet the standards we say we hold.

Key Takeaways

  • The requirement set has five moving parts. A designated CAIO with real governance structures, a maintained and published inventory covering systems in development as well as production, minimum practices for high-risk AI, public reporting to a specified standard, and waivers that are documented and revisited. Build the inventory to serve all five.
  • Confirm the instrument that binds you. Federal AI guidance changes between fiscal years. The substance in this lesson has proved durable, but check with your CAIO or counsel which memo governs before building a compliance process on any summary.
  • The inventory drive is a hunt, not a survey. Run surveys, contract review, cloud and SaaS audits and shadow-AI hunting in parallel, and ask behavioral questions such as "does it predict, score, classify or generate?" rather than "do you use AI?"
  • The determination is a switch, not a label. Flipping a system to rights-impacting or safety-impacting commits the agency to impact assessment, real-world testing, monitoring and meaningful human oversight, which is exactly why under-classification is the most damaging shortcut available.
  • When in doubt, classify up and write it down. An undocumented "not rights-impacting" call is what an Inspector General will challenge. A documented determination shows governance was exercised even if it is later revised.
  • Default to public disclosure, withhold only with a stated basis. Narrow law-enforcement and national-security exemptions can justify withholding; discomfort cannot. Flag every record's disclosure status with its reason.
  • Write every record to survive an audit. Name an owner, date each determination, state the basis rather than only the conclusion, and record who accepted the residual risk. Audits of AI governance start at the inventory and find missing systems, stale entries, questionable classifications, absent monitoring and unnamed officials.
  • The inventory is your NIST AI RMF MAP work, written down. The framework is voluntary, but MAP and the inventory record ask the same questions, so run them as one effort.
  • A bigger inventory is a better inventory, and a stale one is worse than none. Daniel's list grew from six to twenty-three. Register new tools at adoption, re-attest annually, and force a fresh determination whenever a system's use materially changes.

Frequently Asked Questions

What counts as an AI system for inventory purposes? The government-wide definition is broad: a system using machine learning, statistical models or other computational methods to make or support decisions. Two failure modes bracket it. Define it too narrowly and you miss the commercial tool with an embedded AI feature and the contractor deliverable nobody looked inside. Define it so broadly that every spreadsheet formula qualifies and the inventory becomes unusable. The workable line requires a machine learning or statistical model rather than ordinary calculation, applied consistently rather than renegotiated for each awkward case.

Do systems still in development belong in the inventory? Yes. The inventory covers systems in use and systems in development, and the status field exists precisely to distinguish them alongside retired systems. Including a system in development is also the cheapest moment to make its determination, because the answer can still change the design rather than requiring a retrofit after deployment.

Who makes the rights-impacting determination? The determination belongs to the agency's AI governance structure, with the CAIO accountable for it, informed by the system owner and by legal. What matters as much as who decides is that the decision is recorded with its date and its basis. In practice most disputes are not about the definition; they are about whether an output is a principal basis for a decision or merely one input among several, and that is a judgment somebody has to make in writing.

Can we withhold a use case from the published inventory? Only with a stated basis tied to a real exemption, such as compromising sensitive law-enforcement or national-security functions or revealing information otherwise protected from release. Aggregation is sometimes an option where full disclosure is not. What is not acceptable is withholding because publication would be awkward, and an agency that withholds broadly should expect that pattern itself to become a finding.

Is a spreadsheet good enough? For a small inventory, yes, and most agencies start there. It stops being good enough at roughly fifty systems, or earlier if you need linked impact assessments and monitoring reports to stay synchronized with their entries, or the first time someone asks a portfolio-wide question the sheet cannot answer. Move to a database or a governance platform when that happens rather than on principle.

What do we do when we find shadow AI? Get it into the inventory, not into a disciplinary process. The team that adopted a useful tool without going through review has told you something about your review process as much as about their judgment. Classify the tool, document it, apply whatever practices its classification requires, and then fix the intake path so the next adoption registers a draft record before it goes live.