←
AI for Government
Strategic · M12 · lesson 12 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Supply Chain Security: End-to-End
📖
now learning

AI Supply Chain Security: End-to-End

15 min

Patrick Brennan is the Chief Information Security Officer of a federal agency responsible for managing national physical infrastructure records: 1,100 employees, highly sensitive mapping data, and a Chief AI Officer who arrived in Patrick's office in April to report that the agency had just deployed an AI summarization tool. The tool was built on a third party large language model, licensed through a software reseller, running on cloud infrastructure managed by a fourth contractor, with a data pipeline built by a fifth vendor. Patrick asked one question: where is the security assessment for the model itself? The answer was that the model had been assessed as part of the software product's Authority to Operate review, but the model provider's own security practices had never been independently examined. Nobody had mapped the complete chain. Patrick spent the next six weeks building what he now calls the agency's AI supply chain map: every vendor, every contractor, every component, and every security control or gap at each node. What he found was not catastrophic. It was simply not what anyone had assumed.

Where this lesson sits

A companion lesson, AI Supply Chain Risk, covers why AI supply chains behave differently from software supply chains, the policy instruments that already apply, and the choice between open weights and a closed interface. This lesson takes the same territory from the other end. It assumes you accept the risk and asks what the end to end control set looks like in practice: what the chain contains, how you produce a map of it, how you assess each node, what you require from each supplier, how you verify what arrives, and how you respond when something goes wrong in a chain where five organizations can each credibly say the problem was somebody else's.

What the chain actually contains

A government AI system is rarely a single product from a single vendor. It is a stack of components from multiple suppliers, integrated by one or more contractors, operating on infrastructure maintained by yet another provider. Understanding the security posture of the full stack rather than the top layer application is the substance of this discipline. Counting the categories in Patrick's deployment gives five distinct kinds of component, each with its own security profile, its own failure modes and its own set of people who believe someone else is responsible for it.

The foundation model. The base model, whether a large language model, an image classifier or a predictive algorithm, was developed by its creator using training data, software libraries and hardware that may themselves carry supply chain vulnerabilities. A model trained on poisoned data produces outputs that reflect the poisoning. A model built with compromised libraries may carry vulnerabilities that persist through deployment. Agencies using commercial foundation models cannot inspect the training process, but they can require documentation of the developer's security practices and attestations that the model has undergone adversarial testing.

The training data. If the agency is fine-tuning a commercial model on government data, or training a custom model from government datasets, the provenance and integrity of that data is a security matter and not only a data governance matter. Adversarial data poisoning, meaning the deliberate introduction of corrupted or manipulated records into a training set to influence the resulting model's behavior, is an established attack vector. Agencies training on data they did not exclusively control throughout its lifecycle should build integrity verification into the data preparation process rather than trusting the custody chain.

The inference infrastructure. The servers, cloud environments and interfaces through which the model receives inputs and returns outputs are a critical layer, and the one most often assumed to be somebody else's problem. Where inference runs on infrastructure shared with other tenants, side channel effects that let one tenant observe aspects of another's activity are a consideration for sensitive workloads. The inference interface must be protected against unauthorized access, and inputs and outputs should be logged for audit. An application inside an authorized boundary whose inference endpoint sits outside it has an authorization that does not describe the system.

The integration layer. The software connecting the model to existing systems, meaning the pipeline that feeds it inputs and the workflow integration that delivers outputs to users, is usually written by a systems integrator and is a common source of gaps. Integration code is frequently held to a lower standard than the core application, and it is where data leaks actually happen: a pipeline configured to send more fields than the model needs, an output delivery path that writes sensitive content into an unencrypted log, a retry mechanism that stores request bodies. This layer rarely appears in vendor security documentation because no vendor considers it theirs.

The reseller or systems integrator. Many government AI deployments are procured through a reseller or integrator rather than directly from the developer. That party has access to the deployment environment during implementation. Their personnel, their systems and their security practices are part of your chain. Software bill of materials requirements, which oblige contractors to document the components included in a delivered system, should extend to AI deployments and should require disclosure of foundation model provenance, training data sources and third party libraries in the AI component rather than only the conventional software components.

Producing the map

The first step is the artifact Patrick had to build from nothing: a complete record of every component, every supplier and every data flow in the system. Everything downstream depends on it. Risk assessment needs it to know what to assess. Vendor requirements need it to know who to impose them on. Incident response needs it to know who to call outside business hours. Without it, each of those activities operates on somebody's recollection of how the system was assembled, which is reliably wrong in the same direction: simpler than reality.

The map should record, for each element: the name and version of the foundation model and its developer; the training data sources and their custodians; the infrastructure provider and the specific services in use; the integration contractor and the components they built; the reseller and their role; and the interfaces between components, including data formats and authentication mechanisms. A table beats a diagram for this purpose, because a table has empty cells and a diagram hides them.

NodeWho supplies itWhat data it touchesControls in placeAccountable inside the agency
Foundation modelModel developer, reached through the resellerEvery input submitted for inferenceDeveloper attestations, version pinning, integrity verificationNamed system owner
Training or fine-tuning dataAgency, or agency plus contractorWhatever was included, including anything included by mistakeProvenance record, integrity verification, review of what was in scopeData owner
Inference infrastructureCloud or hosting providerAll inputs and outputs in transit and at restAuthorized environment at the right impact level, access control, loggingSecurity office
Integration layerSystems integratorWhatever the pipeline actually moves, which is usually more than requiredCode review, field level scoping, log content reviewProgram office and integrator lead
Reseller or integrator personnelContract holderDeployment environment during implementationPersonnel vetting, scoped and time limited access, access recordsContracting officer representative

The final column is the one agencies leave out and the one that makes the map operational. A node with no accountable name inside the agency is a node nobody will check, nobody will renew and nobody will investigate. Patrick's first version of the map had far fewer names in that column than it had rows, and closing that difference took longer than assembling the technical detail.

Assessing risk at each node

Every node gets two questions. What data does this component have access to, and what controls exist at this component? The answers identify the gaps, meaning the nodes where access is not matched by commensurate control, and they identify the dependencies whose failure would be hardest to absorb. That is a deliberately small method. It is small because it has to be applied to every node by people who have other jobs, and a method that takes a day per node is a method that gets applied to the two nodes somebody was already worried about.

Two refinements earn their cost. The first is asking what the component would let an adversary do rather than what it is supposed to do, which surfaces the integration layer's real risk profile: a pipeline with read access to a case management system is an exfiltration path regardless of what it was built for. The second is recording, for each node, how you would know if something went wrong there. Most nodes in a first pass produce the same honest answer, which is that you would not know, and that answer is the finding. Detection gaps are cheaper to close than they are to discover during an incident.

Where a node's controls are supplied by someone else, record whose assurance you are relying on and what form it takes. There is a real difference between a control you operate, a control a provider operates under an authorization you can read, and a control a vendor says exists. All three can be acceptable. Only the third is frequently mistaken for one of the other two.

What to require from each supplier

Government AI procurement should require foundation model suppliers to attest to their development security practices. The attestations should address the data governance practices applied to training data; whether adversarial testing, meaning deliberate attempts to elicit harmful, manipulated or out of distribution outputs, was conducted before release; the software security practices used in development; and the model's vulnerability history, including publicly disclosed jailbreaks or prompt based attacks. Ask for the same from the integrator about the code they wrote, because the integration layer carries as much of the practical risk and receives a fraction of the scrutiny.

Be clear with yourself about what an attestation is. It is a statement by the party with the strongest incentive to be reassuring, and its value is that it creates a record and a form of accountability if it turns out to be false. It is not verification. An agency that collects attestations and believes it has assessed the chain has replaced an unknown with a document about the unknown. Where a claim genuinely matters to the mission, pair the attestation with something you can check yourself: your own evaluation on your own data, evidence of an independent assessment, or a right of inspection you actually exercise.

The Authority to Operate problem in Patrick's agency is the same mistake in institutional form. The product had an authorization. Everyone read that authorization as covering the model, because the model was inside the product. An authorization covers what its assessment examined, and if the assessment did not reach the model provider's own practices then the authorization does not describe them. When you inherit an approval, read what it assessed before you rely on what it implies.

Provenance verification and its limits

Provenance verification confirms that the model running in your environment is the model that was assessed and has not been altered along the way. The mechanism is a cryptographic checksum or a digital signature: a model tampered with after assessment produces a different fingerprint than the assessed version. This is a basic integrity control, it is cheap, and remarkably few government AI deployments perform it. Require it at acceptance and repeat it after every change to the deployed artifact.

Then be precise about what it proves. A verified checksum establishes that your file matches the publisher's file. It says nothing about whether the publisher's model was trained on poisoned data, whether it carries a backdoor, or whether it behaves acceptably on your mission. Integrity verification and behavioral evaluation answer different questions and neither substitutes for the other. A program that verifies fingerprints and never tests behavior has confirmed that it faithfully downloaded something it has never examined, which is a real improvement over the previous state and is not assurance.

The same distinction applies to the bill of materials. A complete inventory tells you what is in the system. It does not tell you that what is in the system is safe, and a vendor who fills every field has demonstrated cooperation rather than security. The inventory's power is that it converts unknown risk into named risk that someone can then assess, and its most valuable output is often an empty field, which is a finding rather than a formatting problem.

Keeping the map current

A supply chain map is accurate on the day it is finished and decays from then on. Providers update models, integrators patch pipelines, cloud services are reconfigured, resellers are acquired and subcontractors change. The control is a short list of events that trigger re-verification: a base model version change, a change of hosting environment or region, a change of integrator or subcontractor personnel with environment access, a change of ownership at any supplier, a new data source entering the pipeline, and a publicly disclosed vulnerability in any listed component. Each trigger has an owner and a defined response, and the triggers only work if a contract obliges suppliers to tell you when they occur.

Set a floor as well: a scheduled review of the whole map on a defined cycle, independent of whether anything triggered. Triggered reviews catch the changes you were told about. The scheduled review catches the ones nobody reported, and in most agencies the first scheduled review after the initial map finds at least one node that changed without notice. Record what changed and why nobody knew, because that second answer usually points at a missing contract clause rather than at a careless supplier.

Incident response across multiple vendors

An incident in a multi-vendor AI chain is harder to attribute and harder to contain than one in a single system environment. When an anomalous model output or an unauthorized data access appears, the investigation must first establish which component produced it, and the components are operated by organizations with different logging practices, different retention periods, different legal exposure and different willingness to help at speed. Without a map and without procedures designed for this shape of problem, the investigation is slower, more expensive and more likely to end inconclusively.

The response plan should include a notification requirement covering every supplier whose component may be implicated; a defined timeline for supplier cooperation and data provision during an investigation; and pre-negotiated access rights letting your security team obtain logs and forensic data from supplier controlled components while the incident is live. Negotiate all three before award. After an incident begins, every one of them becomes a contract discussion conducted by people who are also trying to contain the incident, and the delay is measured in the days when evidence is still retrievable.

Exercise the plan rather than filing it. Run a tabletop on an anomalous output whose source is ambiguous between the model and the integration layer, because that is the realistic case and it is precisely where attribution stalls. The exercise usually surfaces two findings: that the logging needed to distinguish the two does not exist on both sides, and that nobody knows who at the supplier answers the phone outside business hours. Both are fixable in advance and neither is fixable during.

Assessing readiness and resourcing the work

Before standing this up as a program, assess where the agency is. Current capability: who can actually assess a model supplier's practices today, what gaps exist, and where the largest exposure sits. Organizational readiness: is the agency prepared for a security function to hold up a deployment over an unmapped node, and what would help. Stakeholder alignment: the security office, the AI office, program offices and contracting each have interests here, and they are not identical. Resource constraints: what staff time genuinely exists, and how do you work inside it rather than designing a process for a team you do not have.

Then plan it as an initiative. Goal clarity: what does supply chain assurance mean here, and what would success look like in terms a program executive recognizes. Action planning: which systems get mapped first, in what sequence, and with what resources. Risk management for the effort itself: the most common failure is a map that is built once, so plan for the maintenance before the first map is finished. Stakeholder engagement: the contracting office is the pivotal partner, because every durable control in this lesson ends up as a contract term, and a security requirement that never reaches a solicitation is an opinion.

Running it as a program

Execution turns on process and measurement. Capability building asks what skills the work requires and how they survive turnover. Process design asks how a new AI acquisition gets mapped as a matter of routine rather than because someone remembered, how the assessment is recorded, and how the record is found later by someone who did not write it. Continuous improvement asks what you monitor and what you change as a result. Useful measures include how many deployed AI systems have a current map, how many nodes lack an accountable name, how many supplier notifications arrived through the contract rather than by accident, and how long the last attribution exercise took from alert to identified component.

Coordinate rather than duplicate. Agencies already run supply chain risk functions for conventional information technology, already track foreign supplier concerns, and already hold vendor performance data. Building a parallel AI specific version of each is how this work becomes a burden that gets abandoned at the next reorganization. Extend the existing inventory with the AI specific fields, add the AI components to the existing supplier review, and take the questions that are genuinely new, meaning model provenance, training data and behavioral evaluation, as additions rather than as a separate discipline.

Sustaining it past the first map

Successful pilots have to scale and survive. Scaling asks how mapping one system becomes mapping a portfolio without the effort growing linearly, which mostly means templates, reusable supplier records and a shared map for shared infrastructure. Funding asks who pays for the work and what happens to the reviews when that source changes. Organizational embedding asks whether this is a role or a person, whether it is in the standard acquisition path, and what remains when the person who cared about it moves on.

Patrick's map exists because a CISO spent six weeks on it after a surprise. That origin is also its weakness. Work that begins as one person's response to one incident tends to end when that person's attention moves, unless it is converted into something structural: a required artifact at an acquisition gate, a named owner per node, a scheduled review with a calendar entry, and contract clauses that keep supplying the information after the initiating anxiety has faded. Convert it while the surprise is still recent, because that is the only period in which the conversion is easy to justify.

What Patrick found

The six week mapping exercise produced no breach and no emergency. It produced a list. The model provider's security practices had never been examined by anyone at the agency, and the authorization everyone had relied on did not reach them. The integration pipeline moved more fields than the summarization task required, because the integrator had built against the whole record rather than the needed subset. Nobody had verified that the deployed model matched what had been assessed. Reseller personnel retained environment access from implementation. And no plan existed for who would be called, in what order, if an output turned out to be wrong in a way that mattered.

None of those findings required special tooling. Each of them required someone to ask a specific question about a specific node and write down the answer, which is the entire method. The difference between Patrick's agency before and after was not new technology or new budget. It was a document that told the truth about how the system was assembled, with a name next to every part of it, and a set of contract terms that would keep the document true after the person who built it moved on.

Anti-patterns

  • Inherited authorization as coverage. Reading an approval granted to a product as an assessment of every upstream component inside it. An authorization covers what its assessment examined, and the model provider's own practices are often outside that scope.
  • Attestation as verification. Collecting supplier statements and recording the chain as assessed. An attestation creates accountability if it proves false; it does not tell you anything you checked.
  • Checksum theater. Verifying integrity once at acceptance and never again, or treating a matching fingerprint as evidence of safety. It proves your file matches the publisher's file, which says nothing about poisoning, backdoors or behavior.
  • Inventory mistaken for security. Treating a fully populated bill of materials as assurance. A complete inventory converts unknown risk into named risk, which is where assessment starts rather than where it ends.
  • The unowned node. A map with no accountable name against a component, which guarantees that nobody renews it, checks it or investigates it.
  • The integration blind spot. Scrutinizing the model while nobody reviews the pipeline that feeds it, which is where over-collection, unencrypted logging and retry storage actually leak data.
  • The endpoint outside the boundary. Authorizing an application while the inference endpoint that processes the data sits somewhere the authorization does not describe.
  • The one time map. Building the map once and never defining the events that make it wrong, so a model version change or a subcontractor swap silently replaces the system you assessed.
  • Access that outlives implementation. Reseller and integrator personnel retaining environment access long after the work that justified it ended, because nobody owns the removal.
  • Response plans negotiated during the response. Discovering at incident time that you have no right to supplier logs, no cooperation timeline and no out of hours contact, and negotiating all three while evidence ages out of retention.
  • A parallel AI security program. Building an AI specific duplicate of supply chain functions the agency already runs, which doubles the cost and is abandoned at the next reorganization.

Practice prompts

  1. Take one deployed AI system and build the map: model and version, developer, training and fine-tuning data and custodians, hosting provider and services, integrator and what they built, reseller and role, and the interfaces between them. Mark every cell you had to guess.
  2. Add the accountability column to that map and try to fill it with real names. Count the nodes that end up empty.
  3. For one node, answer the detection question in writing: if this component were compromised or misbehaving, how exactly would you find out, and who would see it first.
  4. Read the authorization your agency is relying on for one AI system and write down what it actually assessed, then compare that against the five component categories in this lesson.
  5. Draft the supplier attestation questions you would require in your next AI acquisition, and mark which answers you could independently check and which you would simply be accepting.
  6. Write the re-verification trigger list for one system, name an owner for each trigger, and identify which triggers depend on a supplier telling you something your contract does not currently require them to tell you.
  7. Run a tabletop on an anomalous output that could plausibly have come from either the model or the integration layer. Note where attribution stalls and what logging would have resolved it.
  8. Assess your readiness honestly: who could assess a model supplier's practices today, what that costs in staff time, and which existing function should absorb this rather than a new one.

Reflection

The most uncomfortable part of Patrick's six weeks was not any single finding. It was discovering that every one of the five organizations in the chain had behaved reasonably and the result was still a system nobody could describe. The reseller sold what it was asked for. The integrator built against the specification it received. The provider published what its customers usually ask for. The security office assessed what was submitted. The program office deployed something that had been approved. Nobody was negligent and the whole was unassessed, because assessment of the whole was not anybody's job. Consider the AI systems your own agency runs, and ask a narrower question than whether they are secure: for each one, who holds the picture of the entire chain, and if the answer is nobody, whose job would it be to notice.

Glossary

  • AI supply chain map. The record of every component, supplier, data flow and control in an AI system, with an accountable name inside the agency against each node.
  • Node. A single component or supplier in the chain, assessed on what data it can reach and what controls exist at it.
  • Foundation model. The base model on which a deployed AI capability is built, developed by a third party using data, libraries and hardware you cannot inspect.
  • Integration layer. The pipeline and workflow code connecting a model to existing systems, usually built by an integrator and usually held to a lower security standard than the application.
  • Inference infrastructure. The hosting environment and interfaces through which the model receives inputs and returns outputs, including the endpoint that must sit inside the authorization boundary.
  • Data poisoning. Deliberate introduction of corrupted or manipulated records into a training set to influence the resulting model's behavior.
  • Provenance verification. Confirming through a cryptographic checksum or signature that the deployed model is the assessed model, unaltered in transit.
  • Adversarial testing. Deliberate attempts to elicit harmful, manipulated or out of distribution outputs from a model before it is relied upon.
  • Software bill of materials. A documented inventory of the components in a delivered system, extended for AI to cover model provenance, training data sources and AI specific libraries.
  • Attestation. A supplier's formal statement about its own practices, which creates accountability if false and is not the same thing as verification.
  • Re-verification trigger. A defined event, such as a base model version change or a supplier ownership change, that obliges the agency to reassess the affected part of the chain.

Closing

End to end supply chain security is less technical than it sounds. The hard part is not cryptography or threat modeling; it is producing an honest document about how a system was actually assembled, keeping a name against every part of it, and writing enough into the contract that the document stays true after the people who built it move on. Patrick's map found no attacker. It found five reasonable organizations, one unexamined provider, one over-collecting pipeline, one unverified artifact, some access that should have ended months earlier, and no plan for the phone call. That is the normal result, and it is worth six weeks in almost every agency that has not yet spent them.

Key takeaways

  • Government AI systems are chains, not products. Foundation model, training data, inference infrastructure, integration layer, and reseller or integrator each carry a distinct security profile, and each is usually assumed to be someone else's responsibility.
  • Build the map before authorization, not after a surprise. Component, supplier, data flow, control and, most importantly, an accountable name inside the agency for every node. A node with no name is a node nobody will check.
  • Assess each node with two questions. What data can it reach, and what controls exist here. Add how you would know if it went wrong, and expect the first honest answer to be that you would not.
  • An inherited authorization covers what it assessed. Read the scope before relying on the implication, because a product level approval frequently never reached the model provider's own practices.
  • Attestations create accountability, not assurance. Require them from the model supplier and the integrator, and pair the ones that matter with something you can check yourself.
  • Verify provenance, and know its limit. A matching fingerprint proves your file is the publisher's file. It says nothing about poisoning, backdoors or behavior, so integrity verification never substitutes for evaluation.
  • Extend bill of materials requirements to AI components. Model provenance, training data sources and AI libraries belong in the inventory, and an empty field is a finding rather than a formatting problem.
  • The integration layer leaks. Over-collecting pipelines, unencrypted logs and retry storage cause more real data exposure than the model does, and no supplier considers that layer theirs.
  • Define re-verification triggers and a scheduled floor. Version changes, hosting changes, personnel changes, ownership changes, new data sources and disclosed vulnerabilities, backed by a contract that obliges suppliers to tell you.
  • Negotiate incident cooperation before award. Notification duties, a cooperation timeline and pre-agreed access to supplier logs are cheap in a solicitation and nearly unobtainable once an incident is live.

Frequently Asked Questions

Our AI product already has an authorization. Is that enough?

Read what the assessment actually examined before answering. Authorizations are granted against a defined scope, and in Patrick's agency the scope covered the software product while the model provider's own security practices were never independently reviewed. That is a common pattern rather than an unusual failure. The question to ask of any inherited approval is which of the five component categories it reached, and the answer is usually available in the assessment documentation rather than in the approval letter everyone forwards.

How is this different from the supply chain risk lesson?

That lesson establishes why AI supply chains are harder to inspect than software ones, which policy instruments already apply, and how to think about open weights against a closed interface. This one is the operating control set: the map, the node assessment, the supplier requirements, provenance verification, re-verification triggers and multi-vendor incident response. Read that one for the framing and the acquisition strategy, and this one when you have to produce the artifact and run it as a program.

The vendor will not tell us what the model was trained on. What now?

Record it as a gap and decide explicitly whether you are accepting it, rather than letting the field be marked complete because something was written in it. Sometimes a category level summary plus a statement of what controls existed against poisoning is a workable compromise. Sometimes the gap is unacceptable for the mission, and that is a source selection outcome, which is far cheaper to reach before award than after deployment.

We are a small agency. Do we need all five categories mapped?

Yes, but the map is much smaller than it sounds when you have three AI systems rather than three hundred. The five categories are questions, not workstreams, and a single page per system answers them for most small deployments. Start with the systems that touch sensitive data or make decisions about people, and reuse the supplier rows, since the same cloud provider and the same reseller usually appear across several systems.

Who should own the AI supply chain map?

The security office typically holds the map, the system owner holds the individual nodes, and the contracting office holds the terms that keep it accurate. What matters more than the specific split is that each node carries one name and that the acquisition path requires the map before deployment. A map owned by a single enthusiastic person is an artifact with an expiry date attached to that person's next role change.

How often should we revisit a completed map?

On defined triggers plus a scheduled floor. The triggers are the events that make the map wrong: base model version changes, hosting or region changes, integrator or subcontractor personnel changes, supplier ownership changes, new data sources and disclosed vulnerabilities in listed components. The scheduled review exists because triggers only fire for changes you were told about, and the first scheduled review after an initial map usually finds at least one change nobody reported.