←
AI for Government
Proficient · M11 · lesson 11 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Supply Chain Risk
📖
now learning

AI Supply Chain Risk

15 min

Lieutenant Colonel Sandra Okeke, an acquisition lead at a Defense Department program office, was about to field an AI tool that helped analysts triage incoming intelligence reports. The contractor had built it on an open weight foundation model, meaning a large AI model whose internal parameters are publicly downloadable, pulled from a popular online model hub. During a final security review, a junior analyst asked a simple question Sandra could not answer: "Who trained this model, on what data, and how do we know nobody tampered with it before we downloaded it?" The contractor's response was a shrug and a license file. Sandra realized she was about to deploy, into a national security workflow, a piece of software whose origin and contents were essentially unknown.

That is AI supply chain risk in one scene. Stated more formally, it is the risk that an agency cannot vouch for where its AI capability came from, what went into it, who controls its updates, and what happens if a link in the chain fails or is compromised. Every AI system you buy or build sits on components someone else made, and you inherit the risk of every one of them, including the ones you will never see.

Why AI supply chains are harder than software supply chains

Government has spent years securing the software supply chain, which means knowing what code goes into systems and being able to prove it. AI adds three links that traditional software does not have, and each one is harder to inspect than source code.

  • The model itself is often a black box of billions of numbers. You cannot read it the way you read source code to see what it learned or what was hidden inside it.
  • The training data is usually invisible to you. You rarely know what the model learned from, whether it included copyrighted, biased or poisoned material, or whether it memorized sensitive information.
  • The behavior is probabilistic. The same input can produce different outputs, so a hidden flaw may surface only rarely, which makes it hard to catch in any finite amount of testing.

This is why the federal AI risk framework from the National Institute of Standards and Technology, which is voluntary rather than binding, treats provenance and third party risk as first class concerns under its governance function, and why the generative AI profile published in 2024 names data provenance, model lineage and third party component evaluation explicitly. You cannot govern what you cannot trace. The discipline this lesson teaches is forcing the whole chain to become one object under governance, instead of a set of unrelated procurements that nobody has ever seen assembled on a single page.

The full chain, end to end

Federal AI systems rarely originate inside agencies. Data comes from vendors, contractors or open sources. Pre-trained models arrive from external developers, through published weights or through a served interface. Fine-tuning pipelines run on commercial cloud platforms. Hardware acceleration comes from a narrow set of suppliers, with custom silicon programs inside the large cloud providers narrowing it further. Systems integrators tie the pieces together, and each of them subcontracts. Written out in full, the chain runs through training data, pre-trained foundation models, fine-tuning datasets, model weights, evaluation benchmarks, inference infrastructure, interface providers, reseller integrators, hardware accelerators, operating systems, security tooling and maintenance support.

Every one of those links can fail in a way you do not control. It can fail silently, be subverted deliberately, be restricted by export control, become subject to acquisition by a foreign owner, or simply become unavailable at the moment the agency needs it most. Hardware is the clearest case: export controls issued in 2022, tightened in 2023 and extended in 2024 now shape which AI accelerators federal agencies and their contractors can obtain at all, which turns a procurement question into a geopolitical one. Data residency, secure enclaves and confidential computing arrangements, and the physical security of accelerator clusters, all belong in the same conversation rather than in a separate infrastructure review.

Walk every AI acquisition through four supply chain links. We will trace Sandra's intelligence triage tool through each of them.

Model provenance: where did this model come from?

Provenance means the documented origin and history of the model. Sandra needs to know who created it, which exact version she has, and that the file she downloaded is the genuine one rather than a tampered copy. The basic control is a cryptographic checksum, a digital fingerprint proving the file matches the publisher's original. No verified fingerprint, no deployment. Provenance also covers licensing, which is a real constraint rather than paperwork: community licenses, provider terms of service, and research only licenses carry different rights, and a license that prohibits your intended use is a finding no amount of technical review will fix.

Training data supply chain: what did it learn from?

You want a documented account of the data used to train or fine-tune the model. Was it scraped from the open internet, including unvetted sources? Could an adversary have planted poisoned examples designed to trigger bad behavior on specific inputs? For Sandra's use, data poisoning is a genuine national security concern, because a model deliberately trained to misclassify certain report types could blind analysts to a specific threat. Federal policy has moved in the same direction from the privacy side: the 2024 executive order on foreign adversary access to bulk sensitive personal data reaches directly into AI training data supply chains.

Third party and dependency risk: whose code is in here?

The model arrives wrapped in software libraries, served from a vendor's platform, running on a cloud. Each is a third party with its own vulnerabilities. Sandra needs to know which cloud the inference runs on and whether that environment is authorized at the appropriate impact level, since a model processing federal data must sit inside an authorized boundary and the inference endpoint has to be in that boundary too. She also needs the list of open source libraries bundled in, because the package ecosystems that federal AI workloads depend on, including PyPI for Python and npm for Node.js, have repeatedly experienced typosquatting and dependency confusion attacks.

Foundation model evaluation: does it behave safely here?

Finally, evaluate the model against your actual mission and threats rather than the vendor's benchmark. Test it on your data, probe it for the failure modes that matter to you, and document the results. A model that scores well on a public leaderboard may behave badly on specialized intelligence reports. The AI specific failure modes belong in this evaluation explicitly: jailbreak techniques cataloged in the federal adversarial machine learning taxonomy and in the ATLAS knowledge base, prompt injection, backdoors, data leakage, and availability limits including rate limiting. Version drift deserves its own test, because a provider updating a model changes the system you evaluated without changing anything in your contract.

You do not just buy an AI system. You adopt everyone who ever touched it, including the ones you will never meet.

The policy backbone you are already inside

The statutory and policy basis for all of this is substantial and it converges. Executive Order 14028, issued in 2021, set the federal software supply chain agenda after a major software update compromise, directed agencies to require software bills of materials and established CISA's role in software supply chain security. OMB Memorandum M-22-18, issued in 2022, required agencies to obtain self attestation from software producers that they follow secure development practices aligned to the NIST Secure Software Development Framework, and M-23-16, issued in 2023, updated and extended that guidance. NIST SP 800-161 Revision 1, published in 2022, formalizes cybersecurity supply chain risk management as a discipline, and NIST SP 800-53 Revision 5 carries the supply chain control family that implements it.

On the AI side, Executive Order 14110, issued in 2023, extended the focus to AI specific supply chain issues including reporting for dual use foundation models, and OMB Memorandum M-24-10, issued in 2024, requires agencies to address third party and supply chain risk as part of AI risk management and vendor oversight for rights impacting and safety impacting AI. FISMA and FedRAMP apply to the systems and cloud services underneath. Section 889 of the FY2019 National Defense Authorization Act restricts covered telecommunications equipment from specified foreign suppliers, and the NDAA Section 1260H list identifies companies for related due diligence. The Federal Acquisition Supply Chain Security Act and the Federal Acquisition Security Council it created govern exclusion and removal orders. None of this is an AI regime bolted on beside the existing one. It is the existing supply chain regime, applied to a new class of component.

The AI bill of materials

The single most useful artifact here borrows from the software world. A software bill of materials lists every component in a piece of software, and standard machine readable formats for it already exist, including CycloneDX and SPDX. An AI bill of materials does the same for an AI system, and it is precisely the document Sandra's contractor could not produce. The concept is still maturing in federal guidance, developed alongside model cards and the datasheets for datasets literature, so expect the field list to evolve. Require it as a contract deliverable now anyway. At minimum it should record the following.

FieldWhat it recordsSandra's check
Base model and versionExact model and release usedNamed and version pinned
Publisher and licenseWho made it and the terms of useIdentified, license permits government use
Integrity fingerprintChecksum proving the file is genuineVerified against the publisher
Training data summarySources, known limitations, poisoning controlsDocumented, not "unknown"
Fine-tuning dataWhat the vendor added on topDisclosed and reviewable
DependenciesLibraries and their versionsScanned for known vulnerabilities
Hosting and authorizationWhere it runs and its security accreditationAuthorized environment at the right impact level
Evaluation resultsMission relevant test outcomesTested on real reports
Update and notification termsWho may change the model, and how you learnWritten notice before any base model change

If a vendor cannot fill this in, that is itself the finding. A blank AI bill of materials means you are accepting unknown risk into a government system, and writing "proprietary" in a field does not convert an unknown into a managed risk. Pair the bill of materials with the software attestation the vendor already owes you under existing memoranda, so the AI specific record and the software record are reviewed together rather than by two teams who never compare notes.

Open weights versus closed interfaces

Sandra's choice between a downloadable open model and a closed commercial one is a supply chain decision, not only a cost one. An open weight model gives you control: you host it, inspect it, and run it inside your own secure environment, which matters for sensitive workloads and for classified or air gapped deployments where an external interface is not an option. In exchange you own the entire chain, including verifying provenance yourself and maintaining the model when nobody else will.

A closed model reached over an interface is maintained and patched by the provider, but you cannot inspect it, your data leaves your environment, and you inherit the provider's security posture, outage risk, rate limits and policy changes. You also inherit version drift, because the model behind the interface can change under you. Neither option is automatically safer. The right call depends on the sensitivity of the workload, your data residency obligations, and your honest assessment of whether you can manage the chain yourself. Whichever you choose, treat the provider as a first tier supplier with a documented contingency plan covering model deprecation, policy change, pricing change and geopolitical disruption.

The case record

The consequences of getting supply chain wrong are documented. The SolarWinds Orion compromise disclosed in December 2020 showed that a single trusted software update mechanism could carry an attack into nine federal agencies and roughly one hundred private sector organizations, and it prompted the executive order that reshaped federal software supply chain policy. Log4Shell, tracked as CVE-2021-44228 and disclosed in December 2021, showed how one open source logging component buried in vendor stacks could expose agencies across every mission area, and the Cyber Safety Review Board report issued in July 2022 documented how long it persisted afterwards. A token forgery campaign in 2023, tracked as Storm-0558, reached State Department and Commerce email.

The AI specific cases follow the same shapes. Engineers at a large manufacturer leaked proprietary source code by pasting it into a public chat service, a data exfiltration pattern that requires policy and tooling rather than a vulnerability fix. Public model hubs have had to remove malicious model weights uploaded under names close to popular releases, which is a direct threat to any agency pulling models for fine-tuning. Clearview AI raised the question of whether agencies were using facial recognition capability built on scraped training data of uncertain provenance, with downstream legal exposure under the Illinois Biometric Information Privacy Act. The IRS ID.me rollout in 2022 surfaced questions about the vendor's subcontractor chain, labor practices at verification centers, and whether the agency could meaningfully audit the end to end identity pipeline.

The European cases make the accountability point most sharply. The Dutch childcare benefits scandal, the SyRI risk indication system, and the United Kingdom's Post Office Horizon matter all showed the same failure: accountability collapses when an institution cannot trace which component produced which output. That is the deepest reason to care about supply chain provenance in government. It is not only that a compromised component might harm you. It is that when a decision is challenged, an agency that cannot say where the output came from cannot defend it, cannot correct it at the source, and cannot promise it will not happen again.

Concentration risk and continuity

Vendor concentration is a supply chain risk that no individual project owner can see. If several initiatives depend on the same provider, a single outage, acquisition, policy change or price change becomes an agency wide event. The mitigations are unglamorous: multi vendor strategies where the workload allows, fallback pathways designed before they are needed, contract termination plans written while the relationship is good, insurance considerations, and an agreed coordinated disclosure route for when a vendor component fails.

Exercise the failure rather than documenting it. Table top the loss of a primary foundation model provider, the compromise of a training dataset, the discovery of a poisoned open source dependency, and a zero day in an AI inference framework. Each of those exercises tends to surface the same finding, which is that the fallback everyone assumed existed does not, or that it exists on paper and nobody has ever run inference through it. Coordinate the supplier side of this with the security, acquisition and export control functions that already track foreign supplier risk, rather than building an AI specific version of a capability the agency already has.

Building supply chain checks into procurement

The leverage point is the contract, before money changes hands. Federal acquisition regulation already carries supply chain clauses, including the clause implementing the Section 889 prohibition at 52.204-25 and the related prohibition at 52.204-27, and acquisition guidance for software consumers exists to help contracting officers use them. Federal acquisition rules already let you require security documentation and the right to inspect and test. Add three requirements to your AI solicitations.

  1. Require an AI bill of materials as a deliverable, with acceptance tied to its completeness rather than to its existence. A form submitted with three fields populated is not a deliverable.
  2. Require the right to independently evaluate the model on your own data before final acceptance, and preserve that right for the life of the contract rather than spending it once at award.
  3. Require ongoing notification. If the vendor swaps the base model, retrains it, or changes the hosting environment, you get told and you re-evaluate, because a model update is a new supply chain wearing the old contract number.

Sandra added all three to her next solicitation. The contractor who shrugged the first time arrived with a completed bill of materials, because the contract had made it the price of doing business. That is the whole mechanism: supply chain transparency is not something vendors volunteer and it is not something a security review can extract after award. It is something the solicitation buys.

The operating checklist for whoever owns this

For the official accountable for AI across an agency, the operational implications are concrete and they form a standing checklist. Maintain an AI vendor and component inventory that crosses the software and AI boundary, holding software bills of materials, AI bills of materials, model cards, dataset documentation and whatever training data provenance is available. Apply cloud authorization and system security boundaries consistently to AI workloads rather than treating an inference endpoint as somebody else's problem. Use contracting vehicles that preserve visibility, and write the supply chain clauses in rather than negotiating them later.

Then treat model providers as first tier suppliers with real contingency plans, coordinate with the security, acquisition and export control bodies that already own foreign supplier risk, exercise incident response against supply chain scenarios rather than only against intrusions, and document all of it in a form that survives personnel turnover. That last item is the one that gets skipped and the one auditors reach for first. The point of the documentation is not the audit; it is that the person who inherits Sandra's program later can answer the junior analyst's question without starting over. Agencies that treat AI supply chain as a procurement checklist lose. Agencies that treat it as an ongoing risk management program integrated with their existing supply chain risk discipline win.

Anti-patterns

  • License file as provenance. Accepting a license document as evidence of origin. A license tells you the terms of use, not who trained the model, on what, or whether the file you hold is the file the publisher shipped.
  • Checksum theater. Recording that a fingerprint exists without ever verifying it against the publisher, or verifying it once at award and never again after an update.
  • Unknown as an answer. Allowing "proprietary" or "unknown" in a training data field and treating the form as complete. An unfilled field is a finding, not a formatting issue.
  • Benchmark as evaluation. Accepting a vendor's public leaderboard score as evidence of fitness for your mission, instead of testing on your own data and your own threats.
  • One time due diligence. Reviewing the supply chain at award and never again, so a base model swap or a retrain silently replaces the system you approved.
  • Boundary gaps. Authorizing the application while the inference endpoint that actually processes the data sits outside the authorization boundary.
  • Concentration blindness. Every program independently choosing the same provider, so no single decision maker ever sees that one supplier now underpins the agency.
  • Fallback on paper. A documented alternative provider that nobody has ever run a workload through, discovered to be non functional on the day it is needed.
  • Separate reviews. Software attestation reviewed by one team and AI documentation by another, so nobody ever compares the two records for the same system.
  • Provenance without traceability. Collecting component documentation that cannot be tied back to a specific output, which is precisely the failure that makes a challenged decision indefensible.

Practice prompts

  1. Pick one deployed AI system and try to answer the junior analyst's question in writing: who trained the model, on what data, and how do you know the artifact is unmodified. Note where you run out of evidence.
  2. Draft the AI bill of materials fields you would require in your next solicitation, and mark which ones your current vendors could fill in today.
  3. List every AI initiative in your organization and the model provider behind each. Count how many stop working if one provider does.
  4. Take one system and locate its inference endpoint. Confirm whether that endpoint sits inside the same authorization boundary as the application, or somebody assumed it did.
  5. Run a table top on losing your primary provider at short notice. Write down what breaks, who decides, and how long the fallback actually takes.
  6. Read the license of one model you have deployed and confirm that it permits the use you are making of it. This exercise ends in a surprise more often than people expect.

Reflection

Sandra's contractor was not hiding anything. They genuinely did not know, because nobody in the chain had ever been asked, and the question had never been a condition of getting paid. That is what makes supply chain risk different from most security problems: the information usually exists somewhere upstream, and it goes uncollected because no contract required it and no reviewer asked. Ask yourself which link in your own chain you would be unable to describe if a committee asked tomorrow, and then ask whether that gap is technical or simply the fact that nobody has ever had to answer. The second kind is far more common, and it is the kind a solicitation can close.

Glossary

  • AI supply chain risk. The risk that an organization cannot vouch for where its AI capability came from, what went into it, who controls its updates, and what happens when a link fails.
  • Provenance. The documented origin and history of a model or dataset, including publisher, version and integrity evidence.
  • Model lineage. The record of what a model was derived from, including base model, fine-tuning steps and the data used at each stage.
  • Open weight model. A model whose internal parameters are publicly downloadable, so it can be hosted and inspected locally.
  • Software bill of materials. A machine readable inventory of the components in a piece of software, exchanged in standard formats.
  • AI bill of materials. The same idea extended to an AI system, covering base model, training and fine-tuning data, dependencies, hosting and evaluation.
  • Data poisoning. Deliberate contamination of training data so the model behaves as the attacker chooses on chosen inputs.
  • Version drift. Change in a served model's behavior when the provider updates it, without any change on your side.
  • Concentration risk. Exposure created when multiple initiatives depend on the same supplier, so one event affects all of them at once.
  • Dependency confusion. An attack that gets a build system to pull a malicious package in place of the intended internal one.
  • Data residency. The requirement that data be stored and processed within a defined jurisdiction or environment.

Closing

The junior analyst asked the best question in the room, and the fact that it came at a final security review rather than at solicitation is the whole problem in miniature. By the time a model is about to be fielded, every cheap answer has already been foreclosed: the contract is signed, the schedule is committed, and the only remaining options are to accept unknown risk or to blow up a program. Move the question to the front, where it costs a paragraph in a solicitation instead of a program review, and the same contractor who shrugged will arrive with the documentation. Supply chain assurance in government is not primarily a technical capability. It is a habit of asking early and writing the answer into the contract.

Key takeaways

  • You inherit every link in the chain. The foundation model, its training data, bundled libraries, the hosting environment and the accelerators all carry risk into your system, including the parts you never see.
  • AI adds three opaque links. A black box model, invisible training data and probabilistic behavior make AI harder to inspect than ordinary software, and finite testing cannot close that gap.
  • Verify provenance before deployment. Know the exact model and version, confirm a cryptographic fingerprint, and read the license for whether it permits your use.
  • Treat data poisoning as a real threat. Demand documentation of training and fine-tuning data, because unknown data is unmanaged risk and for sensitive missions it is a security exposure.
  • Require an AI bill of materials. It is your single best artifact, and a vendor who cannot produce one has just handed you the finding.
  • You are already inside a policy regime. Software supply chain executive orders, attestation memoranda, supply chain risk management publications and AI specific third party requirements all apply to the same system.
  • Open versus closed is a supply chain choice. Open weights give control and inspection with full ownership of the chain; a closed interface offloads maintenance while hiding the model, exporting your data and exposing you to version drift.
  • Watch concentration, and exercise the fallback. A documented alternative nobody has ever run a workload through is not a fallback.
  • Traceability is what makes a decision defensible. The European failures show that when an institution cannot say which component produced which output, accountability collapses regardless of intent.
  • Put the checks in the contract. Mandate the bill of materials, independent evaluation rights and notice of any model change as conditions of acceptance, because none of these can be extracted after award.

Frequently Asked Questions

Is an open weight model more or less secure than a commercial interface?

Neither, and the question hides the real trade. Open weights give you inspection, local hosting and independence from a provider's policy changes, in exchange for owning provenance verification, patching and maintenance yourself. A closed interface gives you a maintained system in exchange for opacity, data leaving your environment, and drift you do not control. Decide on the sensitivity of the workload and on whether your organization can genuinely manage a chain it owns, not on which option sounds safer in the abstract.

What do we do when a vendor calls the training data proprietary?

Record it as a gap rather than as an answer, and decide explicitly whether you are accepting it. Sometimes the honest resolution is a summary rather than the dataset itself: sources at a category level, known limitations, and what controls existed against poisoning. What you should not do is let the field be marked complete because something was written in it. If the gap is unacceptable for the mission, that is a source selection outcome, and it is much cheaper to reach it before award.

Does a checksum actually prove the model is safe?

No. It proves that the file you hold matches the file the publisher published, which is a narrow and genuinely useful fact. It says nothing about whether the publisher's model was trained on poisoned data, whether it contains a backdoor, or whether it behaves acceptably on your mission. Integrity verification and behavioral evaluation are separate controls answering separate questions, and a program that does the first and skips the second has verified that it faithfully downloaded something it has never tested.

Our vendor updated the model. Do we have to re-evaluate?

Yes, and the reason is that a model update is a new supply chain arriving under the old contract number. New weights mean new behavior, potentially new training data and a new evaluation baseline, none of which your original testing covers. This is why written notification of any base model change belongs in the contract, along with your right to re-evaluate before the change reaches production. Version drift on a served interface is the same problem in a form that is easier to miss.

We are a small agency without a supply chain security team. Where do we start?

Start with the inventory, because you cannot manage what you have not listed. Write down every AI system, its model and version, its provider, where inference runs, and who else in the organization depends on the same provider. That single table usually surfaces the findings worth acting on immediately, most often an unauthorized inference path or a concentration nobody had noticed. The bill of materials requirement in your next solicitation is the second step, and it costs a paragraph.

How does supply chain risk connect to the systems we already have authorized?

Directly, because AI components sit inside boundaries that already exist. A model processing federal data belongs in an authorized environment at the appropriate impact level, its inference endpoint has to be inside that boundary, and the supply chain control family in existing federal security controls already covers component provenance and third party risk. The practical failure is not usually an absent requirement; it is an AI system procured as a service by a program office that never went through the boundary conversation at all.