Third-Party and Concentration Risk at Enterprise Scale
The audit and risk committee has forty minutes for AI, and the first thirty go well. The slide shows fourteen AI-touched processes, each with a named vendor, a contract value, a tier, and a green status. Procurement has done its job: every supplier assessed, every clause negotiated, every review dated. Then a non-executive director who spent twenty years in automotive supply chain asks a question that is not on the slide. "How many of these vendors ultimately depend on the same model provider?" The chief information officer says the vendors are independent. The head of procurement says the contracts are separate. The director says, mildly, that she did not ask about contracts. The honest answer, produced eleven days later after somebody actually asked each supplier, was nine of the fourteen. Four logos, one dependency, and a risk register with no column capable of showing it.
The Question the Vendor Register Cannot Answer
You already know how to run vendor risk as a program. Level 4 built it: tiering by what the vendor touches, a dossier per supplier, contract clauses that carry the controls, continuous monitoring, scheduled review, tested exit readiness, and a concentration ledger with tolerances. That program is correct, nothing here replaces it, and an enterprise without it should build it first.
The problem is structural rather than a failure of execution. A vendor register is organized by counterparty: its unit of analysis is the contract, so every column answers one question, who do we buy from and how well are we managing them. That is legitimate, but it is not what a board asks about AI risk. The board is asking: what do we actually rest on, and what happens to us if it moves?
Those questions now have different answers, because the thing you buy and the thing you depend on are no longer the same object. Buy enterprise software in the traditional sense and the vendor is the dependency: the code runs, it behaves as it did last month, and the supplier's obligations are your whole exposure. Buy an AI-enabled service and you buy a position at the top of a stack whose lower layers you did not select, cannot inspect, and were never told about. Your supplier's supplier's supplier is inside your month-end close, and you learned their name, if you ever did, from a footnote.
Every board already understands this shape. It is counterparty concentration, and single-source component risk: three qualified suppliers of a part all buying from the same smelter, so a fire at a facility you never heard of stops three supposedly redundant lines at once. Banks know it as correlated exposure, a loan book diversified by borrower and concentrated by the one industry all those borrowers sell into. The AI version is the same failure of aggregation, dressed in unfamiliar vocabulary.
That familiarity is the strategic opportunity. Presented as a technology worry, AI dependency gets delegated to the people who understand it, and delegation begins an unmanaged exposure. Presented as concentration risk, it is a category the board has governed for decades. Your job is translation: make the AI stack legible in the board's own language, or be the person explaining after an incident why nobody had counted.
The market gives one further reason not to trust the register. Gartner's 2025 assessment found agent washing widespread: of the thousands of vendors claiming agentic capability, only around 130 were judged to be offering anything real, and over 40 percent of agentic AI projects are expected to be canceled by the end of 2027. Read it as a supplier durability finding: many rows on your register are a well-designed interface over somebody else's model, which makes the logo the least durable thing in the arrangement.
The Stack Beneath the Vendor
Before you can map dependency you need a vocabulary for what sits under an AI-enabled process. Four layers is enough; more granularity is unhelpful at board level.
The application layer is the tool your users touch: the contract-review assistant, the claims triage workbench, the service-desk copilot. It contributes the interface, the workflow logic, the integrations, and the configuration your team spent four months tuning. Its quiet failure mode is the dangerous one: a vendor retires a feature your standard operating procedure (SOP, the documented step-by-step for how work is done) depends on, and your process breaks where no dashboard is looking.
The platform layer is the environment the application runs in: the cloud AI service, the orchestration platform, the vector database. It contributes availability, scaling, security boundaries, and increasingly the model routing itself. Its pricing and interface changes arrive on the platform's timetable rather than yours.
The model provider layer is the underlying model doing the reasoning, and it contributes the capability that made anyone want the tool. It is also the least visible layer and the most correlated one, which is the point here. It fails in five ways worth naming separately because they need different controls: version deprecation, where the model you validated is retired on a published schedule; behavior change, where an updated model produces different outputs on the same inputs with no announcement you see; price restructuring, which arrives across every customer at once; capacity limits during demand surges; and terms changes covering data use, retention, sub-processors, or permitted applications. Below any two apparently unrelated AI vendors, the chance this layer is shared is high, because the number of frontier providers is small.
The infrastructure layer is compute and hosting: chips, data centers, regions. It matters through two questions: where is our data physically processed, a compliance question with a jurisdiction attached, and what happens to our service level when demand for compute spikes.
Visibility falls as you descend: the application is on your invoice, the platform is often named and rarely governed by you, the model provider is undisclosed unless you ask, and the infrastructure sits below anything you signed. Exposure runs the other way: the layers you see least are the likeliest to be shared across suppliers you believe are unrelated.
The Discovery Move Most Enterprises Have Never Made
Here is the highest-yield action in this lesson, and it costs almost nothing. Ask every AI vendor two questions in writing, with a named recipient and a date.
- Which models and platforms does your service depend on today? Name the provider, the model family, and the hosting platform. If you route across more than one, say which and on what basis.
- Will you commit to notifying us of changes at those layers? Model version changes, provider changes, material behavior changes, sub-processor changes, and data handling changes, with a stated notice period.
Run that across twenty vendors and you get four kinds of answer, all useful. Some disclose immediately, often with documentation they were surprised nobody had asked for. Some disclose only when pushed, which means no notification process exists. Some route dynamically across several providers and say so with pride: sometimes that is resilience, and it also means behavior can change when routing changes, with no version number moving. And some decline, citing confidentiality. That refusal is a finding: a supplier who will not say what their service rests on has told you the dependency is unmanageable, which belongs in the tier assessment, the renewal conversation, and the risk map.
None of this can be inferred from the register, the security questionnaire, or the SOC 2 report (the standard audit of a supplier's security controls, which says how carefully they operate and nothing about what they are built on). It can only be asked. MIT's 2025 finding that externally partnered AI solutions succeed roughly twice as often as internal builds is a strong argument for buying rather than building; this lesson is its fine print. Buying moves the capability outside your walls, and its price is that you must know what sits behind the wall.
The Four Exposure Types, and How to Measure Each
Once you can see the stack, concentration stops being one worry and becomes four, each with its own measurement and board framing. Treating them as one blurred concern defers the topic; separating them governs it.
Operational continuity: what stops
If a layer becomes unavailable, what work stops, for how long, and what does carrying on manually cost per day? Ask each process owner one question: if this capability disappeared at nine on Monday morning, what is the fallback, how many extra hours per day does it consume, and what service impact do customers see?
Most organizations answer with an estimate delivered in a meeting. An organization that took Level 3 seriously answers with evidence, because the fallback section of every redesigned SOP already contains it: the continuity position is the sum of fallbacks written, costed, and rehearsed process by process. If they have been drilled, you can state your exposure in hours and dollars. If they were never exercised, you have documents rather than a position, and the difference shows up under load. If they do not exist, your honest answer is that you do not know, which is legitimate exactly once. The board framing is business interruption: how many days of degraded operation the enterprise can absorb, and at what cost.
Economic exposure: a repricing that arrives everywhere at once
This is where per-contract analysis fails most visibly. A 40 percent increase on a 30,000 dollar tool is an annoyance a category manager absorbs without escalating. Four of them in one quarter, because they share a cost driver nobody mapped, is a budget event.
Measure it the way the shock arrives: take your most concentrated layer, assume a repricing of a specific size, and total the impact across every dependent process at once. That is a correlated shock stress test, a board-standard instrument elsewhere, since banks stress a rate move and manufacturers a commodity move without anyone calling it alarmist. Framing AI cost this way moves it out of the IT budget variance conversation and into the risk conversation. The purpose is not to predict a price rise; it is to know, before the letter arrives, whether your answer is a shrug or a crisis.
Behavioral exposure: the one with no traditional analogue
The third exposure is genuinely new, and it is why this lesson exists rather than being a paragraph in a procurement policy. A model provider updates a model. Several of your vendors inherit the change on their own schedules, and multiple processes begin performing differently at roughly the same time. Nothing has broken and no invoice changed. The outputs are simply not what they were.
Level 3 taught drift detection at the process level: a canary set of representative cases with known-good outputs, re-run on a schedule, so quality movement is caught by a control rather than a customer complaint. At enterprise scale that control has a blind spot: each process monitors itself. A four-point quality dip reads as noise, because usually it is noise, whether sampling variation, a seasonal input mix, or a new team member, and the monitor is behaving correctly when it stays quiet. Meanwhile the same dip is occurring in three other processes, each also correctly ignoring it, and the common cause is invisible precisely because the controls are decentralized and doing their jobs.
The enterprise-level control is a cross-process canary deck: one evaluation set spanning processes that share a model provider, run on one cadence, reported together and tagged by provider rather than by owning function. Built from material you already have, its cost is coordination rather than construction. Take twenty to thirty cases from each participating process with their agreed acceptable outputs, run them fortnightly, and plot the results on one chart grouped by provider. Set the alert on the pattern rather than the point: simultaneous movement across two or more processes sharing a provider is a signal even when no single movement would have crossed its own threshold.
That inversion runs against ordinary monitoring instinct: small movements correlated in time on a shared dependency are more diagnostic than one large movement in isolation, because the first names a common cause and the second is usually local. Nothing else in your control environment looks for the first kind. Ownership belongs to the enterprise AI function, since a control reporting into a process it monitors loses its cadence within two quarters, and the deck needs a named escalation path.
Contractual and compliance exposure: whether your terms reach the layer that changed
Terms change at layers you do not contract with. A model provider revises its data-use policy, adds a sub-processor, changes retention, or shifts processing to a new region. Your obligations to customers and regulators did not change. The question is whether your contractual protections travel down the stack to the layer where the change happened.
Level 4 gave you the clause set: data handling and residency, model change notification, audit rights, service levels, liability, and exit assistance. Measuring this exposure has two parts. First, clause coverage: which contracts carry them. Second, and almost nobody has done this part, flow-down: does your vendor's contract with their model provider oblige anything equivalent, and does yours require them to pass obligations down and notifications back up?
Ask it plainly: our contract binds you; does anything bind the provider beneath you, and will you tell us when they change something? Some vendors have negotiated genuine enterprise terms and will show you the commitments. Some are on self-service terms that oblige their provider to nothing and can change with thirty days of notice posted on a web page. Under the EU AI Act, whose general-purpose AI (GPAI, the broad foundation models sold for many uses) obligations have applied since 2 August 2025 and whose AI-content transparency requirements arrive on 2 December 2026, "we did not know what was underneath" is not a defense that improves with repetition. For processes handling personally identifiable information (PII, data that can identify a specific person), an undisclosed sub-processor change two layers down is a live compliance issue however well ordered your own house is.
The Artifact: The Dependency Stack Map
The artifact turns four abstract exposures into a governable position. The Dependency Stack Map traces every AI-touched process through its layers to its providers, aggregates exposure by layer rather than by contract, names the correlated-failure scenarios, and states your mitigation position.
One row per AI-touched process, columns in three groups. The trace: process, owner, criticality, application vendor, platform, model provider, infrastructure region, and a disclosure status (disclosed, disclosed on request, refused, unknown). The exposure: annual run cost, manual fallback, whether it has been rehearsed and when, fallback cost per day, clause coverage, flow-down status, canary deck membership. The position: for each concentrated layer, the mitigation stance, the tolerance threshold, the exit estimate with its date, and a named owner.
The map is not the deliverable; the aggregation is. Sort by model provider and count processes, run cost, and criticality. That view, which no vendor register can produce, turns a list of suppliers into a statement of enterprise exposure.
A Worked Example: Fourteen Processes, One Provider
The numbers below are hypothetical, realistic for a mid-sized enterprise and chosen to show the arithmetic. Use the shape, not the figures.
The transformation team maps 14 AI-touched processes across finance operations, customer service, procurement, legal, and marketing. Tracing the application and platform layers takes two days, because that is on the contracts; tracing the model provider layer takes three weeks. The finding: 9 of the 14 rest on the same model provider, through four application vendors that look entirely independent in the register, hold separate contracts, and are managed by three different category managers. Two disclosed only when asked directly, one had it in public documentation nobody had read, and one declined on confidentiality grounds. The refusal is recorded as a finding, flagged for a renewal eight months out, and raised in the governance forum.
- Operational continuity. Aggregating the fallback sections of the nine affected SOPs gives an estimated 34,000 dollars per full day of provider unavailability in additional labor alone, plus two customer-facing queues that would breach published response commitments inside 48 hours. It also finds 3 of the 9 have fallbacks written and never rehearsed, the more valuable finding: those three are drilled within the quarter, and one turns out to reference a report decommissioned last year.
- Economic exposure. A single 40 percent repricing at the shared provider layer, applied across all nine dependent processes at once, models out at an estimated 96,000 dollars of additional annual cost. Each of those nine contracts had been reviewed individually, and every review concluded the price exposure was immaterial: correct at the contract level, wrong at the enterprise level, which is the aggregation failure the map exists to catch.
- Behavioral exposure. A cross-process canary deck is stood up for the shared provider: 25 cases from each of the nine processes, run fortnightly, on one chart. Three months later it earns its cost in one event. A provider model update produces a simultaneous shift in two processes, roughly five points and four points, neither of which would have crossed its own threshold and both of which would have been logged as noise. Because they moved together, in the same week, on the same provider, the pattern is escalated within two days, the vendor confirms the update, and re-calibration happens on the enterprise's schedule rather than after a complaint.
- Contractual exposure. Clause coverage is good on six of the nine contracts and thin on three. Flow-down is worse: only two of the four vendors can point to any obligation binding their model provider, and one is on standard terms permitting change with 30 days of posted notice. Two renewals are re-scoped, and a model change notification clause enters the AI contract template.
The board decision at the next risk committee: accept the concentration deliberately. The rationale is documented, the exposure quantified, tolerance thresholds set (no more than 60 percent of AI-touched processes and no more than two tier-one processes on one provider without an explicit board review), exit estimates refreshed annually, and one alternative assessed but not deployed. The committee spends eighteen minutes on it, agrees, and moves on. That eighteen minutes is what the exercise was for.
The Failure Story: The Invisible Common Cause
Now the version where nobody drew the map. A professional services group runs five AI-supported workflows, bought from four vendors over three years by three budget holders: proposal drafting, engagement scoping, timesheet narratives, research summarization, and clause extraction. Each was procured properly, with an owner, a business case, and a monitoring routine. The register shows four healthy suppliers and no concentration flag, because it has no concept of a layer.
A model provider deprecates a version. This is entirely routine and entirely announced, with notice to its direct customers on a published schedule months ahead. Those customers are the four vendors, not the services group. Three of the four pass the change into their release notes, and none contact their customers with any urgency, because from where each vendor sits it is a maintenance item, not an incident.
Over the following five weeks, three of the five workflows shift behavior. Proposal drafts get longer and blander, and reviewers add a rewriting pass nobody logs. Clause extraction starts missing a category of indemnity language it had previously caught, noticed by a lawyer who assumed she was seeing a bad week. Research summaries change tone enough that two account leads quietly stop using them. Each owner investigates independently and reasonably. Two conclude the problem is their own input quality, the usual cause of this symptom, and their vendors report no incidents. One raises a ticket that returns nine days later asking for example inputs. Nobody looks sideways, because there is no forum in which looking sideways is anybody's job. The common cause surfaces in week six, when an analyst notices three quality lines bending in the same fortnight.
The direct cost is modest: a few hundred hours of rework, one embarrassing proposal, no client lost. The incident review's finding is not. The enterprise had no way to know that five workflows were three dependencies, no monitoring across processes, and no contractual right to notification at the layer where the change originated. Every individual control worked; the aggregate control did not exist. The risk was never that the provider changed something, since providers change things constantly, competently, and with notice. The risk was that the enterprise could not see its own dependency structure, so a routine, announced change arrived as an unexplained five-week quality mystery.
The Four Mitigation Decisions, With Honest Costs
Having measured the exposure, you choose a position. Four are available, and the discipline is to treat them as decisions with costs rather than best practices with halos.
Accept is the most common answer and frequently the correct one: keep the concentration, price the exit, monitor the position, and write down why. It is not passivity but a decision with an owner and a date, available only to an enterprise that measured what it accepts. Chosen concentration has a quantified exposure, a documented rationale (capability, cost, coherence, integration effort already spent), thresholds that trigger a review, and a refreshed exit estimate. Accidental concentration has none of those and looks identical right up until it does not: opposite conditions in governance terms, even though the technical position is the same.
Diversify means deliberately spreading dependency across providers: real, effective, and more expensive than it looks on a slide. The costs are cumulative: more contracts to monitor, more integrations to maintain, more variants of every SOP and training package, more evaluation work because each variant needs its own baseline, and usually worse terms from every supplier because you split the volume that was buying your leverage. The rule: diversify where the exposure is existential or where switching capability is itself the asset. A process whose failure would stop revenue, breach a regulatory obligation, or make the front page justifies duplicate cost as insurance. Routine capabilities do not; diversifying the marketing copy assistant is a cost with a story attached rather than a control.
Abstract means designing so a layer can be swapped: an internal interface between your processes and the model, provider-neutral prompts and configuration in your own repository, evaluation sets you own rather than ones embedded in a vendor's console, and schemas that do not encode one supplier's assumptions. Honesty matters more here than anywhere else in this lesson, because portability is the most over-promised property in the AI market. Abstraction genuinely buys schema-level independence, prompt and configuration portability so months of tuning are not stranded, evaluation sets that transfer so a replacement can be tested against the cases that defined your current quality, and a shorter migration. What it does not buy is behavioral equivalence: two models given identical inputs produce different outputs with different failure modes, so a swap requires re-testing, re-calibration, and usually some SOP adjustment, every time, however clean the interface. Abstraction reduces switching cost; it does not eliminate switching risk, and anyone who calls a model swap a configuration change is using marketing language.
Maintain optionality is the cheapest mitigation and the most neglected: keep exit estimates current, keep evaluation sets portable and owned by you, keep one alternative assessed and roughly priced even with no intention of deploying it, and re-run the exit test annually so the estimate is evidence rather than folklore. Optionality is a discipline rather than an architecture, and it degrades silently: the estimate ages, the alternative you assessed two years ago repriced, the evaluation set drifts, and the person who knew the migration moved on. Nothing visible breaks, so you discover the decay on the day you need the option. Put the annual refresh in the governance calendar with a named owner and treat a stale exit estimate as an open finding.
| Decision | What it costs | When it is right |
|---|---|---|
| Accept | Ongoing measurement and honest board reporting | Routine capabilities where exposure is quantified, priced, and tolerable |
| Diversify | Contracts, integrations, training, variants, weaker terms | Existential exposure, or where switching capability is itself strategic |
| Abstract | Design effort, internal interfaces, owned evaluation sets | Layers you expect to change; reduces switching cost, not switching risk |
| Maintain optionality | A few days a year, and the discipline to keep doing it | Almost always, alongside whichever of the other three you chose |
Taking the Position to the Board
Presentation is not cosmetic: the same position produces different governance outcomes depending on how it arrives. Bring five things, in order. The exposure map by layer: one page, processes aggregated by provider, with criticality and run cost. The correlated scenarios: two or three, quantified, stated as assumptions rather than forecasts. The current position with its rationale, in commercial terms. The tolerance thresholds that bring this back to the board without anyone having to remember. Your recommendation, with its cost. Then the framing sentence that makes it land, said in one breath: we have chosen this concentration, here is what it saves us, here is what it would cost us in the scenarios we consider plausible, and here is what we monitor.
Every clause does work. "We have chosen" establishes agency, the difference between a managed risk and an accident waiting for an owner. "What it saves us" makes the concentration a decision with a benefit. "What we monitor" is the control that lets the board stop thinking about this until something changes.
A board that hears a chosen concentration governs it. A board that discovers an unmeasured one intervenes.
That asymmetry is organizational physics. Boards are not risk-averse; they are surprise-averse. A director given a quantified exposure and a deliberate position asks two or three sharp questions, tests the thresholds, and delegates. A director who discovers in an incident review that nobody had counted will not delegate anything for a year. Given that S&P Global found 42 percent of companies scrapped most of their AI initiatives in 2025, and MIT found 95 percent of pilots delivering no measurable return, the board's default posture is already tighter than you would like.
Present this not as a warning about AI but as a standing position in a familiar risk category, reviewed on a cadence like any other. With dependencies mapped and deliberately governed, one question remains: when someone asks you to prove any of it, what exactly do you hand them? That is the evidence layer, and it closes the chapter.
What to Do Monday Morning
- Send the two disclosure questions to every AI vendor in your register. Which models and platforms do you depend on, and will you notify us of changes at those layers? Give a date for the reply, and record refusals as findings rather than blanks.
- Build the Dependency Stack Map. One row per AI-touched process, traced through application, platform, model provider, and infrastructure. The first two layers come off your contracts in a day; the third is the work and the whole point.
- Aggregate by layer instead of by contract. Sort by model provider and total the processes, run cost, and criticality. If more than half land on one provider, you have found the number your board will ask for.
- Model one correlated repricing. Take your most concentrated layer, assume a stated increase, and apply it to every dependent process at once. Compare that total against the contract reviews that each concluded the exposure was immaterial.
- Stand up a cross-process canary deck for your most-shared provider. Twenty to thirty cases per process, a fixed cadence, one chart, alerts on correlated movement rather than any single threshold, owned by the enterprise AI function with a named escalation path.
- Take a deliberate accept-or-diversify decision to your board. Exposure map, scenarios, position, thresholds, recommendation, in that order. A position you chose and can defend beats one that accumulated while everyone managed vendors correctly, one at a time.
Key Takeaways
- Separate the vendor question from the dependency question: a register organized by contract answers who you buy from, while the board is asking what you rest on.
- Trace every AI-touched process through four layers (application, platform, model provider, infrastructure), remembering that the model provider layer is the least visible and most correlated across suppliers who look independent.
- Ask every AI vendor which models and platforms they depend on and whether they will notify you of changes, treating a refusal as a recorded finding, because none of this can be inferred from a contract or a security report.
- Measure four exposures rather than one blurred worry: operational continuity from aggregated SOP fallbacks, economic exposure from a correlated repricing, behavioral exposure from shared model changes, and contractual exposure from clause coverage and flow-down.
- Build a cross-process canary deck for your most-shared provider and alert on correlated movement, because small shifts occurring together on a shared dependency are more diagnostic than one large shift in isolation.
- Ask the flow-down question plainly, since your contract binds your vendor and whether anything binds the provider beneath them is a specific, checkable fact most enterprises have never established.
- Choose a mitigation position with its honest cost: accept deliberately with the exposure quantified, diversify only where exposure is existential, abstract knowing it reduces switching cost but never removes re-testing, and maintain optionality as a discipline that decays silently if unexercised.
- Present the position in one sentence the board can govern: we chose this concentration, here is what it saves us, here is what it would cost us in plausible scenarios, and here is what we monitor.
Skill.re