←
AI Readiness & Process Transformation
Strategic · M25 · lesson 25 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Vendor and Third-Party Risk Management

15 min

The email arrives on a Tuesday and it is polite. A large client of the professional-services firm is running its own third-party audit, and one of its questions is routine: please confirm the current list of sub-processors handling client material inside your AI summarisation tool, and the retention period applied to uploaded documents. The partner who receives it forwards it to the person who owns the vendor relationship, who forwards it to the person who ran the security review at purchase, who no longer works there. Someone finds the review in a shared drive. It is magnificent: 140 questions, answered in detail, signed off by two functions, dated twenty-two months ago. It describes a vendor that no longer exists in that form. The company was acquired eleven months back, the service migrated to the acquirer's infrastructure in a different jurisdiction, the sub-processor list changed twice, and the default retention setting was quietly reset from thirty days to indefinite. None of it was hidden; all of it was published, in release notes and a portal announcement. None of it was read, because the review was an event, and events do not have successors.

The Questionnaire That Expired

Finish the story, because the cost matters. The firm cannot answer truthfully within the seven days the client allows. It answers approximately, discovers the approximation is wrong, and corrects itself. What follows: a quarter of one general counsel's time reconstructing the vendor's actual configuration, a data-deletion exercise against twenty-two months of uploads, a note to the risk committee explaining how a firm that sells rigour to clients did not apply it to itself, and a client relationship that survives but is never quite as warm again.

Notice what did not go wrong. Nobody skipped a step; the 140-question review was better than most enterprises manage. The failure is structural: the organisation treated vendor risk as an event (a questionnaire at purchase, then silence until renewal or incident) and bought a counterparty that changes continuously. Point-in-time diligence against a continuously-changing counterparty is theatre. It produces a document, a sign-off, and a feeling of control that decays at a rate nobody measures.

Diligence has a shelf life, and the shelf life of AI vendor diligence is roughly one release cycle.

This lesson turns vendor risk from an event into a standing programme: diligence at entry, monitoring during, review on a calendar, and exit readiness maintained rather than discovered. Third-party risk management (TPRM: governing the risks your suppliers carry on your behalf) is not new, and your organisation has a version of it. What is new is that the AI vendor breaks the assumptions it was built on, in three specific ways.

Break one: the product changes without a release you approved

Traditional enterprise software changes on a schedule you can see: a version number, a release note, an upgrade you consent to, often a window to stay on the old version while you test. An AI service frequently has none of that. The model is updated, a prompt template behind the interface is tuned, a retrieval component is swapped, and the outputs shift. Users notice before governance does, if anyone notices at all. Level 3's silent-update incident is the origin story: a tool whose output distribution moved overnight, nine days of confused investigation, and a root cause that was the vendor's routine improvement. The traditional questionnaire has no field for "how will you tell us when the product's behaviour changes," because traditional products did not need one.

Break two: your data lives inside their system in ways the old questionnaire does not probe

A classic software-as-a-service (SaaS) security review asks where data is stored, who can access it, and how it is encrypted. Necessary, insufficient. The AI-specific questions sit a layer deeper. Is our content used to train or improve the vendor's models, and does that restriction bind their own suppliers? Which sub-processors touch the payload (model provider, vector store, observability platform)? Are prompts and outputs retained as logs, for how long, readable by whom during support? Does the deletion promise reach derived artefacts (embeddings, caches, fine-tuning sets) or only the original file? Personally identifiable information (PII: data that identifies a specific human) can persist in an embedding index long after the source document is deleted, and a questionnaire written for a file-storage product will never ask about it.

Break three: the market itself is unstable

Gartner's 2025 review of the agentic AI market found that of thousands of vendors claiming agentic capability, only around 130 were judged real: the rest were what the analysts called agent washing: conventional automation or chat interfaces relabelled for the moment's vocabulary. The same research forecast that over 40 percent of agentic AI projects will be cancelled by the end of 2027. Read those findings as a base rate for your supplier list rather than as commentary on other people's mistakes. A meaningful share of the AI vendors you contract with this year will not exist in three years, or not in this form: acquired, pivoted, sunset, or absorbed into a platform with a different roadmap. That is arithmetic about a young market, and it changes what diligence is for. You are not only asking "is this vendor safe today," you are asking "what happens to us when this vendor changes, and how quickly will we know."

None of this argues against buying. MIT's GenAI Divide research found externally partnered solutions succeed roughly twice as often as internal builds, and Chapter 4.2 turned that into sourcing policy. But partnership advantage is not free: it is a return you earn by managing the partner, available to organisations that govern the relationship and unavailable to those that sign, file, and forget. Add the regulatory clock: EU AI Act general-purpose AI (GPAI) obligations have applied since 2 August 2025, AI-content transparency arrives 2 December 2026, high-risk Annex III obligations 2 December 2027, and embedded Annex I systems 2 August 2028. Each milestone raises the documentation a deployer must produce about the systems it uses, including ones it did not build. A vendor file twenty-two months stale is not just a risk position. Increasingly, it is a compliance gap.

The Vendor Risk Register: Four Stages and a Calendar

Here is the artifact. The Vendor Risk Register is one row per AI vendor, maintained by the strategist, reviewed on a published cadence. It is deliberately narrow: it does not replace procurement's supplier master or security's control catalogue. It is the page that answers, for any AI vendor, the questions an executive or an auditor will actually ask.

ColumnWhat it holdsWhy it earns its place
Tier1, 2, or 3, set by exposureSets diligence depth, review frequency, and evidence rights
Dossier dateWhen the entry diligence file was last refreshedMakes staleness visible; a date is harder to ignore than a feeling
Monitoring signalsWhich signals are watched, and the named human watching themConverts "we keep an eye on them" into an owned duty
Concentration positionPortfolio items carried, annual spend, processes touchedAggregates the single-vendor exposure individual contracts hide
Exit positionExit cost in money and weeks, plus the date last measuredAn exit estimate ages; the measurement date is half the information
Next reviewA date, already in a calendar, with an ownerTurns intention into a meeting that happens

Behind the register sit four stages. Most organisations run the first well, the second partially, the third not at all, and the fourth only when forced.

  1. Entry diligence. Assemble a dossier proportionate to exposure and keep it as a living file rather than an artefact of the purchase.
  2. Contract as control. Convert the risks you found into clauses, before negotiation starts.
  3. Continuous monitoring. Watch a few cheap signals, each with a named recipient.
  4. Scheduled review and exit readiness. Re-open the file on a calendar, re-price the exit, and test that the exit actually works.

Stage one, first move: tier by exposure, not by spend

Procurement's default tiering is spend. Reasonable for office furniture, dangerous for AI, because it systematically under-scrutinises the small purchase that matters. A 9,000-a-year transcription tool ingesting recorded customer calls full of account details, health disclosures, and complaint narratives outranks a 240,000-a-year forecasting platform touching nothing but aggregated internal numbers. Spend-based tiering sends the forecasting platform through a full review and lets the transcription tool through on a manager's credit card, because 9,000 sits under the threshold that triggers anything. Every enterprise with a shadow-AI problem built it this way.

Tier on three factors instead, taking the highest applicable answer on each.

  • Data class. The most sensitive category the vendor can see: public and aggregated at one end; customer PII, employee records, health data, privileged legal material, or regulated financial data at the other.
  • Action authority. What the system does without a human in the loop: draft only, draft plus route, or execute (send, post, transact, update a system of record). Level 3's agent-control distinction, applied to the supplier rather than the use case.
  • Process criticality. What happens if the service stops for two weeks: mild inconvenience, degraded service, or a stopped process with customer or regulatory consequences.

High on two of three is tier 1. High on one is tier 2. Low on all three is tier 3. Write the rule down before applying it to a specific vendor, for the same reason Chapter 4.2's sourcing policy is written before anyone knows whose proposal it favours: a tiering rule invented during an argument produces the tier that ends the argument.

Stage one, second move: the dossier, assembled once and kept

The dossier is the entry evidence file. Level 1's twenty-question due-diligence sheet for a single vendor demo and Level 3's seven agent-control interrogation questions both belong in it, and neither is repeated here. What a tier-1 AI vendor additionally requires:

  • The sub-processor list. Who else touches our data. This is the fourth-party problem: your vendor's suppliers are transitively your exposure, and their changes arrive with no notice unless a clause creates one. Ask for the current list in writing, ask how changes are communicated, and read the reissued data processing agreement (DPA: the annex governing how a supplier handles your data) every time. Somebody has to actually read it, which sounds trivial and is the most commonly skipped step in the programme.
  • Model provenance and change policy. Which models are used, whose they are, whether they are fine-tuned on your data or anyone else's, in which regions inference runs, and what happens when their upstream provider deprecates a version. A vendor who cannot answer has not thought about it.
  • Security attestations, with their scope read carefully. An attestation such as SOC 2 Type II or ISO 27001 certification is a real signal, and its scope statement is where the information hides. It is common, and consequential, for an attestation to cover the vendor's corporate systems or legacy product while the AI service you are buying sits outside the audited boundary because it shipped after the audit period. Read the scope section, name the systems it lists, and ask whether your service is inside it. "We are SOC 2 certified" is a company-level statement; you need a service-level one.
  • Incident history and notification commitments. What has gone wrong, when, and how customers found out. A vendor with no incidents in three years is either very good or not counting.
  • Financial viability, tier 1 only. Funding stage, runway if disclosed, customer concentration, ownership. Asking a startup carrying a critical process whether it will exist in two years is not rude, it is business continuity, and mature vendors answer without flinching. An uncomfortable answer is not automatically a reason to walk away; it is a reason to price the exit and keep the estimate current.
  • Bias and evaluation documentation, where the use case touches people. Level 3's fairness lesson set the standard: for anything affecting hiring, credit, pricing, service levels, or access, ask what the vendor tested, on what populations, with what results, how often re-run. A vendor with nothing to show has not necessarily built something unfair; they have definitely built something you cannot defend.

Assemble the dossier once, store it where the next person will find it, and stamp it with a date. That date is the field the register cares about, because it tells you when the questionnaire started expiring.

Stage Two: The Contract Is the Control

Diligence tells you what the risks are. The contract is where you do something about them, and it is the only stage with a deadline you cannot move: at signature, your leverage collapses to whatever the renewal date offers. Late-arriving strategists inherit contracts they would not have signed, then spend two years working around a missing clause that would have cost nothing to include.

You do not write the contract. Legal writes it, procurement negotiates it, and both do it better than you would. Your job is upstream: supply the risk requirements list before negotiation starts, ranked, with the reason attached to each item. Legal cannot ask for a model-change notification window if nobody told them such a thing exists. The highest-return habit in vendor risk is being in the room two weeks earlier than you were invited.

ClauseWhat it says, in operator termsWhat it prevents
Model-change notificationAdvance notice of material model or behaviour changes, with a defined testing window before the change reaches our production tenantThe Level 3 silent update: a behaviour shift found by users, investigated for nine days, explained by a release note
Data-use restrictionOur content is not used to train, fine-tune, or improve any model, stated affirmatively rather than as an opt-out, and binding on sub-processorsYour material becoming someone else's capability, and a client question you cannot answer
Retention and deletion with verificationDefined retention periods, deletion on request within a stated window, extending to derived artefacts, with written confirmationThe quiet default reset: retention drifting to indefinite because nobody re-read the settings page
Export in documented formatsData, configuration, and history exportable in named machine-readable formats, on demand, within a stated turnaroundAn exit estimate built on hope; this clause makes the exit number real
Incident notificationNotice within a defined window (hours, not "promptly") for security, availability, and data incidents affecting usLearning about your vendor's breach from a journalist or a client
Evidence and audit rightsTiered: annual attestation refresh and questionnaire response at tier 2, plus evidence rights and named escalation at tier 1A monitoring programme with no right to ask for anything
Price-change protectionCaps and notice periods on renewal pricing, especially usage-based componentsThe renewal at 2.4x because usage grew and the discount expired

The table is a menu, not a mandate: a tier-3 vendor does not need evidence rights, and asking wastes goodwill you will want later. And the clause you cannot get is itself information: a vendor who refuses model-change notification is telling you their release process cannot support it, which tells you what monitoring must carry alone.

Stages Three and Four: Monitoring, Review, and Exit Readiness

These two stages do not exist in most organisations, and they are where the questionnaire-that-expired failure actually happened. Both are cheap. Neither is difficult. They simply require someone to own them on an ordinary week when nothing is wrong.

The signals worth watching

Monitoring does not mean a dashboard or a tool. It means five signals, each with a named human and a routing rule.

  • The vendor's own change communications. Release notes, product emails, status-page subscriptions, and DPA reissues, routed to a named person, not a shared inbox nobody reads. The cheapest control in the programme, and the one whose absence caused the opening story. Fifteen minutes a fortnight per tier-1 vendor.
  • Your own canary suite. Level 3's regression tests are the most valuable monitoring asset you own, being the only signal that does not depend on the vendor telling you anything. A fixed set of representative inputs run monthly, outputs compared against known-good results: when the numbers move and no release note explains it, the model changed. That makes the suite your independent change detector.
  • Incident and status history. Not the marketing uptime figure: the pattern of how often, how long, how communicated. Three incidents handled with fast honest notice beats one explained six days late.
  • Concentration drift. The Chapter 4.2 ledger changes without anyone deciding to change it, because the third and fourth use cases adopt the vendor already in place: usually the right local decision and an unwatched aggregate one. Re-count items and spend per vendor each quarter.
  • Market signals, tier 1 only. Funding, acquisitions, leadership changes, layoffs in the team building your product, pricing shifts. An acquisition is a roadmap event: the acquirer's plan now governs your next two years, and it is rarely published. Fifteen minutes a quarter of deliberate searching.

The review calendar

Tier-1 vendors are reviewed semi-annually, tier 2 and 3 annually, and every vendor at renewal regardless, because renewal is the only moment your leverage is real and arriving uninformed wastes it. Put the dates in a calendar with an owner, in the same governance rhythm that carries your stage gates. A well-maintained register makes a tier-1 review a forty-five minute meeting: five agenda items and one output.

  1. Dossier refresh. What changed in the sub-processor list, attestation scope, model provenance, ownership? Re-date the file.
  2. Monitoring findings. What the signals showed, and what each cost in re-test days or incident hours.
  3. Concentration position. Items, spend, and where that sits against the stated tolerance.
  4. Exit estimate, refreshed. Re-measured, not read out from last time.
  5. Exportability, tier 1 annually. Did the test run, and what did it find?

The output is a verdict, recorded with a date: continue, remediate (with named actions and an owner), or plan exit (with a target date). A review ending without one of those three words is a status update wearing a review's clothes.

Exit readiness is a property, not a project

Two disciplines keep exit from being something you discover during an emergency. The first is re-measuring the exit estimate. Exit cost grows silently while you do nothing: every integration adds a connection to unpick, every month adds data to migrate, every rule someone tunes becomes knowledge embedded in a system you do not own, every person trained is a person to retrain. An estimate written at signature and never touched understates reality within a year and badly understates it within two, so the option you thought you had has quietly expired without anyone deciding to give it up. Re-measuring turns that from an accident into a choice: you may well decide deepening lock-in is worth it, and that is fine as long as it is a decision.

The second is the annual exportability test for tier-1 vendors. Actually export the data. Actually open it. Confirm it contains what you need, in a form a successor system could consume, including the parts nobody thinks about (history, overrides, configuration, audit records). This is the retrieval-test discipline from Level 3's audit-trail lesson pointed at the vendor relationship: an untested export is a contractual promise, not a capability. Contracts are what you sue over; capabilities are what you use on the Monday you need them.

The concentration discipline, with teeth

Chapter 4.2 started the concentration ledger. The register gives it consequences. State a tolerance in advance, at portfolio level, where the steering committee can see it. Illustratively: no single vendor carries more than five portfolio items or more than 20 percent of AI spend without explicit board-level acknowledgement. The numbers matter less than the existence of a threshold and a named consequence for crossing it. A tolerance with no consequence is a preference.

Frame diversification honestly, because reflexive multi-vendor sprawl is its own failure mode. It costs real money: more contracts, more integrations to monitor, more training curricula, more dossiers to refresh, more relationships each too small to command vendor attention. Concentration buys leverage, coherence, and lower running cost. So the recommendation is not "spread the risk" by default. It is: concentrate deliberately, with a priced and re-measured exit, acknowledged at the level that would have to fund the consequences. Level 5 takes this to enterprise scale, where concentration becomes a board-reported risk category alongside the organisation's other supplier dependencies.

The Register in Practice: One Vendor, One Review

Take the enterprise from this chapter: three approved platforms, eleven portfolio items, a monthly governance forum. Its Vendor Risk Register has seven rows. All figures are hypothetical, chosen to show the shape of the arithmetic rather than to serve as benchmarks.

VendorTierWhyItems / spendReview
A: workflow platform1Customer-adjacent data, routes and updates records4 / 210kSemi-annual
B: agent vendor1Customer PII, executes actions in a system of record2 / 95kSemi-annual
C, D, E: document, analytics, support tooling2Internal data or draft-only authority4 / 130k combinedAnnual
F, G: transcription, meeting summaries3Low data class, no action authority1 / 22k combinedAnnual

Note what tiering by exposure did. Vendor B, at 95k, is tier 1 while vendor C at 78k is tier 2, because B executes actions on customer records and C drafts documents a human sends. Spend-based tiering would have inverted that. Now vendor A's semi-annual review, narrated as it runs.

Dossier refresh. The vendor reissued its DPA in the spring. Someone read it, which is the entire trick. A new sub-processor appeared: an observability provider handling prompt and output logs, in a region the original review did not cover. Nothing improper, nothing announced beyond the reissue. Routed to privacy the same day, assessed in a week, accepted with a note. Elapsed effort: three hours. Elapsed effort had it surfaced through a client's audit twenty-two months later: see the opening story.

Monitoring findings. Two model changes in six months, both notified under the model-change notification clause won after the silent-update incident, each with its two-week testing window used. Re-testing cost roughly four working days each, eight in total, which is exactly what the team budgeted when the clause was negotiated. Say plainly at the review that this is the clause working, because controls that work are invisible and invisible controls get cut. The comparison point is nine days of blind investigation for one unannounced change, plus the credibility damage of users finding it first.

Concentration position. Vendor A now carries four portfolio items at 210k against a stated tolerance of five items and 250k. The fourth was adopted three months ago by a team making an entirely sensible local decision: the platform was already there, already integrated, already trained on. Nobody chose to concentrate; concentration happened. Flagged amber and noted for the next governance meeting, so the board hears it before the fifth item arrives rather than after.

Exit estimate, refreshed. This is the finding that changes the conversation. The signature-date estimate was 90k and ten weeks. Re-measured now: 140k and fourteen weeks, driven by two new integrations (roughly 25k of unpicking), eighteen months of accumulated data and configuration (15k more in migration and validation), and a longer parallel-running period because two of the four items now touch a customer-facing process that cannot go dark. A 56 percent increase over eighteen months, and not one line of it was decided by anyone. The register records the new number, the date, and the drivers, and the honest question follows: do we accept deepening lock-in deliberately, and what would we want in the renewal to make that acceptable?

Exportability test. The export succeeded, on time, in the documented formats, which is better than many organisations discover. But the override history (which AI outputs humans changed and how, the data behind the audit trail and every future model comparison) came back in an undocumented internal format a successor system could not consume without custom parsing. Here is why the calendar matters: it was found four months before renewal, not during an exit. Documented export of override history went onto the renewal risk requirements list and was agreed, because the ask was small, specific, evidenced, and made at the one moment in the year when the vendor wanted something too.

Verdict: continue, with two remediation actions. Export-format clause into renewal (owner: procurement, by the renewal date). Concentration position to the governance forum with a recommendation on whether item five may proceed (owner: strategist, next meeting). Register updated, next review dated, meeting closed inside the hour.

Count what a programme costing perhaps two hours a month across the whole vendor list produced: an unnoticed sub-processor caught, a clause proven to earn its keep, concentration drift surfaced before it became a dependency, an exit option re-priced before it silently expired, a contractual gap fixed with leverage instead of discovered without any. None of it requires a tool. All of it requires a calendar and a named human.

What to Do Monday Morning

You need no budget and no mandate to start: a spreadsheet and about three hours.

  1. List and tier your AI vendors by exposure, not spend. One row each, scoring data class, action authority, and process criticality. Expect an uncomfortable surprise: a cheap tool that turns out to be tier 1, sitting on someone's expense card.
  2. Assemble one tier-1 dossier and date it. Take the vendor carrying most exposure. Gather the sub-processor list, the attestation with its scope statement read properly, the model-change policy, the incident history. Where you cannot get an answer, write "not provided" and the date you asked. Gaps recorded are gaps managed.
  3. Route vendor release communications to a named human. Today. Subscribe that person to the status page and product notes for every tier-1 and tier-2 vendor, and write the name in the register. Fifteen minutes of work, and the control the opening story was missing.
  4. Refresh your oldest exit estimate. Re-measure integrations, data, configuration, retraining, parallel running. Record both numbers and both dates. The delta is the finding, and it is what gets a steering committee's attention.
  5. Run one exportability test before your next renewal negotiation. Actually export, actually open the file, and check the awkward parts: history, overrides, configuration, audit records. Whatever you find becomes a line on the risk requirements list you hand to legal and procurement before negotiation starts.
  6. Put the review dates in a calendar with owners. Semi-annual for tier 1, annual for the rest, plus every renewal. An unscheduled review is an intention, and intentions are what the 140-question questionnaire was made of.

Vendor terms govern the perimeter: who may touch your data, what they may do with it, when they must tell you something changed. The next lesson closes this chapter by governing what crosses that perimeter, in the security and privacy architecture an AI programme needs when the data moves continuously and the systems acting on it are not entirely yours.

Key Takeaways

  • Run AI vendor risk as a four-stage programme (entry diligence, contract as control, continuous monitoring, scheduled review and exit readiness) rather than a questionnaire completed at purchase, because point-in-time diligence against a continuously-changing counterparty is theatre.
  • Tier vendors by data class, action authority, and process criticality instead of spend, since spend-based tiering systematically under-scrutinises the cheap tool touching customer PII with authority to act.
  • Read attestation scope statements rather than attestation logos: an audit covering corporate systems but not the AI service you are buying is a common and consequential mismatch.
  • Name the fourth-party problem by demanding the sub-processor list, binding data-use restrictions on sub-processors, and reading every reissued DPA, because their changes arrive without notice unless a clause creates one.
  • Hand legal and procurement a ranked risk requirements list before negotiation starts (model-change notification with a testing window, data-use restriction, verified deletion, documented export, incident windows, tiered evidence rights, price protection), because at signature your leverage collapses to the renewal date.
  • Build monitoring from five cheap signals with named owners, and treat your own canary suite as the most reliable: regression tests reveal that the model changed even when the vendor never says so.
  • Re-measure exit cost at every review and run an annual exportability test for tier-1 vendors, because estimates age as integrations deepen and an untested export is a contractual promise rather than a capability.
  • Set a concentration tolerance with a named consequence, then concentrate deliberately rather than diversifying reflexively, since sprawl carries real costs and the defensible position is eyes-open concentration with a priced, current exit.