←
AI for Government
Proficient · M34 · lesson 34 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Managing AI Vendor Performance
📖
now learning

Managing AI Vendor Performance

15 min

Tomas Reyna, a procurement officer at a mid-sized city, ran a flawless competition for an AI chatbot to handle 311 service requests. The vendor won on a polished demo and a strong price. The contract was signed. Eighteen months later, residents were furious. The chatbot resolved barely a third of requests without a human, routinely misrouted reports of downed power lines to the parks department, and the vendor's quarterly report said only "system operating nominally." When Tomas tried to hold them accountable, he discovered the contract had no performance standards, no right to the data needed to verify performance, and no remedy short of termination, which would have left the city with nothing. He had run a perfect procurement and built a system he could not manage.

The win was at signing; the loss was every day after. Most government AI failures are not procurement failures. They are performance management failures that happen after the contract is signed, when nobody is watching the system closely enough to catch it drifting. Signing a contract only starts the accountability clock. The multi-year period of performance that follows is when an agency either delivers the public benefit promised in the acquisition plan or watches a well-funded system degrade into a liability.

This lesson is about the discipline of managing an AI vendor's performance for the life of the contract, so the system you bought stays the system you need. Two framing points carry through all of it. First, none of the protections described here exist by default. Service levels, credits, audit rights and remedies are terms an agency negotiates into its own contract before signature, and an agency that did not negotiate them does not hold them. Second, the obligation does not transfer with the work. Under federal AI policy, agencies must maintain ongoing monitoring for any use deemed rights-impacting or safety-impacting, which means your agency remains on the hook even when the vendor operates the model day to day.

Why AI vendors need different management

Traditional IT vendor management asks a simple question: is the system up and running? For AI, "up and running" is dangerously incomplete. An AI system can be fully operational and quietly wrong. It can degrade silently as the world changes around it, a phenomenon called model drift, where a model trained on last year's patterns slowly loses accuracy as conditions shift. Tomas's chatbot did not crash. It just got worse, invisibly, while the dashboard stayed green.

Three traits make AI vendors distinct. Their systems can fail without any outage. Their performance depends on data and conditions that change over time. And their inner workings are often opaque, so you cannot verify quality by inspection; you have to measure outcomes. That means your management approach must be built on measurable outcomes, regular monitoring, and contractual rights to the information you need. Without those, you are trusting a vendor to grade their own homework. An AI system can pass every uptime check and still be failing the public, because you manage what you measure and uptime is not what matters here.

Government adds a fairness dimension most commercial clients never face. When a fraud detection model starts disproportionately flagging returns from a particular area, that is not a customer experience issue; it is potential civil rights exposure that can draw a Department of Justice referral. When a border screening system's performance degrades for travelers from specific national origin groups, that becomes an inspector general finding. When a benefits adjudication system produces wrong denials, it directly harms people who hold statutory due process rights. This is why federal AI policy requires disaggregated performance monitoring for rights-impacting uses, and why the EU AI Act treats ongoing post-market monitoring as a mandatory obligation for high-risk systems.

The KPI architecture

A key performance indicator is a measurable target you hold the vendor to. For AI, good indicators measure the outcome the public experiences, not the technology's vital signs. A defensible architecture for a government AI contract has four layers, and every layer must be written into the contract as measurable acceptance criteria with a named test method and a named consequence for a miss. The weighting below is the pattern the source guidance describes, and the four weights sum to 100.

LayerWhat it measuresWeight
Mission-criticalWhether the system delivers the public benefit that justified the acquisitionAbout 50 percent
OperationalWhether the vendor runs the system competently inside your infrastructureAbout 25 percent
GovernanceWhether the vendor meets the transparency, documentation and audit-support obligations you depend onAbout 15 percent
Equity and safetyDisaggregated performance, calibration parity, appeal outcomes and safety-critical failure ratesAbout 10 percent

Mission-critical indicators are specific to the mission. For a benefits triage model it is accuracy against adjudicator review on a statistically representative sample, measured monthly by an independent evaluator, with a floor set at the performance the vendor committed to in its proposal. For a fraud detection model it is precision and recall on a held-out test set refreshed quarterly, plus the disparity ratio across protected classes. For a threat detection model it is the true positive rate at a fixed false positive budget, measured against a corpus generated by a red team. A miss should trigger an immediate root cause analysis request, a cure notice if it is not corrected within the agreed window, and a service credit if the contract provides for one.

Operational indicators cover latency at the 95th percentile, uptime measured by independent probes rather than vendor self-report, data pipeline success rate, and integration health with your authoritative data sources. Governance indicators cover timely delivery of model cards updated to reflect drift, continuous monitoring artifacts, quarterly fairness audit reports, incident disclosure within the 72-hour window many agency contracts now specify, and responsiveness to inspector general document requests. Governance carries the smallest weight of the three upper layers but a miss has outsized consequences, because it impairs your ability to oversee everything else.

Equity and safety indicators carry the smallest weight and deserve the most ruthless attention, because they are the ones most likely to trigger external scrutiny. Disaggregated accuracy by demographic group, calibration parity, the rate at which human adjudicators sustain model decisions on appeal, and safety-critical failure mode rates all live here. For a rights-impacting use these are not optional; they are a precondition for continued operation. The AI management system standard at clause 8.3 and the EU AI Act at Article 9 both require, as the source guidance states, that such metrics be tracked throughout the operational life of the system rather than only at acceptance testing.

Translated to Tomas's chatbot, the four layers ask four plain questions. Is it doing its job? Containment rate, the share of requests resolved without a human, plus routing accuracy and resolution time. Tomas had a vague goal of "good service" and no number at all; a workable contract would have set an explicit containment floor and routing accuracy target, illustratively something like 70 percent containment with routing accuracy above 95 percent, chosen from the agency's own baseline rather than from a slide.

Is it working equitably? Accuracy broken out by neighborhood, language and demographic group, which is where a quiet civil rights problem hides. Is it reliable? Uptime, latency and incident frequency, which still matter and are simply not sufficient. Is it holding up over time? Whether accuracy is trending down, how often the model is retrained, and whether performance holds against a baseline. That last one is the metric that would have caught the chatbot's decline.

Each indicator needs three things: a target, a measurement method, and a consequence if it is missed. An indicator with no consequence is a wish.

The independent monitoring stack

The central failure mode of government AI vendor management is reliance on vendor self-reporting. A vendor that tunes and reports its own performance metrics will, predictably, report favourable numbers. The deepest flaw in Tomas's contract was not the missing targets; it was that he could not independently verify anything, so even a target would have been graded by the party being measured. Independent measurement rather than vendor self-report is the standard the NIST AI Risk Management Framework's MEASURE function points toward, and it is the practical difference between oversight and correspondence.

Three rights make it possible, and all three have to be secured in the contract because you cannot add them later. First, the right to the data: access to the underlying performance data, transaction logs and outcomes, not just a summary the vendor chooses to share. Second, the right to independent measurement: you or a third party must be able to calculate the indicators yourselves rather than accepting the vendor's math. Third, the right to audit: you must be able to examine the system, including bias testing, on a schedule and after any significant change.

With the rights in place, build the stack in five layers. Application performance monitoring captures request-level latency, error rates and throughput, with the monitoring agent running inside your agency boundary rather than inside the vendor's reporting environment. Infrastructure monitoring for uptime and resource utilization is typically consumed from the vendor's continuous monitoring feed but verified by your own synthetic probes, exercising the system at known intervals from endpoints you control. Data pipeline monitoring checks automatically that feature distributions at inference time match the distributions at training time, alerting when drift exceeds an agreed threshold; open-source validation and drift tooling can run inside the agency data lake and warn your team before degradation becomes visible on the vendor's side.

Fairness monitoring is separate from and additional to pipeline monitoring, because fairness drift can occur while aggregate performance looks stable. A fraud detection model can hold 94 percent overall accuracy while its false positive rate on one demographic group doubles. Fairness toolkits and agency-built dashboards should produce disaggregated metrics at least quarterly for rights-impacting uses, with a standing review by the agency's AI governance board. The fifth layer is incident and change management telemetry: every model version, configuration change, deployment and incident generating an immutable audit record your team can reach without vendor gatekeeping. That is a configuration management requirement under federal information security law and a control family many agencies under-enforce for AI systems. Route alerts through your security operations centre or on-call paging so staff can respond within the agreed window.

Set a cadence and put it in the contract so it is an obligation rather than a favour: automated metrics reviewed monthly, a formal performance review with the vendor quarterly, and a deeper audit including fairness testing annually and after any major model change. The acceptance criterion for the whole stack is a single sentence worth memorizing. You should be able to discover any serious issue within 24 hours regardless of whether the vendor chose to tell you.

Escalation tiers: who acts, and when

A government AI contract needs a pre-negotiated escalation ladder with explicit thresholds, response times and authorities. Each tier must have a named authority, a time budget and a documented artifact, or the ladder becomes a diagram nobody climbs.

  1. Tier one, routine support. Handled by the contracting officer's representative and the vendor's account team through the ticketing channel the statement of work specifies. A tier one issue is a minor performance deviation or a discrete bug. The vendor is expected to acknowledge within one business day and to resolve or provide a workaround within five.
  2. Tier two, program management. If tier one does not resolve within the committed window, the representative promotes the issue. This brings the vendor's program manager in, triggers a formal root cause analysis deliverable within 48 hours, and begins accumulating facts for a potential cure notice. Sustained service level misses, fairness anomalies above the agreed threshold and repeated integration failures start here. The contracting officer is informed in writing even when no contractual action is taken.
  3. Tier three, the contracting officer. The officer issues a formal cure notice under FAR 49.402-3 or the equivalent commercial item remedy, specifying the deficiency, the corrective action required and the period in which the vendor must cure. A cure notice is not a termination. It is a defined tool that preserves the government's remedies and puts the vendor's legal team on notice. If the vendor does not cure, the officer may issue a show cause letter demanding an explanation of why the contract should not be terminated for default.
  4. Tier four, the decision point. Termination for default, termination for convenience with documented cause, or transition to a Performance Improvement Plan with a firm deadline.

The remedy ladder and the remediation workflow

Escalation tiers say who acts. The remedy ladder says what the government can do, and it is a different list serving a different purpose. Tomas had only one lever, termination, which was so drastic he never pulled it. A well-managed contract has graduated remedies so that pressure can be applied in proportion to the problem rather than in a single irreversible step, and each rung has to have been negotiated into the contract before signature to exist at all.

The rungs run from notice and cure, where the vendor is formally notified of the missed target and given a defined period to fix it and where most problems should resolve; to a corrective action plan with milestones you approve and track, if the problem recurs; to financial consequences in the form of service credits or withheld payment tied to sustained underperformance, because money focuses attention faster than meetings do; to enhanced oversight through increased reporting frequency and a senior escalation contact while performance recovers; and finally to termination for cause, which is only credible if you hold transition rights and your data is portable. The ladder only works if the lower rungs exist, and Tomas had only the top one, which is exactly why he was stuck with it.

Inside any tier, the remediation workflow follows the same ten steps. Detect, through monitoring or a user report. Confirm that the signal is real and not a sensor artifact. Isolate the scope, identifying which users, use cases or demographic groups are affected. Escalate to the correct tier. Instruct the vendor to perform root cause analysis with a specified deliverable format. Review the proposed fix for contract compliance. Witness or run acceptance testing in a non-production environment. Approve deployment with a rollback plan. Monitor the fix in production for a defined burn-in period. Document the full sequence in the contract file.

Never skip the documentation step. Missing contract file entries are the single most common finding in government AI acquisition audits, and the file is what a successor will rely on when you have moved on. One authority question belongs in the contract from the start and is easy to forget: who may revert a model version without vendor consent, and how quickly. Rollback governance is a configuration management matter, and an agency that has not settled it in writing will discover it during the incident.

Diagnosing a missed target before you escalate

Escalating the wrong way is almost as costly as not escalating, because a remedy aimed at the wrong cause consumes the goodwill and the calendar you will need for the right one. When a committed performance level is missed, the first question is not which rung to reach for. It is which of three things went wrong, because each points at a different tier and a different remedy.

A data problem means the world moved. Input distributions have shifted, an upstream source changed its business rules, or the population the system serves is no longer the population it was trained on. The signal is drift in your pipeline monitoring, often preceding the accuracy decline by weeks. The right response is a root cause analysis deliverable and a retraining commitment, not a cure notice, because the vendor did not breach anything by failing to predict a change in your data. An infrastructure problem means the system is fine but its environment is not: latency, integration failures, capacity or a configuration change. The signal is in your application and infrastructure telemetry, and the fix usually sits at tier one or tier two with a defined remediation window.

A commitment problem is different in kind. The vendor has not staffed the work, has not delivered contract artifacts, has missed remediation deadlines, or is disputing obligations it accepted at signature. The signal is a pattern across incidents rather than any single metric, which is why it is so often missed by teams watching dashboards. This is the case that belongs with the contracting officer, and treating it as a technical problem is how agencies spend a year on root cause analyses for a situation that was never technical. Write the diagnosis down each time, because the accumulated record is what distinguishes a run of bad luck from a pattern, and the pattern is what a cure notice has to describe.

Relationship tiers and what you owe upward

Healthy vendor management runs at three tiers simultaneously, and confusing them is a common source of stalled issues. The operational tier is daily: your representative and the vendor's delivery team, working tickets and metrics. The management tier is the monthly and quarterly rhythm: performance reviews against the scorecard, trend discussion, upcoming changes. The executive tier is periodic and deliberate: the agency's senior accountable official and the vendor's executive sponsor, meeting on strategy, risk and the relationship itself rather than on incidents. When an issue that belongs at the executive tier is raised only at the operational tier, it circulates without resolution, and both parties conclude the other is not serious.

You also owe reporting upward. The senior official accountable for AI in your agency cannot sign an annual use case inventory update in good faith unless you can produce independent telemetry showing the system is meeting its committed performance levels. Brief that official on trends rather than incidents: where each system sits against its scorecard, which indicators are moving in the wrong direction, what has been escalated and to which tier, and what remains unverifiable. Keep the format stable so that the briefings compare across quarters, because a trend line is what makes a two-point decline legible before it becomes a six-point one. The same material answers an oversight inquiry, which is not a coincidence.

Underperformance compounds, and that is the argument for all of this. A small accuracy drop noticed in month one is a conversation with the vendor's program manager. The same drop ignored for six months becomes an audit finding, then an oversight inquiry, and potentially litigation, with the remediation cost rising at every step while the options narrow. Agencies that manage this well have institutionalised the rhythm rather than relying on vigilance: monthly performance reviews, quarterly fairness audits, and a culture in which a cure notice is understood as a professional management tool rather than as a nuclear option.

Performance Improvement Plans with teeth

A Performance Improvement Plan is the structured, time-boxed process used when vendor performance has slipped below committed levels but the government has not yet decided to terminate. In federal practice a plan typically runs 60 to 90 days and contains four elements: specific measurable targets tied to contract indicators, a named executive sponsor on the vendor side who attends weekly check-ins, committed additional resources such as dedicated senior engineers or a named technical account manager, and clear consequences if milestones are missed.

The consequence language is what separates a real plan from a polite letter. It must state that failure to meet milestones will result in termination for default under FAR Part 49, activation of any withholding provisions in the payment schedule, and a negative past performance report in the Contractor Performance Assessment Reporting System, which will follow the vendor into every future federal competition.

A worked example from the source case studies makes the shape concrete. An image classification vendor supplying a secondary screening model has seen accuracy drift from 93 percent at acceptance to 87 percent eight months into operation. A well-structured plan reads: within 30 days, the vendor delivers a root cause analysis identifying whether the drift is due to distributional shift in input imagery, infrastructure changes or a model update. Within 60 days, the vendor delivers a remediated model that has passed agency-run evaluation on a fresh test set at 91 percent or above, with disaggregated metrics inside committed tolerance.

Within 90 days, the model has been in production for 30 days at 92 percent or above with no new fairness anomalies. Weekly check-ins include the vendor's engineering executive and the agency's contracting officer, representative and governance board liaison. Missed milestones trigger, in sequence, a withholding of 10 percent of invoiced amounts, a cure notice and a show cause letter.

Two things are worth noticing in that example. The timeline is internally consistent, since a model delivered at day 60 and observed for 30 days lands exactly at day 90. And the endpoint sits one point below where the system started: the plan restores performance to 92 percent against an acceptance level of 93 percent. That is a deliberate choice a contracting officer should make consciously rather than inherit, because accepting a recovery target below the committed level quietly renegotiates the contract's performance floor.

Two tactical notes. First, do not negotiate the plan into ambiguity. If the vendor wants "substantial progress" instead of a numeric target, the plan has already failed. Second, loop in your agency's small business advocate and competition advocate early, because if the plan fails you will need a transition that likely involves a recompete or an expedited bridge contract. Agencies that treat these plans as the beginning of a possible transition, rather than only as an attempt to rescue the incumbent, protect mission continuity when the incumbent cannot be saved.

Managing change within the rules

AI systems and missions change, so contracts must too, within the bounds of acquisition rules. Under federal acquisition regulations, modifications must stay within the original scope of the competed contract; a change so large it would have attracted different bidders generally requires a new competition. The practical discipline is to anticipate likely changes at the outset and build the flexibility in: priced options for added capabilities, a defined process for adjusting indicators as you learn what good performance looks like, and clear terms for retraining the model on new data. Document every modification with its justification, because that record is what protects the agency if the change is later questioned by an inspector general or an auditor.

The AI vendor performance management kit

Stand up this lightweight kit for every AI contract and review it on the monitoring cadence above. None of it requires specialist tooling, and all of it fails the moment it is maintained by exactly one person.

  • KPI scorecard. One page listing each indicator, its target, how it is measured, who measures it, its layer weight, and the consequence for a miss, covering mission, operations, governance and equity.
  • Data and audit rights register. A record confirming the contract grants data access, independent measurement and audit rights, with the clause references written down so you are not searching for them during a dispute.
  • Monitoring calendar. Monthly metrics, quarterly vendor reviews, annual fairness audit, plus a trigger to audit after any major model change.
  • Escalation and remedy map. The four escalation tiers with named authorities and time budgets on one side, the graduated remedies with their triggers on the other, so nobody has to invent either during an incident.
  • Drift log. A running record of accuracy over time against baseline, retraining events, model versions and any unexplained decline.
  • Modification log. Every contract change with its justification and a scope check confirming it stays within the competed scope.
  • Contract file discipline. Every communication, cure notice, plan milestone and remediation artifact filed in the agency's contract writing system within five business days.
  • Exit readiness. Confirmation that your data is portable and a transition plan exists, so termination is a real option rather than an empty threat.

Anti-Patterns

  • Trust-based monitoring. The agency relies entirely on vendor-provided dashboards, which is the pattern that recurs most often in oversight reports on failed AI acquisitions. A 2024 Department of Homeland Security inspector general report on a border technology vendor found actual availability of 91 percent against a reported 99.7 percent, a gap of 8.7 percentage points, because the agency had no independent uptime probes. The remedy is the independent monitoring stack: never accept a vendor metric you cannot reproduce.
  • Silent fairness drift. Technical monitoring catches nothing because aggregate accuracy is stable, while disaggregated monitoring would have revealed a widening gap across groups. This is the failure that reaches the public through a complaint or a records request rather than through your dashboard. Mandate disaggregated monitoring as a contract deliverable, not as a courtesy the vendor extends when asked.
  • Frozen escalation. The representative identifies an issue but the contracting officer, citing relationship sensitivity, declines to issue a cure notice even as the vendor misses successive remediation deadlines. The remedy is process discipline. Cure notices are not relationship-ending; they are the defined tool for preserving the government's remedies, and an officer who cannot bring themselves to issue one should be rotated off the contract.
  • The toothless improvement plan. Milestones are soft, consequences are vague, and the vendor correctly perceives that nothing will happen if they miss. The remedy is numeric targets, named executives, explicit withholding and termination language, and a transition plan the vendor knows exists.
  • The lost contract file. Turnover on the agency side means the successor officer cannot reconstruct the sequence of decisions, which converts a defensible management history into an audit finding. Continuous contract-file discipline is not paperwork; it is the evidence base your agency will rely on in the review that eventually comes.
  • Treating uptime as performance. Availability answers whether the system responded, not whether it was right. An AI system can post flawless uptime for a year while its accuracy erodes, its routing degrades and its error rate on one group doubles. Every one of those is invisible to an availability metric, which is why the reliability layer carries a modest weight rather than the headline.
  • Reading a green dashboard as assurance. Monitoring detects what you configured it to watch, on the dimensions you thought to instrument, at the frequency you chose. A quiet dashboard is evidence that the configured checks passed, not evidence that the system is healthy. Ask periodically what failure your current instrumentation would not catch, and treat the answer as your monitoring backlog.
  • Assuming service credits are an entitlement. Credits, withholding, audit rights and cure periods exist only where they were negotiated into your contract before signature. No memorandum grants them to you, and no framework supplies them by default. If the register of clause references is empty, so is the remedy ladder, and discovering that during a dispute is the most expensive way to learn it.
  • Accepting a quarterly report as measurement. A vendor summary is an assertion about performance, produced by the party whose performance is at issue, from data you have not seen. It may be entirely accurate. It is not independent, and a management regime built on it produces exactly the outcome Tomas got: eighteen months of "operating nominally" over a system the public had already given up on.

Practice Prompts

  1. Audit one contract for the three rights. Take an AI contract your agency holds and find the clauses granting data access, independent measurement and audit rights. Write down the clause reference for each, or write "absent." Then write the one-paragraph note you would send the contracting officer describing what you can and cannot verify today.
  2. Build the four-layer scorecard. For one AI system, write indicators across all four layers with a target, a measurement method, a measurer and a consequence for each. Assign weights. Then mark which of those indicators you could actually calculate this month without asking the vendor for anything.
  3. Draft the escalation map. Write your four escalation tiers with the named authority, the time budget and the required artifact at each. Circulate it to the people named and see whether they agree that they hold that authority. Disagreements found now are cheaper than disagreements found during an incident.
  4. Write a plan that would actually bite. Draft a Performance Improvement Plan for a hypothetical accuracy decline: numeric milestones at 30, 60 and 90 days, a named vendor executive, committed resources, and the specific consequences of a miss. Check the arithmetic of your own timeline, and decide deliberately whether your recovery target restores the committed performance level or settles below it.
  5. Test your exit. Assume you terminated today. Write what you would be operating in 90 days: where the data is, what format it is in, who could run the alternative, and what the transition clause obliges the vendor to provide. If the answer is uncomfortable, your termination remedy is theoretical and your whole ladder is shorter than it looks.

Reflection

Think about an AI system your agency currently runs under contract. If its accuracy had declined by a few points over the last six months, what specifically would have told you, and when? Who calculates the numbers in your last vendor performance report, and have you ever reproduced one of them independently? If you needed to apply pressure short of termination tomorrow, which rung of the ladder would you reach for, and is it actually in the contract? What would a disaggregated view of that system's performance show, and are you confident enough in the answer that you would rather see it than not? And if you left the role next month, could your successor reconstruct from the file what has happened on this contract and why?

Glossary

  • Model drift. The gradual loss of accuracy as conditions change around a model trained on earlier patterns. It produces no outage and no error, only worse decisions.
  • Containment rate. The share of requests an automated system resolves without escalating to a human. A headline effectiveness measure for service-delivery AI.
  • Disaggregated monitoring. Measuring performance separately for each group of interest rather than in aggregate, which is the only way fairness drift becomes visible while overall accuracy looks stable.
  • Calibration parity. Whether a system's stated confidence corresponds to its actual hit rate equally well across groups, rather than being well calibrated for some and misleading for others.
  • Synthetic probe. A test transaction your agency sends through the system at known intervals from endpoints you control, so uptime and latency are measured independently of the vendor.
  • Service level. A contractual performance commitment with a defined measurement method and period. It binds only to the extent it was negotiated into your contract.
  • Service credit. A financial reduction the vendor owes when it misses a committed service level. An incentive, not a substitute for a termination right, and available only where the contract provides for it.
  • Cure notice. A formal notice under FAR 49.402-3, or the equivalent commercial item remedy, specifying a deficiency, the corrective action required and the time to cure. It preserves the government's remedies and is not itself a termination.
  • Show cause letter. The follow-on notice demanding the vendor explain why the contract should not be terminated for default when a cure notice has not been answered.
  • Termination for default. Ending the contract under FAR Part 49 because the vendor failed to perform, as distinct from termination for convenience.
  • Performance Improvement Plan. A time-boxed structured process, typically 60 to 90 days, with numeric milestones, a named vendor executive, committed resources and explicit consequences for a miss.
  • Contractor Performance Assessment Reporting System. The federal record of past performance. A negative report follows a vendor into future competitions, which is what gives improvement plans leverage.
  • Rollback governance. The pre-agreed answer to who may revert a model version without vendor consent, and how fast. A configuration management question that must be settled before the incident.
  • Burn-in period. The defined window during which a deployed fix is monitored in production before the issue is considered closed.
  • Root cause analysis. The vendor deliverable identifying why a failure occurred, in a format the contract specifies, rather than a narrative assurance that it has been handled.

This lesson picks up where the acquisition ends. Federal Acquisition of AI: FAR/DFARS establishes the regulatory frame, Writing AI Requirements in RFPs and SOWs is where measurable performance standards have to originate, Evaluating AI Vendor Claims and AI Vendor Evaluation Methodology cover the selection that precedes signature, and AI Contract Negotiation is where the audit rights, service levels and remedies described here are actually won. For the monitoring craft, see Continuous Monitoring Fundamentals, AI Metrics and KPIs for Government and Bias Detection and Mitigation at Scale. For what happens when performance failure becomes an incident, see AI Incident Response Planning and Enterprise AI Risk Management. Third-Party AI Risk Management extends this to subcontractors and the supply chain, Vendor Lock-In Prevention addresses the portability your exit remedy depends on, and Oversight Mechanisms: IG, GAO, Congress describes the reviewers your contract file is ultimately written for.

Closing

Vendor performance management is unglamorous work that decides whether a well-run procurement produces a well-run system. The competition is a moment; the period of performance is years. Everything that protects the public during those years, the measurable targets, the independent telemetry, the escalation tiers, the graduated remedies and the contract file, has to be negotiated in before signature and then actually exercised. None of it arrives by default, and none of it survives being treated as a formality.

Tomas Reyna did not fail at procurement. He failed at the thing nobody had told him was part of the job: staying close enough to a system he had bought to notice it getting worse. Build the scorecard, secure the three rights, put the cadence in the contract, and be willing to use a cure notice as the professional management tool it is rather than as a relationship-ending act. A vendor who is measured honestly and escalated to promptly usually improves. A vendor who is never measured has no reason to.

Key Takeaways

  • The contract is the start, not the finish. Most AI vendor failures happen after signing, when nobody is watching closely enough to catch the system drifting.
  • Uptime is not performance. An AI system can run flawlessly and still fail the public through silent accuracy loss. Manage outcomes, not vital signs.
  • Use a four-layer scorecard. Mission-critical, operational, governance, and equity and safety, weighted roughly 50, 25, 15 and 10 percent, each with a target, a measurement method and a consequence.
  • Nothing here is an entitlement. Service levels, credits, audit rights and cure periods exist only where your agency negotiated them into its own contract before signature.
  • Secure the three rights. Data access, independent measurement and audit rights. You cannot add them once a dispute has begun, and without them you are grading nothing.
  • Measure independently. Probes and agents inside your boundary, drift checks on your own data, fairness metrics computed by you or a third party. Never accept a vendor metric you cannot reproduce.
  • Fairness drift hides behind stable aggregates. Overall accuracy can hold steady while performance for one group collapses, so disaggregated monitoring has to be a contract deliverable.
  • Build both ladders. Escalation tiers say who acts and when; graduated remedies say what the government can do. Termination alone is a lever nobody pulls.
  • Make improvement plans numeric. Specific targets, dated milestones, a named vendor executive, explicit withholding and termination language, and a decision made consciously about whether the recovery target restores the committed level.
  • Document everything, within days. Missing contract file entries are the most common finding in government AI acquisition audits, and the file is what your successor inherits.

Frequently Asked Questions

We inherited a contract with none of these terms. What can we do now? Considerably less than you could have done before signature, which is the honest answer, but not nothing. Start by writing down exactly what you can verify independently today, because that defines your real position. Then look for leverage at natural boundaries: option year exercises, renewals, scope modifications and any change the vendor wants. Each is an opportunity to add audit rights, reporting obligations or measurable targets in exchange for something the vendor values. In the meantime, build whatever independent telemetry you can from your own side of the boundary, since synthetic probes and outcome sampling rarely require vendor cooperation.

Is a cure notice too aggressive for a vendor we otherwise like? No, and the belief that it is causes the frozen escalation anti-pattern. A cure notice is a defined contractual tool that specifies a deficiency, states the corrective action required and sets a period to cure. It preserves the government's remedies and it puts the issue in front of people at the vendor who can allocate resources to it. Agencies that use cure notices routinely, as a professional management step rather than a rupture, generally get faster remediation and have healthier vendor relationships than agencies that avoid them until the situation is unsalvageable.

How do we monitor fairness without becoming data scientists? The hard part is access and cadence rather than mathematics. Fix the groups you will report on in advance, in the contract, along with the metric definitions and the reporting frequency, so nobody is choosing them after seeing results. Then either require the vendor to deliver disaggregated metrics on a schedule while retaining your right to reproduce them, or have a third party compute them from data you control. Open fairness toolkits handle the computation. What they cannot do is decide which groups matter for your program or notice that a quarterly report stopped arriving.

The vendor says our monitoring requirements are unreasonable and expensive. Are they? Sometimes the cost objection is genuine, particularly for a small vendor being asked for bespoke telemetry. Test it by separating what you need from how they proposed to deliver it. Much of the stack, including synthetic probes, outcome sampling and your own pipeline checks, runs on your side of the boundary and costs the vendor nothing. Where the requirement is genuinely on them, price it as a line item rather than debating it as a principle. A vendor who resists any independent verification, at any price, is telling you something about what verification would reveal.

Our vendor's report says performance is fine and our users say it is not. Who is right? Assume both are describing something real and find the gap between them, because that gap is usually the finding. Common causes: the vendor measures a different population than the one complaining, measures at a different time granularity that averages away a bad period, excludes cases that error out before reaching the model, or reports aggregate accuracy while the problem is concentrated in one group or one request type. Ask for the definition, the population and the exclusions behind each reported number. If you cannot reconstruct their figure from data you can see, you have identified your first monitoring requirement.

When should a Performance Improvement Plan become a transition instead? Treat it as both from the day you start it. Improvement plans typically run 60 to 90 days, which is rarely enough time to stand up a replacement afterwards, so the transition planning has to begin in parallel rather than sequentially. Bring in your competition advocate and small business advocate early, understand what a bridge contract or recompete would take, and know your data portability position. If the vendor meets its milestones, you have lost a little planning effort. If it does not, you have preserved mission continuity, which is the outcome the public actually experiences.