Working with AI Vendors and Contractors
Cornelius Pham remembers the demo like a highlight reel. He is a contracting officer at a mid-sized federal civilian agency, roughly 2,800 employees, a $620 million annual budget, and an obligation to process more than 400,000 benefit applications per year. The vendor arrived with two engineers, a clean dataset, and slides showing 94 percent accuracy on fraud detection. The demo took ninety minutes. The contract was $4.2 million over three years. Eighteen months later the production system was misclassifying one in five applicants in rural districts, the vendor's engineering team had turned over, and Cornelius was sitting across from an Inspector General auditor trying to explain why the agency had paid $1.6 million in milestone payments for a system that never met its accuracy targets. The showroom model had been immaculate. Nobody had looked under the hood.
Most government AI projects involve vendors or contractors, because most agencies will not build everything in house. That creates a governance problem rather than merely a purchasing one. You need external parties to build systems that meet your fairness, accuracy, and governance standards, and you cannot contract for "an AI system" and assume that what arrives will meet them. This lesson covers evaluating vendors, asking the questions that separate maturity from salesmanship, negotiating terms that protect government interests, and managing the relationship through to its end.
Why Vendor Management Is a Governance Function
Three conditions make government AI procurement harder than ordinary IT procurement. First, the AI market is young and immature, vendors make inflated claims about capabilities, and vendor promises about accuracy and fairness frequently do not materialize in production. Second, government carries obligations around fairness, transparency, and accountability that not all vendors understand or are willing to accept, and a vendor who has only served commercial clients may never have been asked to demonstrate subgroup performance to anyone. Third, agencies need independence: when the contract ends, you must be able to maintain and update the system yourself.
That third point deserves emphasis, because vendors sometimes build systems that are impossible to maintain without them, and the incentive to do so is not subtle. Layered on top of all this is the procurement framework itself. You cannot simply pick the vendor you like best. You follow FAR and DFARS requirements, run competitive selection processes, and document your decisions in a form that will be reviewed. Understanding how those constraints interact with AI-specific risk is the actual skill, and it is why bad vendor selection produces poor systems, cost overruns, and governance failures in roughly that order.
Why AI Demos Are Systematically Misleading
Buying an AI system from a vendor is like buying a used car at a dealership. The car on the showroom floor has been detailed, photographed under flattering light, and presented in the smoothest conditions. The vendor controls everything you see, and that control is the problem.
Vendors curate demo environments carefully. They select data their model handles well, choose scenarios that showcase strengths, and run inference on a dedicated server rather than the shared cloud environment your production workload will actually use. The accuracy figures they show are measured on a test set they chose, using evaluation criteria they defined. None of that may match your agency's data, your users, or your operating conditions, and none of it is dishonest in a way you could point to in a protest. It is simply a different question being answered than the one you need answered.
When Cornelius's vendor showed 94 percent accuracy, they were measuring performance on their own holdout set. When the system processed real benefit applications from rural applicants who filed on paper with spotty address histories, accuracy dropped to 79 percent. Nobody in the procurement process had required the vendor to test on data resembling the agency's actual caseload, which is the specific omission that made everything downstream possible.
The fix is a government-controlled proof of concept: a structured pilot using a sample of your own data, evaluated against metrics your team defines. It is the equivalent of taking the used car on a real road before signing. It adds four to eight weeks to procurement timelines and is not optional for high-stakes AI systems. Note what the proof of concept does and does not establish. It shows how the system performs on your data under your evaluation, which is far more than a demo shows. It does not guarantee production performance, because a pilot sample is not a year of live caseload, and it is not a substitute for the monitoring obligations that follow deployment.
Assessing Vendor Capability
Before you sign, assess what the vendor can actually do rather than what they say they have done. Five questions do most of the work. What similar projects have you completed, and may we call the references? Who specifically will do the work, and what happens if that person leaves? What methodology will you use, and is it standard practice or something you invented? What tools and platforms will the system use, and are they stable and well supported or experimental? And how will you handle the discovery that your initial approach does not work, meaning do you have a process for pivoting or only a plan for succeeding?
The answers carry recognizable warning signs. "We have built this type of system before," offered without comparable projects to show. All work performed by a subcontractor, leaving you one layer removed from the actual builders and one layer further from any remedy. "We use proprietary methodology," meaning one that no one else can evaluate. Aggressive timeline estimates, claiming to deliver in six months what comparable efforts take twelve to deliver. And a proposal with no mention of testing, validation, or monitoring, which tells you which parts of the work the vendor considers optional.
Fairness and Governance Maturity
This is where many vendors fall short, and it is rarely because they are unskilled. They have strong technical capability and limited governance maturity, because their previous clients never required any. Ask directly: how do you address fairness and bias in your projects, and is that a defined process or handled case by case? What fairness metrics do you measure, and can you name them? Have you ever had to address a fairness concern, and how did you handle it? How do you approach transparency and explainability? What monitoring and alerting do you build into the systems you deliver?
The red flags in this area are quotable, and evaluators should treat them as near-disqualifying rather than as points deducted. "Fairness is not something we focus on, our systems are technically accurate" is wrong on its own terms, because accuracy and fairness are separate properties and a system can maximize one while failing the other. "We have never had fairness issues" is not reassurance; every system serving a diverse population has fairness considerations, so this answer says the vendor has not looked. "Fairness is your problem, not ours" states an unwillingness to take responsibility that will not improve after award. A proposal with no mention of demographic performance analysis or bias testing says the same thing more quietly. And "the system is objective so it cannot be biased" misunderstands how bias enters systems in the first place, which is a knowledge gap you would then be paying to fill.
Critical Questions for AI Vendors
During evaluation, work through these five groups. The quality of the answers, not their enthusiasm, indicates vendor maturity.
| Group | Questions |
|---|---|
| Requirements and collaboration | How will you work with us to refine requirements during the project? What will you do if the success criteria we defined turn out to be unachievable with our data? How much discovery time do you budget before committing to a timeline? What happens if the data we provide is lower quality than expected? How will you involve our team in decision-making? |
| Testing and validation | What testing methodology will you use? How will you test for fairness and bias? Will you do stratified testing across demographic groups? How will you validate that the system is ready for production? What documentation will you provide on methodology and results? Who validates the test results, you or an independent third party? |
| Operations and monitoring | What monitoring will the system have once deployed? How will you alert us to performance degradation? How will humans override or review system decisions? What happens when the system encounters data it was not trained on? How do you plan for retraining and updates? Will we be able to maintain this system ourselves after the contract ends? |
| Transparency and accountability | How will we explain to the public how the system works? Can we see the factors that influence system decisions? How do you handle audit requests? What documentation will you provide for regulatory compliance? If the system is challenged legally, what evidence will we have that it was fair and accurate? |
| Failure and exit | What happens if the system fails after deployment? What is your rollback procedure? How do you handle incidents where the system makes harmful decisions? What is your contingency if the contract is terminated early? How will you support us if we need to fix problems after you have left? |
Red Flags in Vendor Proposals
Most warning signs appear in writing before the contract is signed. Evaluators who know what to look for can filter out weak vendors before any money changes hands, and the filtering is cheaper at this stage than at any later one.
Vague accuracy claims with no methodology come first. "Over 90 percent accurate" is meaningless without specifying what was measured, on what dataset, using which metric. Precision, recall, and area under the curve, a measure of how well a model separates categories, are different things that can move in opposite directions on the same system. A vendor who cannot specify which metric they are reporting likely does not understand their own system well enough to support it in production. An unconditional claim such as "our system will achieve 95 percent or better accuracy on your use case" is worse, because accuracy is conditional on data quality and should be stated as such. Ask for the basis, the comparable projects, and the comparison to your context, and push back when it does not come.
No model card is a serious warning. A model card is a standardized document describing what a model was trained to do, what data it used, how it performs across demographic subgroups, and what its known limitations are. A vendor who has never produced one has probably never pressure-tested their system for equity failures. No bias testing documentation means the vendor has not measured whether their system performs equally across the populations your agency serves. In a federal benefit context, differential performance by race, disability status, or language preference is not just a technical flaw; it is potential civil rights exposure under Title VI of the Civil Rights Act.
Turn-key claims are the next pattern. "We have a pre-built solution for your problem, we will install it and you are done" misdescribes how AI deployment works. Every deployment requires customization, validation, and governance setup, and a vendor claiming otherwise has not understood the complexity or is hoping you have not. Ask how they will customize to your data and how they will validate the result, and assume you will need three to four months of integration work after handoff regardless of the answer.
Offshore data processing without disclosure is a sovereignty risk. If training or inference runs on servers outside the United States, or through undisclosed subcontractors, your agency may be transmitting sensitive constituent data to jurisdictions outside U.S. legal control. Its absence from a proposal is a procurement disqualifier. No data rights provisions in the draft contract means the vendor's default position is likely that they own the model, the training data you provided, and the inference outputs. That is a recoverable mistake only if you catch it before you sign.
Three further shapes round out the list. A proposal that is vague on scope, of the form "we will build an AI system for fraud detection, timeline six months, cost $500K," lacks the detailed scope, phased timeline, and decision points that a serious plan contains; require sprint plans naming who does what work in which weeks, what constitutes a minimum viable product, and what full rollout means. A proposal built around a permanent retainer, "you will need to keep us on for ongoing maintenance and updates," describes dependence rather than delivery; insist on knowledge transfer and budget for your own staff to take over. And a proposal that ends at deployment, with no mention of monitoring, alerting, incident response, or ongoing maintenance, has defined "done" in a way that leaves the hardest part with you and none of the budget for it.
FAR and DFARS Constraints on AI Procurement
The Federal Acquisition Regulation governs how federal agencies buy goods and services. The Defense Federal Acquisition Regulation Supplement adds defense-specific rules on top. Both create constraints that directly shape AI procurement, and both operate whether or not the technology is novel.
FAR Part 6 requires competition for most federal contracts. Sole-source justifications, contracts awarded without competitive bidding, are permitted only in narrow circumstances. Vendors sometimes argue their AI system is uniquely capable; demand documented technical evidence, not sales materials, and remember that competition gives you alternatives to compare rather than a guarantee of the best outcome. Comparison is the mechanism; running one does not relieve you of evaluating what it produced.
Lowest Price Technically Acceptable, or LPTA, awards contracts to the cheapest vendor who clears a minimum technical bar. It was designed for commodities. Applied to AI procurement, it tends to select the vendor who made the most optimistic claims at the lowest price, because the technical bar is a pass or fail gate and optimism is free. Best value trade-off evaluations take longer but produce defensible awards for complex systems, and the extra time is the same time a proof of concept needs anyway.
Contract Terms That Protect Government Interests
Once a vendor is selected, the contract is your maintenance history: the written record of what the system must do and who is responsible when it does not. Standard government contract terms need AI-specific additions, organized below by what each set protects.
Fairness and validation requirements. The system must pass fairness analysis before deployment, with the contract defining what "pass" means rather than leaving it to the vendor. Demographic performance documentation is required. Bias testing methodology must be documented. Government approval is required before deployment. Bias testing cadence belongs in the performance work statement, the contract section defining what the vendor must actually do: testing at contract award, then quarterly for high-volume benefit systems, with results delivered to the agency rather than retained internally by the vendor.
Documentation requirements. Test plan and test results documented. Model documentation including the features used. Data quality assessment. Fairness analysis and results. Monitoring plan. Incident response procedures. Require a current model card at award, updated model cards whenever the model is retrained, and technical documentation sufficient for an independent third party to evaluate the system's behavior, which is a different artifact from a marketing summary and should be described as such in the requirement.
Independence and maintenance. All code and models must be government-owned. The vendor must provide documentation sufficient for independent maintenance. Knowledge transfer, meaning actual training of your staff, is required rather than offered. Tools must be standard rather than proprietary. These four together are what makes the exit rights below enforceable in practice instead of only on paper.
Data rights, ownership, and sovereignty. Data rights clauses must state explicitly that government data provided for training remains government property. The clause should prohibit the vendor from using that data to train models for other clients, retaining copies after contract expiration, or transferring data to undisclosed subcontractors. FAR 52.227-14 provides a baseline; extend it to AI training data with your agency counsel. Specify what happens to the data when the contract ends. Data sovereignty clauses must specify that all training, inference, and storage occur in government-controlled or FedRAMP-authorized environments, with any deviation requiring written contracting officer approval.
Be precise about what FedRAMP authorization covers. It is the federal cloud security approval framework: it addresses the security posture of the environment. It is not a restriction on how a vendor may use your data, and a fully FedRAMP-authorized service can still have contract terms permitting the provider to use customer data for model improvement. Those are two separate protections requiring two separate clauses, and treating the security authorization as though it covered data use is a common and expensive conflation.
Performance standards and change management. Specify accuracy metrics and thresholds, fairness metrics and thresholds, reliability and uptime requirements, response time requirements, and how long the vendor will support the system. Define the process for model updates and retraining, the process for responding to performance degradation, the process for making changes after deployment, and a stated minimum period of vendor support following deployment.
Liability and indemnification. Address what happens if the system causes harm, meaning how liability is allocated. Address what happens if the system violates civil rights, which is where indemnification language matters most in a benefits or enforcement context. Set insurance requirements. Review any limitation of liability terms the vendor proposes with counsel, because a cap negotiated for a commercial software licence may not be appropriate for a system making determinations about people's benefits.
Exit rights. This is the clause most agencies forget until they need it. At contract expiration or termination, the vendor must deliver model weights, training data, API documentation, and integration specifications within thirty days. Understand what that clause is: a contractual commitment you can hold a vendor to, not a technical mechanism that makes the system portable. Weights delivered in a proprietary format, with no documented training pipeline and no staff who have ever run it, satisfy the letter of the clause and leave you exactly as stuck as before. This is why the independence and maintenance terms above are the ones that give exit rights their force.
Managing Vendor Performance
Service Level Agreements, the contractual commitments defining acceptable performance, are the speedometer on your used car. Most vendor SLAs measure uptime and response latency. A system that is available 99.9 percent of the time but produces inequitable outcomes for fifteen percent of applicants is meeting its SLA while failing the mission, and every one of those applicants is a person with an appeal right.
Write equity metrics into the SLA. Require the vendor to report accuracy separately for each major demographic subgroup your agency serves. Set performance floors, so that accuracy for any subgroup cannot fall more than three percentage points below the overall rate. Tying milestone payments to equity metric compliance gives the vendor a commercial reason to manage the metric, which is a real improvement over having no such incentive. It is not the same as improving the underlying outcome, and a payment-linked metric is also a metric the vendor has reason to optimize narrowly, so read the subgroup definitions carefully and check that the categories reported are the ones your population actually divides into.
Beyond the contract, manage the relationship actively. Hold weekly or bi-weekly check-ins. Establish clear escalation procedures and defined decision-making authority, so it is written down who approves a change. Monitor progress against commitments: are they hitting sprint milestones, are they engaging in requirements conversations or just building, are they doing validation testing, is fairness analysis happening or quietly deferred to the end? Insist on transparency throughout. You should see test results rather than summaries of test results, see the fairness analysis rather than assurances that the system is fair, participate in validation decisions, and have visibility into monitoring data once the system is live.
Write audit rights explicitly into the contract. The agency must have the right to commission an independent technical audit at any time, with access to model documentation, inference logs, and bias test results. Vendors sometimes resist this clause, and resistance is worth probing rather than interpreting. Ask what specifically they object to and why: an objection about protecting a third party's licensed component is a different problem from an objection to any external review at all, and the two lead to different negotiating positions. What should not happen is quietly dropping the clause because the conversation became uncomfortable.
Preparing for the Relationship to End
Set expectations about what "done" means before the vendor reaches it. Do not accept "here is the model, good luck." Completion means a working system, complete documentation, monitoring in place, and knowledge transfer performed, with post-deployment support terms defined in writing. A handoff accepted before those exist transfers the remaining work to your team without transferring the budget for it.
Then prepare for the end from the beginning. Ensure your team has actually learned to maintain the system rather than attended a session about it. Ensure the code and documentation are clear enough that someone who has never met the vendor can work with them. Do dry runs of maintenance tasks while the vendor is still available to answer questions, because a dry run is where you discover which parts of the documentation are fiction. And do not extend the vendor contract indefinitely out of convenience, which is how a delivery relationship becomes a dependency without anyone deciding that it should.
What Happened to Cornelius
Cornelius's agency ultimately terminated the contract for default, a formal remedy under FAR Part 49 when a contractor fails to perform. The vendor challenged the termination. The agency prevailed, but the process consumed fourteen months and two agency attorneys, and during those fourteen months the applicants who had been misclassified were still working through appeals.
The counterfactual is worth stating plainly. A proof of concept requirement would have added four to eight weeks at procurement and would have tested the system on rural paper applications before award, which is exactly where it failed. Tighter contract terms would have made the accuracy shortfall a documented performance failure early enough to withhold milestone payments rather than recover them afterwards. Neither measure was expensive. Both were skipped because the demo was persuasive and the timeline was tight, which is how nearly every version of this story begins.
Anti-Patterns
- Signing on leadership enthusiasm. A proposal arrives, leadership likes it, and the contract is signed without detailed technical review. Work begins and the promises turn out to be unrealistic. Have technical experts review proposals, ask the hard questions, and do not accept vague scope because the timeline is tight.
- Taking the first capable-sounding vendor. The first vendor says they can build what you want and you sign without considering alternatives, then learn later that others would have been better. Run the competitive process and evaluate multiple approaches, not just multiple prices.
- Treating an exit clause as portability. A delivery obligation for model weights is a commitment you can enforce, not a technical guarantee that you can run the system without the vendor. Without standard tools, documentation, and staff who have practiced maintenance, the clause is satisfied and you are still locked in.
- Reading FedRAMP authorization as a data-use restriction. It authorizes the security posture of an environment. It says nothing about whether the provider may use your data to improve their models, which is a separate clause you have to negotiate separately.
- Assuming a commercial tier's terms carry over. Protections attach to the specific product and contract you signed, not to a vendor's name or to a similarly named tier. Verify the terms against the offering you are actually buying, in writing, before award.
- Becoming dependent by default. The vendor is the only party who understands the system, so they can price support at whatever the renewal will bear. Invest in knowledge transfer, insist on documentation, and build internal capability from the start rather than after the first renewal quote.
- Deferring fairness testing to the end. Accuracy testing runs throughout the project while fairness testing waits until deployment as a checkbox, so fairness problems surface when fixing them is most expensive. Test for fairness every sprint.
- Accepting a handoff that is not complete. The vendor says the model is ready, but monitoring is not configured, documentation is partial, and your team has not maintained anything. Define "done" in the contract and hold the line at acceptance.
- Optimizing the paid metric. Tying payment to a fairness number gives the vendor a reason to manage that number. Check that the reported subgroups match your actual population and that improvements show up in outcomes, not only in the reported metric.
Practice Prompts
- Develop a scorecard for evaluating AI vendor proposals. Which criteria matter most, and how would you score technical capability, fairness and governance maturity, approach to testing and validation, clarity of timeline and scope, and risk mitigation? Weight the criteria and decide how you would aggregate the scores.
- You are negotiating with a vendor who proposes a six-month timeline for a complete system, ongoing support you will need to keep buying, optional fairness testing, and model documentation that will be "proprietary." Write your response and the specific terms you would insist on.
- Write out the questions you would ask in vendor interviews across all five question groups. For each, note what answer would make you confident and what answer would concern you.
- Take a contract your agency already holds for an AI or analytics system and check it against the term categories in this lesson. Which categories are present, which are absent, and which are present but too vague to enforce?
- Draft the acceptance criteria for a system currently in development, defining "done" precisely enough that a disagreement at handoff would be settled by reading the document rather than by negotiating.
Reflection
Reflect on a vendor relationship you have been part of, in government or elsewhere. What went well, and what was difficult? Which of the difficulties would better requirements or better contract terms have prevented, and which were about how the relationship was managed after award rather than how it was written? Then ask the uncomfortable version: if that vendor disappeared tomorrow, what could your team still operate, and what would stop? The gap between those two answers is the dependency you accepted without deciding to.
Glossary
- Vendor capability assessment: Evaluation of whether a vendor has the technical skills, governance maturity, and relevant experience to deliver an AI project successfully.
- Key person risk: The risk that a project depends on one specific individual, so that success is threatened if that person leaves.
- Turn-key solution: A pre-built system sold as needing minimal customization or integration, a claim that often indicates unrealistic vendor expectations.
- Knowledge transfer: The process by which a vendor shares understanding of system architecture, maintenance procedures, and operations with government staff so they can maintain the system independently.
- Proprietary systems: Tools, code, or approaches owned and controlled by the vendor that cannot be modified or used independently by the client.
- Model card: A standardized document describing what a model was trained to do, what data it used, how it performs across subgroups, and what its known limitations are.
- Proof of concept: A government-controlled pilot on a sample of the agency's own data, evaluated against metrics the agency defines rather than the vendor's.
- Performance work statement: The contract section defining what the vendor must actually do, and the right home for bias testing cadence and delivery obligations.
- Service Level Agreement: The contractual commitments defining acceptable performance, conventionally uptime and latency, and the place where equity metrics belong.
- LPTA: Lowest Price Technically Acceptable, an award method that selects the cheapest offeror clearing a minimum technical bar. Designed for commodities.
- FedRAMP: The federal cloud security approval framework. An authorization of an environment's security posture, not a restriction on how a provider may use customer data.
Related Lessons
- AI Vendor Evaluation Methodology formalizes the scorecard this lesson sketches.
- Evaluating AI Vendor Claims goes deeper on accuracy claims and how to test them.
- AI Contract Negotiation takes the term categories here into the negotiation itself.
- Federal Acquisition of AI: FAR/DFARS covers the acquisition framework in full.
- OMB M-24-18 and AI Procurement Governance addresses the federal policy layer over AI acquisition.
- FedRAMP and AI Cloud Authorization explains what the authorization does and does not cover.
- Vendor Lock-In Prevention is the deeper treatment of exit rights and independence.
- Managing AI Vendor Performance continues the post-award relationship in detail.
- Testing and Validating AI Systems is the methodology behind the validation clauses.
Closing
The moment you realize you are locked in is the moment your vendor knows they can charge anything for the renewal. Everything in this lesson is an attempt to make that moment arrive later, or not at all: the proof of concept that tests on your data, the questions that reveal governance maturity before award, the contract terms that keep the code and the data yours, and the knowledge transfer that makes the exit clause worth something.
Treat the vendor relationship as a means to a capable system, not as an end in itself. Agencies that manage vendors well get better systems and keep more control over them. Agencies that do not end up in expensive relationships with systems they cannot maintain, explaining to an auditor why the money was spent. Cornelius learned the whole curriculum in eighteen months and fourteen months of litigation. The terms in this lesson cost a few weeks at procurement.
Key Takeaways
- Require a government-controlled proof of concept. Never award a high-stakes AI contract on vendor-run demos alone. Run a structured pilot on your own data against your own evaluation criteria, understanding that it shows performance on that sample rather than guaranteeing production behavior.
- Evaluate governance maturity, not just technical capability. Ask how the vendor addresses fairness, which metrics they measure, what monitoring they build in, and who validates their test results. "We have never had fairness issues" means they have not looked.
- Treat missing model cards and missing bias testing as disqualifying. A vendor who cannot document training data, subgroup performance, and known limitations has not tested rigorously enough for a government benefit context, where disparity is civil rights exposure under Title VI.
- FAR competition requirements apply to AI, and LPTA is the wrong tool. Push back on sole-source justifications backed only by capability claims, and use best value trade-off for complex procurements. Competition gives you alternatives to compare; it does not by itself produce the best award.
- Data rights and data sovereignty clauses are non-negotiable. Government training data must remain government property, must not train models for other clients, and all processing must occur in government-controlled or FedRAMP-authorized environments. Undisclosed offshore processing is a procurement disqualifier.
- FedRAMP authorization is a security control, not a data-use restriction. An authorized environment can still carry terms permitting the provider to use your data. Negotiate the data-use clause separately and in writing.
- Exit rights are a commitment, not a mechanism. Require delivery of model weights, training data, and API documentation within thirty days, and pair it with government ownership of code and models, standard tools, documentation, and real knowledge transfer. Without those, the clause is satisfied and you are still locked in.
- SLAs must include equity metrics, not just uptime. Set per-subgroup accuracy floors and report by subgroup; tying payment to the metric creates a real incentive and also a reason to optimize the metric narrowly, so verify the subgroups and check the outcomes.
- Cover liability, indemnification, and change management explicitly. Allocate liability for harm, address civil rights violations, set insurance requirements, review limitation-of-liability caps with counsel, and define the retraining and post-deployment support processes in the contract.
- Write audit rights in, and define "done" before handoff. Audit access should cover model documentation, inference logs, and bias test results; acceptance should require a working system, documentation, monitoring, and completed knowledge transfer, not a model and good wishes.
Frequently Asked Questions
The vendor says their enterprise agreement means our data will never train their models. Is that enough?
It is a contractual commitment you can hold them to, which is worth having, and it is not a technical mechanism that makes training impossible. Get it in the contract you are signing rather than accepting a statement about the product line, confirm it applies to the specific tier and configuration you are buying, and do not assume it transfers from one product to another because the names are similar. Then pair it with the audit right that lets you verify compliance rather than trusting it.
We are a small agency with no leverage. Can we really negotiate these terms?
Leverage is highest before award and effectively zero afterwards, so the practical answer is to put the terms in the solicitation rather than raise them in negotiation. A requirement stated in the solicitation is a condition of bidding, and vendors who cannot meet it either decline or adjust. Prioritize if you must: data rights, government ownership of code and models, documentation sufficient for independent maintenance, and audit access buy you the most protection per clause.
What if the vendor's accuracy claim is true but was measured differently than we would measure it?
That is the normal case rather than the exceptional one, which is why the proof of concept exists. Ask for the metric name, the test population, and the evaluation criteria, then run your own test on your own data using your own definition of a correct outcome. Where the two disagree, yours describes the system you will operate. Write your definition into the performance standards so that the number in the contract is the number you care about.
Is a subcontractor arrangement automatically a problem?
Not automatically, but it changes what you must verify and it should never be undisclosed. If the actual builders are one layer removed, you need the same capability and governance answers from them, the same documentation obligations flowed down in writing, and clarity about who holds the data and where processing occurs. Undisclosed subcontractors handling government data are a distinct and serious problem, separate from the question of whether subcontracting is acceptable at all.
How much of this survives if the vendor is acquired or goes out of business?
Only the parts you can operate without them. Contract terms bind a successor in ways your counsel can explain, but a firm that has ceased to exist cannot perform knowledge transfer, and a clause requiring delivery within thirty days does not help if there is no one left to deliver. This is the strongest argument for government ownership of code and models, standard rather than proprietary tools, documentation written for an outsider, and maintenance dry runs performed while the vendor is still answering the phone.
Skill.re