←
AI for Government
Proficient · M29 · lesson 29 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Federal Acquisition of AI: FAR/DFARS
📖
now learning

Federal Acquisition of AI: FAR/DFARS

15 min

Raymond Castellano, a contracting officer at a federal civilian agency, signed a $2.4 million firm-fixed-price contract for an "AI-powered case triage system." The vendor demoed beautifully. Six months in, the program office discovered the model had been trained on the vendor's own synthetic data, the agency had no rights to retrain it, and switching vendors meant rebuilding from scratch. The contract said nothing about any of this, because Raymond had bought AI the way he bought office furniture. The Federal Acquisition Regulation gave him every tool he needed to avoid the trap. He just had not adapted his playbook to what makes AI different. This lesson is that adaptation.

You already know acquisition. What changes with AI is not the rulebook but the risk profile: the "product" keeps learning, the data is the asset, performance drifts, and the thing you are buying can fail in ways a chair never will. The Federal Acquisition Regulation, or FAR, and its defense supplement, DFARS, are flexible enough to handle all of it, if you plan for it up front.

Why AI breaks the standard playbook

A normal IT buy delivers a fixed thing that does a fixed job. AI violates four assumptions baked into that model:

  • It changes over time. A model retrained on new data behaves differently next quarter. "Acceptance testing once at delivery" no longer proves much.
  • The data is the real asset. Who owns the training data, the fine-tuned model, and the outputs determines whether you are free or trapped. This is the single most expensive thing to get wrong.
  • It can be confidently wrong. A model can produce fluent, plausible, incorrect output. Your contract needs accuracy and oversight requirements, not just uptime.
  • It carries bias and rights risk. If the tool touches benefits, enforcement, or hiring, fairness and the ability to explain a decision become contractual obligations, not afterthoughts.

In a normal IT buy you acquire a product. In an AI buy you acquire a relationship with data, drift and risk that outlives the delivery date.

The specification tension at the heart of it

Federal procurement is built around clear specifications and measurable deliverables. Procure a building and you specify dimensions, materials and performance standards, and the contractor delivers what was specified. Procure conventional software and you specify functional requirements and acceptance criteria, and if the software meets them you accept it. AI procurement is messier, because AI systems are trained rather than fully programmed. You cannot perfectly specify how the system will behave. You cannot demand perfect accuracy, because accuracy involves trade-offs against speed and sometimes against fairness. You cannot guarantee performance on data the model has never seen, because machine learning generalises imperfectly.

That creates a genuine tension: federal procurement requires clear acceptance criteria, and AI systems are inherently uncertain. The art of AI procurement is managing the tension honestly rather than pretending it away. Be precise about what you can specify, which is more than people assume: data requirements, testing protocols, fairness thresholds, monitoring obligations, documentation deliverables and exit terms. Be explicit about what you cannot specify, which is exact accuracy on every edge case. A contract that pretends to certainty it does not have will produce a dispute at acceptance, and the government usually loses that dispute because the specification was unmeetable.

Acquisition planning for AI

Planning is where you win or lose an AI buy. Three planning decisions matter most. Define the use case and its risk tier before you write a word of the requirement. Is this AI making or influencing decisions about people, or just helping staff draft documents? A rights-impacting use needs testing, oversight and appeal provisions written into the contract. A back-office productivity tool does not. Right-sizing here keeps you from over-buying compliance you do not need, or under-buying protection you do.

Choose your contract type to match the uncertainty. Raymond's mistake was a firm-fixed-price contract for something whose performance he could not yet specify. When the AI's behavior is well understood, fixed-price works. When you are still learning what "good" looks like, a phased approach, with a paid pilot or evaluation period before full commitment, lets you test before you bet the budget. Plan for the whole lifecycle. Budget for monitoring, retraining, model updates and an exit. AI is not a one-time purchase; it is an ongoing operation, and the money you did not budget for year two is the money that funds the vendor's leverage over you.

Buy, build, or hybrid

The first structural decision is whether to procure an off-the-shelf product, contract for a custom build, or start commercial and customise. Buying is faster and cheaper to start, the vendor provides support, and the vendor maintains its own security and privacy compliance posture. Read that last advantage carefully: the vendor maintains compliance for the vendor's environment. Your agency remains accountable for whether your use of the system complies, and no vendor attestation transfers that. The trade is less control and a system that may not match your requirements exactly, so buying suits standard capabilities like language processing, image recognition or classification that commercial products already do well.

Building means contracting a development firm for a custom system. You get something tailored to your requirements, you can own the result, and you can optimise for your use case. You also pay more, wait longer, carry more risk and have to manage the vendor closely. Build when your requirements are genuinely unique, commercial products do not meet them, and you need complete control. Hybrid starts with a commercial product and customises it under contract. It suits the common case where a commercial product provides most of what you need, in the range of seventy to eighty percent, but requires significant configuration or extension. Hybrid reduces risk against a full custom build and demands careful vendor management, because the boundary between "product" and "our customisation" is exactly where data rights arguments start.

Task order or blanket purchase agreement

The second structural decision is the vehicle. A single task order issues a contract for a specific AI project with defined scope, timeline and deliverables. The scope is clear, it is easier to manage, and it suits well-defined projects. The cost is that similar work later has to be recompeted, which is inefficient when you know the work is ongoing. A blanket purchase agreement establishes terms with one or more vendors so you can issue task orders repeatedly without recompeting each one. Orders move faster, the vendor learns your environment, and ongoing work becomes practical.

The risk in a blanket agreement is concentration. If only one vendor sits on the agreement, the convenience that made it attractive becomes the mechanism of lock-in, and by the third task order the switching cost is doing your negotiating for you. Where the subject is AI, that risk compounds, because each order deepens the vendor's hold on your data and your fine-tuned models. Put more than one capable vendor on the agreement where you can, and keep the data-rights terms identical across all of them so a switch is a commercial decision rather than a technical rebuild.

Fixed-price, cost-plus, or both

Fixed-price means the vendor commits to specific deliverables for a fixed sum. The vendor absorbs cost overruns and you absorb performance risk. It suits well-defined projects with clear requirements and low technical uncertainty. The problem for AI is that AI projects are inherently uncertain, so a vendor may underbid to win and then deliver a weak system rather than lose money on it. That is precisely what happened to Raymond.

Cost-plus means the vendor bills actual costs plus a fee. You absorb cost risk and the vendor absorbs performance risk. It accommodates the uncertainty of AI development, which is why it is often the better fit. Be careful how you sell that internally, though. Cost-plus removes the vendor's incentive to underbid; it does not create an incentive to finish, and it moves the cost exposure onto the government. It works when it is paired with real oversight: a government technical representative close to the work, scheduled technical reviews and staged acceptance. Without those, cost-plus buys you time and nothing else. A hybrid often beats either: fixed-price for well-defined deliverables such as training data curation and system documentation, cost-plus for the development work whose requirements will evolve.

FAR Part 12 and the commercial-item question

Most AI tools are sold as commercial products, which points you to FAR Part 12, the streamlined path for buying commercial items. Part 12 is faster and lighter, and for genuine commercial AI it is often the right route. But it has a catch worth naming.

Commercial items come "as is," on the vendor's standard terms. Those terms typically favor the vendor on exactly the things AI makes critical: data rights, liability for harmful output, and the right to change the model. Using Part 12 does not mean accepting boilerplate. You can and should negotiate the terms that matter for AI: ownership of your fine-tuned model and outputs, accuracy and bias requirements, notice before model changes, and your right to your own data on exit. The streamlined path is about process speed, not surrendering protections.

Which existing clauses reach an AI buy

Neither the FAR nor DFARS was written with machine learning in mind, and the regulations do not address fairness, explainability or model provenance as such. What they do contain are clause families that reach AI procurements squarely, and the practical skill is knowing which family to pull and what it will and will not do for you. Acquisition regulation also changes, so treat the list below as the shape of the problem and confirm the current clause text and prescription with your contracting office before you cite anything in a solicitation.

Four clauses are worth knowing by number. FAR 52.212-4, Contract Terms and Conditions for Commercial Items, governs the terms for commercial AI products and services, and is where an off-the-shelf AI buy lands. On the defense side, DFARS 252.204-7012, Safeguarding Covered Defense Information and Cyber Incident Reporting, obliges the contractor to safeguard covered information and to report cyber incidents, which reaches an AI vendor whose system suffers adversarial attack or data poisoning. DFARS 252.227-7013, Rights in Technical Data, and DFARS 252.227-7014, Rights in Noncommercial Computer Software and Noncommercial Computer Software Documentation, set the government's rights in technical data and in software developed under contract, which for an AI development effort is where the argument over models, code and documentation is actually resolved.

Four further clause families matter and are best raised with your contracting officer by subject rather than by number, because the citations attached to them in circulation are unreliable and this lesson will not supply a corrected one. Service contract labour standards can apply when you contract for the services of people who will build an AI system, as distinct from buying a product. Patent and invention rights clauses govern who owns an invention made with government funding, which for a custom model is a live question. Rights-in-data clauses govern what data you hold after the contract ends, and this is the one that decides whether you can retrain or audit the system later. Supply chain risk management clauses reach the provenance of models, components and subcontractors, and NIST cybersecurity standards are pulled in by reference through the safeguarding requirements. Ask for these by name and let the contracting office attach the current citation.

Writing the clauses the regulations do not supply

Because no clause covers the AI-specific ground, you write it. Three pieces of contract language do most of the work, and they should be drafted before the solicitation goes out rather than negotiated after award.

Data rights. Specify what happens to the data at contract end. Working language requires the contractor to provide the government with all training data used to develop the system, organised and documented to a stated standard; sufficient documentation to understand data sources, quality and characteristics; all test data used to evaluate performance; documentation of preprocessing and feature engineering steps; and all data governance and quality assurance documentation. That is five distinct deliverables, and vendors will try to concede the first and skip the rest. If you do not specify, the vendor may retain the training data, and without it you cannot retrain or audit the system.

Model rights and access. Specify your rights to the model itself: complete model code, whether open source or proprietary; model weights and parameters in a standard format; documentation sufficient to deploy and retrain the model; complete training logs and hyperparameter documentation; all source code used in development; and rights to use the model for government purposes including audit and testing. Six deliverables again, and the fourth and fifth are the ones that determine whether a successor team can actually operate what you bought. You cannot audit, improve or redeploy a system you do not have access to.

Performance and fairness metrics. Specify what the system must achieve and how it will be measured, in terms someone can test. Accuracy stated against a named test dataset, a fairness standard with a defined investigation trigger, transparency requirements such as returning the factors that drove a classification, and testing requirements including user acceptance testing by named government staff with government approval required before deployment. Make fairness testing mandatory rather than optional, specify the metrics in the contract, require the vendor to document methodology and results, reserve the right to an independent bias audit, and do not accept an argument that a discovered bias is unfixable. Requiring remediation is the point of writing the clause.

An AI acquisition clause checklist

Use this as a planning artifact. For any AI buy, confirm your requirements document and contract address each item before release. Adapt the depth to the risk tier.

AreaThe question your contract must answer
Data rightsWho owns the training data, the fine-tuned model, and the outputs? Can the vendor reuse your data?
Model accessDo you receive weights, code, training logs and hyperparameters in a usable format, with the right to retrain?
PerformanceWhat accuracy, error-rate, or quality thresholds must the system meet, against which dataset, and how are they measured over time?
Drift and updatesMust the vendor give notice before changing the model? How is performance re-verified after updates?
Bias and fairnessFor rights-impacting uses, what fairness testing and reporting is required, and what happens when a test fails?
TransparencyCan the system explain its decisions enough for staff oversight and citizen appeal?
SecurityWhat safeguards apply to your data, and does the vendor meet the required cybersecurity standards?
ProvenanceAre open-source components, dependencies and their versions documented, with supply chain risks identified?
ExitOn termination, do you get your data and a clean off-ramp without rebuilding from zero?
LiabilityWho is responsible when the AI produces harmful or wrong output that affects a citizen?

Two acquisitions, worked

A civilian agency integrating commercial sentiment analysis for citizen feedback writes a modest, testable set of requirements: the system returns a sentiment score plus the top three factors contributing to each classification, so a staff member can see why an item was categorised as it was; the vendor conducts user acceptance testing with five government staff and the government must approve before deployment; a fairness standard applies with a defined trigger for investigation. Management runs on monthly performance reviews covering accuracy, processing volume and user feedback, quarterly fairness audits broken out by feedback source, and a thirty-day warranty under which the vendor fixes shortfalls against specification at no cost after initial deployment. In this scenario the integration completed on schedule and the agency got systematic analysis of citizen feedback it had never had.

A defense department needing custom analysis of satellite imagery for equipment detection runs the opposite structure: a twenty-four-month cost-plus development contract with a fixed management fee, estimated at two million dollars plus fee, with a computer vision firm. The clauses do the heavy lifting. Data rights: the government owns all training imagery, the model, code and documentation. Model provenance: the vendor documents every open-source model and dependency with versions, and documents supply chain risks. Performance: 95 percent detection accuracy on a labelled test set, false positive rate below 2 percent, latency below 100 milliseconds per image. Security: resilience to adversarial image perturbation, with documented defences against model extraction. Robustness: testing across diverse geographies, lighting conditions and seasons, with performance documented per condition.

Two features of that second contract deserve copying. It keeps the vendor engaged after delivery, with twelve months of post-deployment monitoring, improvement recommendations and quarterly retraining if accuracy degrades. And it stages acceptance rather than betting everything on a final delivery: prototype at month six, alpha at month twelve, beta at month eighteen, final at month twenty-four, with weekly technical reviews, monthly accuracy reporting against the test set, quarterly robustness assessments, and a government technical representative embedded with the vendor team. Staged acceptance is what converts an uncertain development effort into four smaller decisions you can still walk away from.

DFARS and defense AI

If you are in the defense space, the Defense Federal Acquisition Regulation Supplement adds requirements on top of the FAR. Two areas dominate AI buys. First, cybersecurity: defense contractors handling controlled unclassified information must meet specified safeguarding and reporting requirements, and AI vendors processing your data are squarely in scope. Adversarial attacks and data poisoning against a fielded model are cyber incidents, and your contract should say so explicitly rather than leaving the vendor to decide whether a degraded model counts.

Second, supply chain and provenance: where the model, its components and its data come from matters intensely when the application is mission or safety related. A modern AI system is assembled from pretrained components, open-source libraries and third-party datasets, most of which arrived without a bill of materials. Build provenance documentation in as an evaluation criterion, not paperwork you check at award, because a vendor that cannot tell you what its model is built from at proposal time will not be able to tell you after a supply chain advisory lands either. Where the system runs in a cloud environment, the environment's authorization is a separate question from the contract terms, and one does not answer the other.

The prerequisite nobody puts in the acquisition plan

Every control in this lesson assumes someone on the government side can tell whether a deliverable is any good. Independent technical review, rejection of inadequate work, staged acceptance, verification of fairness testing: each of these requires a person who understands what they are looking at. Agencies that lack that person do not fail loudly. They accept everything, because the alternative is to reject a deliverable they cannot articulate a problem with, and no contracting officer wants to defend that position.

Building government technical capacity is therefore a procurement control, not an HR initiative, and it should appear in the acquisition plan alongside the budget for monitoring. In practice it takes one of three forms. You hire or detail a technical representative who sits close enough to the work to see it, which is what the staged-acceptance model in the defense example depends on. You buy independent technical review separately from the development contract, so the reviewer's incentives are not the developer's. Or you build a small shared pool of AI-literate staff who support multiple acquisitions, which is how smaller agencies make the economics work.

Whichever route you take, decide it before award rather than discovering late in the effort that the only people who can evaluate the vendor's work are the vendor's. Write the review role into the contract as a government responsibility with named access to the work product, because a technical representative without contractual standing to see the training logs is a job title rather than a control. And be honest in the acquisition plan about what happens if the role goes unfilled: the correct answer is a simpler system with fewer things to verify, not the same system with unverified deliverables.

Raymond's contract, rewritten

With this checklist, Raymond's $2.4 million buy looks different. He would have run a paid evaluation period before full commitment rather than committing the whole budget on a demo, written a data-rights clause securing the training data, the fine-tuned model and the outputs, required notice before retraining, set a measured accuracy threshold against a named test set and checked it quarterly, staged acceptance rather than accepting once, and guaranteed a clean data export on exit. None of that requires a new authority. It is the FAR he already knew, applied to the risks AI actually carries. The discipline is in the planning, where it has always been.

Anti-Patterns

Four of these are documented failure modes in AI acquisition. The last two are the ways a well-drafted contract still ends badly.

  • Inadequate technical review during execution. The government lacks the AI expertise to evaluate the work, the vendor says the deliverable is complete, and the government accepts. The system does not meet requirements and nobody knows; corners were cut on testing and documentation and nobody catches it; the agency inherits a weak system that fails later, and rebuilding is expensive. Require a government technical representative throughout, conduct independent technical reviews monthly or quarterly, build acceptance testing with government sign-off into the contract, and be willing to reject an inadequate deliverable even when rejection extends the timeline.
  • Over-specifying technical requirements. Specifying exact algorithms, data sources and implementation details feels like control and produces the opposite. The vendor cannot adapt when requirements shift, will not propose a better approach the specification forbids, and optimises for meeting the letter of the spec rather than delivering a good system. Specify what you need, not how to build it, and put the precision into performance requirements instead.
  • Inadequate data rights and IP provisions. Data rights are boring next to performance specifications, so they get skipped. After the contract ends the vendor will not hand over the model or the training data, you cannot retrain or audit, you remain dependent for support, and any new use requires renegotiating from a position of no leverage. Include a comprehensive data rights clause, specify the format deliverables must arrive in, and treat data and model rights as non-negotiable rather than accepting the argument that they are vendor property.
  • Inadequate fairness and testing requirements. Fairness testing gets treated as nice to have, the system deploys with unknown bias, and the agency discovers it when citizens complain. Then the system comes down for repair, operations are disrupted, and legal exposure follows. Make fairness testing mandatory, specify the metrics, require documented methodology and results, and reserve an independent audit right.
  • Reading a vendor's compliance posture as your own. A vendor's certifications and authorizations describe the vendor's environment and controls. They are useful evidence and they are not a transfer of your agency's accountability for how the system is used, what decisions it supports, or whether that use complies with the statutes binding you. Ask what the certification covers, then ask separately what your obligations are.
  • Treating acceptance as the finish line. An AI system accepted against a delivery-day test will drift, and the contract that has nothing to say about month fourteen has handed the vendor every subsequent negotiation. Budget and write terms for monitoring, retraining, re-verification after model changes, and an exit, and check at award that someone is funded to do the monitoring you just required.

Practice Prompts

  • Design an AI procurement strategy for a capability you know: buy, build or hybrid and why; fixed-price, cost-plus or a split and why; the evaluation criteria that matter most; and the biggest risks with the mitigation for each.
  • Draft the fairness and data rights clause. Specify the fairness metrics the system must meet, the data rights the government will hold, the testing the vendor must conduct, and what happens when testing reveals bias. The last of those four is the one most drafts omit.
  • Develop a monthly vendor performance assessment form: which metrics you will track across schedule, quality and testing; how each is measured; what counts as acceptable and what counts as concerning; and what specifically triggers escalation.
  • Create a technical evaluation framework for AI vendor proposals: the technical criteria that matter, how each is scored, the red flags that would downgrade a proposal, and the questions you will ask every vendor about their approach.
  • Take a contract your office already holds for an AI or automated system and run it against the ten-row checklist in this lesson. Count the rows the contract does not address, and identify which of those you could still fix through modification.
  • Write the staged acceptance schedule for a hypothetical multi-year AI development effort, defining what would have to be true at each stage for you to authorise the next one, and what you would do if it were not.

Reflection

If your agency had to walk away from its most important AI vendor next month, what exactly would you be able to take with you? Name the artifacts: the training data, the model weights, the training logs, the documentation, the outputs. For each one, point to the clause that entitles you to it and the format it would arrive in. Then ask the harder question: if the clause exists but you have never tested it, do you know whether the vendor can actually produce what it promises? Data rights that have never been exercised are a hypothesis about the contract, not a property of it, and the moment you discover otherwise is the worst possible moment to find out.

Glossary

  • FAR: The Federal Acquisition Regulation, governing procurement by federal agencies generally.
  • DFARS: The Defense Federal Acquisition Regulation Supplement, adding procurement requirements for the Department of Defense on top of the FAR.
  • FAR Part 12: The streamlined path for acquiring commercial items, faster in process while still permitting negotiation of the terms that matter.
  • Fixed-price contract: The vendor commits to deliverables for a set sum and absorbs cost overruns; the government carries the performance risk.
  • Cost-plus contract: The vendor bills actual costs plus a fee; the government carries the cost risk and must supply the oversight that a fixed price would otherwise have forced.
  • Blanket purchase agreement: A standing arrangement allowing repeated task orders without recompeting each one, efficient for ongoing work and a concentration risk if only one vendor holds it.
  • Task order: A specific procurement request issued under a pre-existing contract or blanket agreement.
  • Data rights: The contractual specification of what data the government holds and controls once the contract ends, and the clause that determines whether you can retrain or audit the system.
  • Acceptance criteria: The objective specifications a system must meet for the government to accept the deliverable, stated against a named dataset where accuracy is involved.
  • Staged acceptance: Structuring delivery as a sequence of gated milestones so an uncertain development effort becomes several smaller decisions rather than one irreversible one.
  • Model provenance: Documentation of which pretrained models, components, dependencies and versions a system is built from, and the supply chain risks each carries.

Closing

Federal AI procurement is in transition. The current regulations were written before AI became prevalent and do not explicitly address fairness, explainability or model provenance. Effective procurement officers are adapting existing rules to cover AI-specific concerns while staying inside them, which is a drafting problem more than an authority problem. Expect the regulatory picture to keep moving, and expect your contracting office to be the authority on where it currently stands.

The most useful shift in mindset is to treat the contract as a collaboration tool rather than a legal artifact. A well-drafted AI contract clarifies expectations, makes the vendor's obligations testable, and protects the agency's ability to keep operating the system after the vendor leaves. A poorly drafted one produces disputes, weak deliverables and a system you cannot maintain. Raymond's contract was legally sound and operationally useless, and the gap between those two is the entire subject of this lesson.

Key Takeaways

  • The FAR already handles AI. You do not need new authority, you need to apply existing acquisition discipline to AI's distinct risks: drift, data ownership and confident errors. Neither the FAR nor DFARS was written for AI, so confirm current clause text with your contracting office rather than citing from memory.
  • Specify what you can and be honest about what you cannot. Data requirements, testing protocols, fairness thresholds, monitoring and exit terms are all specifiable. Exact accuracy on every edge case is not, and a contract that claims otherwise fails at acceptance.
  • Set the risk tier first. Whether the AI affects decisions about people determines how much testing, oversight and appeal protection your contract must require.
  • Match contract type to uncertainty. Fixed-price invites underbidding on uncertain AI work; cost-plus accommodates the uncertainty but shifts cost risk to the government and only works with embedded technical oversight and staged acceptance. Splitting the two by deliverable often beats either.
  • Data and model rights are the buy. Training data, test data, weights, code, training logs, hyperparameters and documentation, in a specified format, with the right to retrain and audit. Treat these as non-negotiable rather than accepting a vendor-property argument.
  • Fairness testing is mandatory, and so is the remedy. Specify the metrics, require documented methodology and results, reserve an independent audit right, and write in what happens when a test fails. Do not accept the claim that a discovered bias cannot be fixed.
  • Part 12 speed is not surrender. Buying commercial AI streamlines process but you still negotiate data rights, accuracy, change notice and exit terms off the vendor's boilerplate.
  • DFARS adds safeguarding and provenance. Defense buys must address protection of covered defense information, cyber incident reporting that expressly covers adversarial attack and poisoning, and documented provenance of models, components and dependencies as an evaluation criterion.
  • Stage acceptance and stay engaged after delivery. Gated milestones with technical reviews turn one irreversible bet into several reversible ones, and post-deployment monitoring with retraining obligations is what keeps an accepted system working.
  • A vendor's compliance is not your compliance. Certifications describe the vendor's environment. Your agency answers for how the system is used and what it decides.

Frequently Asked Questions

Are there AI-specific FAR clauses we should be citing?

Not in the way people hope. The regulations contain clause families that reach AI procurements, particularly around commercial items, safeguarding information, and rights in technical data and software, but they were not drafted with machine learning in mind and they do not address fairness, explainability or model provenance. That is why the AI-specific terms in this lesson are ones you draft yourself. Acquisition regulation also changes, so the reliable move is to bring your contracting officer the subject and let them attach the current citation rather than citing a clause number you read somewhere.

Can we require a vendor to hand over model weights?

You can require it, and whether you get it depends on what you are buying and what you are willing to pay for. For a custom development effort funded by the government, comprehensive rights in the resulting model, code and documentation are a normal ask and should be non-negotiable. For a commercial product built on a vendor's proprietary base model, full weights are usually not on offer at any price, and the realistic negotiation is over your fine-tuned layer, your data, your outputs and a documented exit. Decide which of those two situations you are in before you write the requirement, because asking for the impossible in a commercial buy just narrows your field of bidders.

Our vendor says fixed-price is cheaper. Should we take it?

Only if you can actually specify the performance you are buying. Fixed-price is genuinely cheaper for well-understood work, and it is a trap for work whose success criteria you are still discovering, because the vendor prices the uncertainty into the bid or absorbs it by cutting quality. Test it with a question: can you name the dataset the accuracy will be measured against and the number it must hit? If you cannot answer today, you are not ready to fix the price, and a paid evaluation period before the main award will cost less than the dispute you would otherwise be having at acceptance.

What if the fairness testing comes back showing a disparity?

Then the clause you wrote earns its keep, which is why it must state the consequence in advance. Require the vendor to document the finding and the cause, propose remediation with a timeline, and re-test. Do not accept an assertion that the disparity is inherent to the problem and cannot be addressed; that is a claim the vendor should have to evidence, and the contract should give you an independent audit right to test it. Set the threshold for what counts as significant in writing before testing starts, so nobody is negotiating the definition after seeing the result.

How do we avoid Raymond's lock-in on a commercial product?

By pricing the exit at the start, when you still have competitive leverage. Establish in the contract what you receive on termination, in what format, and within what period: your data, your fine-tuned model where the arrangement permits it, your configuration, and documentation sufficient for a successor to stand the capability up elsewhere. Then test the export once during the contract rather than trusting the clause. Where you are on a blanket agreement, keeping more than one capable vendor on it and holding the data terms identical across them is worth more than any single exit clause, because it keeps switching a commercial decision rather than a rebuild.