←
AI for Government
Proficient · M37 · lesson 37 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
OMB M-24-18 and AI Procurement Governance
📖
now learning

OMB M-24-18 and AI Procurement Governance

15 min

Angela Foster, a contracting officer at a federal transportation agency, had bought software for 15 years and never lost sleep over it. Then she ran the agency's first major AI acquisition: a $4.2 million vendor system to predict bridge maintenance needs. The proposal was glossy, the demo dazzling. Six months after award, the agency's engineers asked the vendor a simple question, how does the model decide a bridge is high-risk, and the vendor refused to answer, citing trade secrets. They could not test it for bias, could not audit it, could not even confirm what data it was trained on. The agency had paid millions for a black box it was contractually forbidden to open. "I bought it like it was a copier," Angela said afterward. "I should have bought it like it was a decision-maker."

The fix was not technical. It was in the contract she never wrote. OMB Memorandum M-24-18 is the procurement companion to the federal government's broader AI governance direction, focused specifically on acquisition, meaning how agencies buy AI. Its core insight is that most government AI is not built in-house; it is purchased. So the moment of real control is the contract. The memo pushes agencies to bake governance requirements, transparency, testing rights, performance standards, ownership, into the acquisition itself, before money changes hands. For a procurement officer or program manager this reframes the job: you are not buying a product, you are buying a governable capability, and the contract is where governance either exists or does not.

Why the Contract Is the Only Real Lever

Angela's hard lesson was about timing. Once the contract was signed, her leverage was gone. The vendor had her money and no obligation to be transparent, because transparency was never required in writing. In government acquisition, the requirements you put in the solicitation are the requirements you get. Everything else is a hope. This is why the memo puts governance at the front of the process: by the time a system is deployed and failing, it is far too late and far too expensive to add protections a clause would have secured for nothing.

The underlying reason AI procurement differs from traditional IT procurement is worth stating plainly. Traditional IT procurement focuses on functionality, schedule and cost, and those three are largely observable at delivery. AI procurement adds a dimension that is not observable at delivery: whether the system is safe, fair, secure and explicable. A vendor who builds accurate systems is not enough. They must build responsible ones, and responsibility is a property you can only verify if you negotiated the right to look. With purchased AI, your governance is only as strong as your contract. After award you do not have a model. You have whatever rights you remembered to negotiate.

The Four Vendor Obligations to Demand

Angela rebuilt her approach around four categories of obligation that her black-box bridge contract lacked entirely. These become requirements in the solicitation and clauses in the award.

  • Transparency and documentation. The vendor must disclose, in usable terms, what the system does, what data it was trained on, its known limitations, and how it performs. Not the secret recipe, but enough to govern it.
  • Testing and audit rights. The agency, or an independent third party, must have the contractual right to test the system, including for bias and accuracy by group, before and during use. This is the single clause that would have saved Angela.
  • Performance and accountability standards. The contract must define what acceptable performance means in numbers and what happens when the system falls short, including remedies and the agency's right to stop using it.
  • Data rights and exit. The agency must know who owns the data and the outputs, that its data will not be used to train other customers' models without permission, and that it can leave with its data intact.

What the Requirements Actually Cover

Underneath those four obligations sits a more granular set of expectations, and the difference between an enforceable contract and a decorative one is whether the granularity made it into writing. The requirements group into four families.

Transparency and documentation. The vendor should explain what the system does and where its limits are; what data it uses and to what quality standard; how it was tested for accuracy and fairness; what safeguards are in place; and how it is monitored in production. In the contract, specify exactly what documentation you expect, when it is due, and in what format. "Vendor shall provide an Algorithmic Impact Assessment before system deployment" is a requirement. "Vendor shall document system" is a conversation you will lose later.

Rights and safety safeguards. For systems affecting people's rights, the vendor should implement meaningful human review, fairness and bias testing, data quality standards, explainability mechanisms, and audit trails for decisions. In the contract, name which systems trigger these and define the terms rather than assuming shared meaning. "Meaningful human review" is a phrase that hides at least three different designs: review before the decision, review after it, or statistical sampling of a proportion. Pick one, write it, and require evidence that the safeguard works rather than evidence that it exists.

Security and supply chain. The vendor should address how the model is developed securely, how data poisoning is prevented, how the model is protected against theft, how third-party dependencies are managed, and how security incidents are handled. In the contract, require security certifications such as ISO 27001, set incident response commitments, and require documentation of the supply chain for any third-party component. Notice that the last one is the hardest to obtain and the most revealing: a vendor who cannot enumerate what its system is built from does not fully know what it sold you.

Monitoring and accountability. The vendor should enable monitoring rather than merely permit it: access to performance metrics, alerts when performance degrades, data on how system decisions compare against expectations, detailed audit logs, and support for government audits and testing. In the contract, specify which metrics, how often they must be reported, whether you need dashboard or programmatic access, and who bears the cost of an audit. Vague monitoring language produces a monthly slide deck. Specific monitoring language produces data.

Turning Requirements Into Clause Language

The gap between knowing what you want and writing it is where most AI procurements fail. The clause framework below is drafting scaffolding, not a legal instrument; your contracting shop adapts it to standard acquisition format, and your counsel reviews it. Where a bracketed or placeholder value appears, that is the agency's own number to set, and setting it deliberately is part of the exercise.

System documentation. The vendor shall provide, within 30 days of contract award, a system description covering what problem it solves; data requirements and quality standards; accuracy, fairness and robustness metrics; a list of all third-party components; testing and validation procedures; and deployment and monitoring procedures. The delivery deadline is the part that does the work. Documentation owed "as required" arrives when the vendor has time; documentation owed within a stated window after award arrives while you still have unspent money and therefore still have leverage.

Fairness and rights safeguards. For systems identified as affecting civil rights, the vendor shall conduct fairness testing before deployment; document the methodology and results of that testing; implement meaningful human review for X% of decisions; provide explainability for individual decisions; maintain audit logs of all decisions; and monitor fairness metrics after deployment. Two of those obligations are frequently traded away in negotiation and should not be. Post-deployment fairness monitoring is what catches the drift that pre-deployment testing cannot see, and the audit log is the only artifact that lets anyone reconstruct a contested decision a year later. A vendor willing to test once but not to monitor has offered you a snapshot of a system that will keep changing.

Data management. The vendor shall document all data sources and their quality; implement data quality monitoring; document any data limitations or biases; implement access controls and encryption; maintain data retention schedules; support government access for auditing and testing; and notify the government within 24 hours of a data breach. Retention schedules deserve a second look before signature, because an AI vendor's default is often to keep everything indefinitely, which is convenient for model improvement and directly at odds with the records and privacy obligations your agency already carries.

Security. The vendor shall maintain SOC 2 certification or equivalent; implement model versioning and rollback procedures; maintain detailed audit logs of all model changes; conduct annual penetration testing; implement incident response procedures with 4-hour notification to the government; maintain documented supply chain for all third-party components; and refrain from using the model for purposes other than the contracted use. Model versioning and rollback are the two that engineers ask for and contracts routinely omit. Without versioning, the agency cannot say which model produced a given decision; without rollback, a bad update has no safe reversal and the only available remedy is to stop using the system entirely.

Monitoring and reporting. The vendor shall provide daily reports on system accuracy, performance and errors; alert the government immediately if accuracy drops below X%; provide a monthly fairness metrics report; provide a quarterly deep-dive analysis of system performance; maintain a real-time dashboard of key metrics accessible to the government; and support government testing and auditing of the system. Set the alert threshold yourself rather than accepting the vendor's, and set it before you have seen the system's actual performance, because a threshold chosen after the fact will always be chosen just below whatever the system happens to do.

One structural point applies to that whole family of clauses. A dashboard is a convenience; the underlying obligation is access to the data behind it. Vendors will sometimes offer a rich dashboard in place of a data feed, and a dashboard cannot be independently analysed, cannot be joined to your own case records, and disappears at the end of the contract. Ask for the metrics in a form your analysts can work with, and treat the dashboard as the presentation layer rather than the deliverable.

Accountability and remedies. If the vendor fails to meet its obligations, the contract should provide notice and an opportunity to cure within X days; permit the government to reduce payment; permit the government to terminate; keep the vendor liable for harms caused by system failures; and require the vendor to maintain insurance covering AI system liability. Read that last pair carefully. Vendor liability for harms and vendor insurance are terms you negotiate, not rights you hold by default, and a contract that is silent on them has allocated that risk to you.

Remedies also need to be proportionate to be usable. Termination is the strongest instrument and frequently the least available one, because an agency that has built a service around a system cannot simply switch it off without harming the people it serves. That asymmetry is why the intermediate remedies matter: a payment reduction, a cure period with a defined clock, and an explicit right to suspend use of a particular function are all things you can actually invoke on a Tuesday. Write the ladder, not just the top rung, and make sure at least one rung is something your program office would genuinely be willing to pull.

Building Governance Into the Acquisition Lifecycle

Angela mapped where each obligation has to enter the process, because a clause added late is a clause fought over. AI procurement governance runs across four stages.

  1. Planning. Before writing the solicitation, classify the AI's risk. Will it affect people's rights or safety? Angela's bridge tool informed maintenance decisions affecting public safety, and that classification sets how stringent the requirements must be.
  2. Solicitation. Write the obligations into the requirements and the evaluation criteria, so vendors compete on governability rather than only price and features. If transparency is not scored, vendors will not offer it.
  3. Evaluation and award. Score proposals against the governance requirements. A dazzling demo from a vendor who will not allow testing should lose to a plainer system that will.
  4. Administration. After award, exercise the rights you negotiated: run the tests, hold the performance standards, monitor over time. Rights unused are rights wasted, and an audit right that is never exercised is functionally identical to one you never obtained.

Evaluating Vendors on More Than the Demo

Evaluation for AI acquisitions runs across three dimensions, and the weakest one usually predicts the outcome better than the strongest. Technical evaluation asks whether the system meets accuracy requirements, whether the vendor can demonstrate fairness testing rather than assert it, whether they hold security certifications, what comparable systems they have delivered, and how robust their own testing regime is. Proposal evaluation asks whether the approach is sound, whether they actually understand the requirements, whether the timeline is realistic, whether they identified risks and mitigations without being prompted, and whether they have the staffing to deliver. Past performance asks whether they delivered on previous AI contracts, whether references speak well of their fairness practices, whether they have had security incidents and how they responded, and whether they document lessons learned.

Those dimensions have to be weighted, and the weighting is a policy statement whether or not you intend it as one. A scoring matrix might set technical capability at 40%, proposed approach at 30%, cost at 20% and past performance at 10%, which sums to a complete allocation. What matters is not the specific split but that you chose it deliberately and can defend it, because the split is what determines whether a cheap, opaque system can beat an expensive, testable one. Weight the factors for your mission, write the weights into the solicitation, and score against them consistently.

Reference calls are the most underused instrument in this list. Past performance questionnaires produce polite answers; a conversation produces the useful one, which is what happened the first time the previous customer asked to test the system. Ask that question directly. The answer distinguishes a vendor whose transparency commitments are cultural from one whose commitments are contractual, and you want the first kind bound by the second.

An AI Procurement Governance Clause Checklist

Angela distilled her rebuild into a checklist she now attaches to every AI solicitation. Each item is a requirement to include and later a clause to enforce. The bracketed phrasing is plain-language clause language her contracting shop adapts to standard federal acquisition format.

  1. Risk classification. Is the AI's risk tier determined, and do the requirements scale to it, strictest for rights- or safety-affecting systems?
  2. Transparency clause. "The Contractor shall provide documentation of the system's intended use, training-data sources and characteristics, known limitations, and performance metrics in a form the Agency can independently review."
  3. Audit and testing clause. "The Agency or its designated third party shall have the right to test the system, including for accuracy and disparate performance across groups, before deployment and at defined intervals thereafter."
  4. Bias and disparity standard. Is the vendor required to test for and disclose performance differences across affected groups, with the thresholds defined in the contract rather than left to the vendor?
  5. Performance and remedy clause. Are acceptable-performance levels defined numerically, with stated remedies, a cure period, and the Agency's right to suspend use if standards are not met?
  6. Data rights clause. "The Agency retains ownership of its data and outputs; Contractor shall not use Agency data to train models for other customers without written consent."
  7. Exit and portability clause. Can the Agency terminate, retrieve its data in usable form, and avoid being locked in by proprietary formats or exclusive interfaces?
  8. Human-oversight requirement. Does the contract require the system to support meaningful human review for any decision affecting a person, with that phrase defined rather than assumed?
  9. Security and supply chain. Are certifications, incident notification timelines, model versioning and third-party component documentation all required in writing?
  10. Monitoring obligations. Are the metrics, reporting frequency, alert conditions and access method specified, rather than left as a general duty to cooperate?
  11. Liability and insurance. Does the vendor remain liable for harms caused by system failures, and is it required to carry insurance covering that liability?
  12. Evaluation weighting. Are these governance terms scored in source selection, so vendors compete on them?

Scaling the Process to the Size of the Buy

The same principles apply at every size; the apparatus should not. Two contrasting acquisitions show the range.

A small, low-risk buy. An agency needs automated document classification. Requirements: accuracy above 95%, handles all document types, logs all classifications. The contract is a five-page statement of work carrying the key requirements. Evaluation splits technical approach at 40%, cost at 30% and past performance at 30%. Monitoring is monthly reports, a quarterly review, and immediate escalation if accuracy drops below 95%. Ongoing management costs roughly ten hours a month. It stays simple because the system is not high-risk and nothing it does lands on a person's rights.

A large, high-impact buy. A federal agency needs a benefits eligibility system affecting hundreds of thousands of citizens. Requirements run to a forty-page detailed specification. Evaluation splits technical at 40%, approach at 30%, past performance at 20% and cost at 10%, deliberately pushing cost down the list. The contract carries fifty pages of terms and conditions. Monitoring means weekly technical meetings, monthly compliance reports and quarterly audits, with a 24-hour response commitment for critical issues and an executive steering committee above it. Ongoing management costs roughly forty hours a month. It is complex because the system decides who receives a benefit.

The comparison is worth internalising because the wrong lesson is easy to draw. The second acquisition is not better governed because it is bigger. It is governed proportionately, and the small one is governed correctly too. What would be wrong is running the light process on the heavy system, which is the most common and most consequential procurement error in government AI.

Contracting Does Not End at Award

The rights you negotiated only become governance when someone exercises them, and that requires a management rhythm rather than good intentions. A kickoff meeting establishes that the vendor understands the requirements, clarifies expectations on documentation, reporting, safeguards and escalation, and sets the communication cadence. Quarterly business reviews cover progress against schedule, system performance, emerging risks, open issues and next-quarter plans.

Monthly technical reporting from the vendor should carry system performance metrics, fairness and accuracy reports, security status and incidents, training and support metrics, and planned updates and changes. That last item matters more than it looks: a model update is a new system, and a change log that arrives after deployment rather than before it defeats every test you ran. Ad-hoc escalations need a defined clock, commonly 24 to 48 hours depending on severity, with critical issues such as security breaches or major accuracy loss requiring immediate notification. An annual compliance audit then checks the whole contract: test the systems, review the documentation, assess whether the safeguards work, and plan improvements for the following year.

Working With Vendors, Not Against Them

Angela worried these demands would scare off good vendors. Mostly they did not. Reputable AI vendors expect transparency and testing requirements and can meet them, and a vendor who refuses outright has told you something valuable before award rather than after. Be careful with the inverse, though: agreement is not assurance. A vendor who signs an audit clause has given you a right, not a result, and the only thing that converts one into the other is running the audit. Treat a refusal as disqualifying and an acceptance as the beginning of the work.

She also learned to use the existing acquisition system rather than fight it. The federal acquisition rules already let agencies define requirements and evaluation factors, so AI governance clauses fit inside the normal process her contracting shop already runs. And she leaned on colleagues, the privacy, civil-rights and technical staff who could tell her what to actually test for, drafted into the acquisition team early rather than consulted after the solicitation closed. The frameworks her agency already used for risk, the voluntary NIST AI Risk Management Framework chief among them, gave her the vocabulary to specify what governable meant in a clause.

The deepest shift was cultural. Angela stopped seeing procurement as the paperwork after the real decision and started seeing it as the decision, the one place where an agency, before spending a dollar, can require that the AI it buys can be opened, tested and trusted.

Anti-Patterns

  • Running an AI buy on a traditional IT template. The contract comes out silent on fairness, transparency and monitoring, and the vendor is therefore not accountable for any of them. Nothing in the process flags the omission, because the process was designed for products whose behavior is fully observable at delivery. Add the AI-specific language explicitly to every AI acquisition, and require vendors to demonstrate fairness and security practices rather than assert them.
  • Awarding on lowest price. Price is the easiest factor to compare and the least informative about whether a system will behave. A vendor bidding low has to find the savings somewhere, and fairness testing, security work and documentation are exactly the line items that are invisible in a demo. Use weighted evaluation criteria, weight technical capability and past performance to match the risk, and remember that a low price on an unusable system is not a saving.
  • Locking yourself in. Custom model weights, proprietary data formats and exclusive interfaces make switching vendors impractical, and impracticality after award is indistinguishable from being trapped. Require access to model weights, data and code; use standard file formats; establish data portability requirements; and include exit strategies in the contract, when you still have the leverage to ask for them.
  • Treating a signed audit right as an audit. The clause creates the possibility of oversight, not oversight itself, and a right that is never exercised produces exactly the same evidence as no right at all. Schedule the first test before award so it is already in someone's calendar, and treat a year with no exercised audit rights as a finding about your own contract administration.
  • Leaving "meaningful human review" undefined. The phrase is agreed to easily because everyone reads it differently. Review before the decision, review after it, and sampling a fraction of decisions are three different controls with three different costs, and the vendor will implement the cheapest reading. Define it in the contract, state the proportion, and require evidence that reviewers can and do overturn outcomes.
  • Accepting documentation that is really marketing. A glossy model overview satisfies a vague documentation clause and tells an auditor nothing. Specify the content, the format and the due date, and check on receipt that what arrived would let an independent reviewer form a judgment about the system rather than a favorable impression of the vendor.
  • Bringing privacy and civil-rights staff in after the solicitation closes. By then the requirements are fixed and their expertise can only produce objections. Draft them onto the acquisition team at planning, when what they know can still become a clause.

Practice Prompts

  • Define your requirements before you look at any vendor. Write down what your agency actually needs on fairness, security, documentation and monitoring for one upcoming acquisition, and notice which of those four you find hardest to state precisely.
  • Draft clause language for two or three of the contract sections in this lesson, adapted to your mission area. Then hand the draft to your contracting officer and ask which parts would survive review, and why.
  • Build a weighted scoring matrix for a real or planned AI acquisition. Set the weights, then run a thought experiment: which vendor would win if the cheapest bidder scored poorly on transparency? If the answer is the cheapest bidder, your weights do not say what you meant.
  • Write the vendor management plan. Which metrics will you require, at what frequency, delivered how, and who on your side reads them? A metric nobody is assigned to read is a metric that will not be produced for long.
  • Design your lock-in protections. What would it actually take to move off this system in two years? Write the answer down before award, and turn each obstacle you identify into a clause.
  • Take an AI contract your agency already holds and audit it against the twelve-item checklist. Mark each item present, partial or absent, and note which absences could still be fixed at the next option year.
  • Call a reference for a vendor you are considering and ask one specific question: what happened the first time you asked to test the system yourself.

Reflection

Think about the AI system your agency depends on most heavily today. If you asked its vendor tomorrow for the training data characteristics, the group-level accuracy breakdown and permission for an independent test, what would happen? Not what should happen, what would. Then read the contract and find out whether the answer you expect is the answer you are entitled to. The distance between those two things is your governance gap, and it was set on the day of award by whoever wrote the requirements. The useful follow-up question is who that person was, and whether they had anyone from privacy, civil rights or engineering in the room with them.

Glossary

  • OMB Memorandum M-24-18: The federal memorandum addressing AI acquisition, which sets expectations that agencies translate into contract requirements when buying AI rather than building it.
  • Meaningful human review: A process requiring a human to review important AI decisions. The term covers several distinct designs, review before a decision, review after it, or sampling, and must be defined in the contract rather than assumed.
  • Fairness testing: Systematic testing to identify and measure bias in an AI system, including performance broken out across affected groups rather than reported only in aggregate.
  • Audit trail: A detailed record of system decisions, retained so that accountability questions and later audits can be answered from evidence.
  • Algorithmic Impact Assessment: A documented assessment of a system's likely effects, commonly required as a named deliverable before deployment.
  • Vendor lock-in: The situation where switching vendors is impractical or prohibitively expensive, usually created by proprietary formats, exclusive interfaces or inaccessible model artifacts.
  • Supply chain risk: The risk that third-party components or data incorporated into a system could compromise it, which is why component inventories are a contract deliverable.
  • Model weights: The numerical parameters that define how an AI system makes decisions, and one of the assets whose accessibility determines whether you can leave a vendor.
  • Data portability: The ability to move data between systems and vendors in a usable form, secured by clause rather than by goodwill.
  • Cure period: The contractual window a vendor is given to fix a failure before the government exercises remedies such as payment reduction or termination.
  • Source selection: The evaluation and award process in which proposals are scored against stated criteria, and therefore the only place governance terms can be made competitive.

Closing

AI procurement is the work of translating policy expectations into contract language and then making sure vendors understand and meet what was written. The specific terms depend on your mission, your budget and your risk tolerance, but the core principles do not vary: transparency, fairness, security, monitoring and accountability. Good procurement practice makes everything downstream easier. Poor procurement practice creates problems that outlast the people who created them, because a contract signed today constrains an agency for years.

Angela's bridge system is still running. She could not renegotiate the trade secret provisions mid-term, so the agency lives with a system it cannot fully inspect until the contract ends, and the engineers work around it. Every acquisition since has carried the checklist. That is the honest shape of this lesson: the clause you forget is not a mistake you fix, it is a condition you inherit. Procurement is where policy becomes practice, and it is the one moment when an agency holds all the leverage it will ever hold.

Key Takeaways

  • The contract is the lever. Most government AI is bought rather than built, so governance lives in the acquisition. What you require in the solicitation is what you get; everything else is a hope.
  • M-24-18 is the procurement companion. Its subject is acquisition, meaning how agencies buy AI and what they must require of vendors, which is why the whole lesson lands on clause language rather than on system design.
  • AI procurement differs because the important properties are not observable at delivery. Functionality, schedule and cost you can see. Safety, fairness, security and explicability you can only verify if you negotiated the right to look.
  • Demand four vendor obligations. Transparency and documentation, testing and audit rights, defined performance standards with remedies, and data rights with an exit.
  • Audit and testing rights are the saving clause. The contractual right to test a system, including for bias by group, is what turns a black box into a governable capability, and it costs nothing before award.
  • Specificity is the whole game. "Vendor shall provide an Algorithmic Impact Assessment before deployment" is enforceable. "Vendor shall document system" is not. Define undefined terms, especially meaningful human review, and state the numbers.
  • Score governability in source selection. If transparency and testing are not evaluation factors, vendors will not offer them. Set the weights deliberately and check that they would actually let a testable system beat a cheaper opaque one.
  • Liability and insurance are negotiated, not automatic. Vendor liability for harms caused by system failures and vendor insurance covering AI liability are contract terms. A contract silent on them has assigned that risk to the agency.
  • Scale the process to the risk, not to the budget. A small low-risk buy is correctly governed by a short statement of work and monthly reports. Running that light process on a benefits eligibility system is the most consequential error in government AI procurement.
  • Tough clauses filter rather than deter, and acceptance is not assurance. Reputable vendors meet these terms and the ones who refuse reveal risk before award. But a signed audit right is only a possibility of oversight until someone runs the audit.
  • Procurement is the decision, not the paperwork. Classify risk in planning, embed clauses in the solicitation, score them at award, and exercise them during administration. Rights unused are rights wasted.

Frequently Asked Questions

The vendor says its model is a trade secret. Are we stuck?

Before award, no. Transparency requirements do not oblige a vendor to hand over the secret recipe; they oblige disclosure of intended use, training-data sources and characteristics, known limitations and performance metrics in a form the agency can independently review, which is a different and much smaller ask. A vendor who cannot meet that has told you something useful. After award, you are largely stuck with whatever you wrote, which is exactly what happened to Angela, so the answer to this question is decided months before anyone asks it.

Will these requirements shrink our competitive field?

Some, and that is mostly the point. Vendors with mature AI practices expect transparency and testing terms and can satisfy them, so the field narrows toward suppliers who can actually govern what they sell. Where it becomes a real problem is with small or novel suppliers who could deliver good work but cannot yet meet enterprise security certifications; the remedy there is scaling requirements to risk, not dropping them, so that a low-risk buy carries a proportionate burden.

Do we need all of this for a small, low-risk purchase?

No, and applying it uniformly is its own failure mode. A low-risk document classification tool is properly governed by a short statement of work, a stated accuracy requirement, monthly reports and an escalation trigger. Reserve the forty-page specification and the quarterly audit for systems that touch people's rights or safety. The error to avoid is the reverse one: running the light process on the heavy system because it moves faster.

What if we already signed a contract without these clauses?

Work the openings you have. Option years, modifications and renewals are all points where terms can be added, and a vendor who wants the follow-on has a reason to negotiate. In the meantime, document what you cannot verify, so the gap is recorded rather than assumed away, and build the testing you can do from the outside, on the outputs, since output-level analysis needs no permission. Then make sure the next solicitation carries the clauses from the start.

The vendor agreed to our audit clause. Are we covered?

You have a right, not a result. An audit right that is never exercised produces the same evidence as no audit right at all, and vendors are not obliged to remind you it exists. Schedule the first test before award so it lands in a calendar rather than a good intention, define who pays for it in the contract, and treat a year in which no negotiated right was exercised as a finding about your own contract administration rather than a sign that everything is fine.

Who should be on an AI acquisition team?

The contracting officer and program manager, plus privacy, civil-rights and technical staff, and they need to be there at planning rather than at review. The people who can tell you what to actually test for are the only ones who can turn a governance intention into a requirement, and once the solicitation closes their expertise can only produce objections. Angela's single most useful change was moving that conversation from after the draft to before it.