Writing AI Requirements in RFPs and SOWs
Denise Okoro, a contracting officer at a federal grants agency, learned the hard way what a vague requirement costs. Her last AI procurement, a document-classification system, had a statement of work saying the contractor would deliver "an accurate, fair, AI-powered solution." Eighteen months and $2.3 million later, the system sorted grant applications at 71 percent accuracy, the vendor refused to explain how it decided anything because the model was "proprietary," and no clause required them to fix the racial skew an auditor found. The contract was technically fulfilled. Denise had written a wish, not a requirement.
Drafting the request for proposals for a replacement tool, she set herself one test for every line: could a third party read this sentence, run a test, and say yes or no? The core principle is that simple. If you cannot test it and enforce it, it is not a requirement. "Accurate" is a wish. "At least 90 percent precision on the government's held-back test set, measured at acceptance and quarterly thereafter" is a requirement. Your requirements are the contract between you and the vendor, and everything downstream, proposal evaluation, negotiation, performance monitoring, and dispute resolution, gets easier or harder based on how well you wrote them.
Why AI Requirements Are Different
Traditional software requirements focus on functionality. "The system shall process 1,000 transactions per second" is testable, unambiguous, and complete. You either hit the number or you do not. AI requirements have to carry the same testability while covering behavior that is statistical rather than deterministic, that varies across the population the system touches, and that decays over time without anyone changing a line of code.
That adds layers a traditional specification never had to address. Transparency: can the system explain its decisions, and to whom? Fairness: does it treat different groups equitably, measured how? Data: what data is used, where did it come from, and how is it validated? Testing: how is accuracy verified, and how is bias tested? Monitoring: how will anyone know if performance degrades after go-live? Writing effective AI requirements means being specific about all of these, not just the first one. Skip a layer and you create the gap a vendor will use against you later.
There is a second reason to take the drafting seriously, and it operates before a single proposal arrives. The quality of your solicitation directly determines the quality of the proposals you receive. Poor requirements attract poor solutions, because bidders can only respond to what you asked for. Vague requirements attract disputes, because every bidder resolves the ambiguity in the direction that favors their own product. At this level of the profession the work is translating policy and program needs into procurement language that shapes vendor behavior, which means the document is a design decision about the system, not paperwork that follows one.
These map to federal guidance you already operate under, including OMB Memorandum M-24-10, "Advancing Governance, Innovation, and Risk Management for Agency Use of Artificial Intelligence," issued in 2024, and the NIST AI Risk Management Framework, which is voluntary rather than binding. You do not need to cite either one in the contract. You need to translate them into terms a test can settle. For the acquisition rules that govern the vehicle itself, see Federal Acquisition of AI: FAR/DFARS; for the procurement governance layer above the individual solicitation, see OMB M-24-18 and AI Procurement Governance.
The Five Pillars of an AI Requirement
Every enforceable AI specification covers five areas. Denise built her second solicitation around them, and each one closed a gap the first solicitation had left open.
1. Capability: what it must do, in measurable terms
Describe the task, the inputs, and the outputs concretely. Not "process applications" but "ingest PDF grant applications up to 80 pages, extract 14 named fields listed in Attachment B, and classify each into one of six program categories." A reader should be able to picture the system running. If two engineers at two different vendors would build materially different systems from your paragraph, the paragraph is not finished.
2. Performance: how well, on whose data
State the metric, the threshold, and the test set. The trap Denise fell into was letting the vendor grade their own homework. Require performance to be measured on a government-held-out test set the vendor never sees during training or tuning. Specify both error types, because a system can hit 95 percent overall and still fail the cases you care about most. State latency separately with the ceiling your operation actually needs, since a correct answer that arrives too late is a failed answer.
3. Fairness: equitable performance across groups
Require the vendor to measure and report performance broken out by relevant population groups, and set a limit on the gap between them. This is where most "fair AI" language is empty. Fairness means different things to different people, so a bare instruction to be fair invites the vendor to implement one metric while you expected another. Specify exactly which fairness metrics matter, which groups, and what threshold. Make it numeric and testable or leave it out and admit you did not specify it.
4. Transparency and documentation: what you get to see
"Proprietary" is not automatically a wall against accountability, but it becomes one if you did not reserve the right in writing before award. Require model documentation, often called a model card, an explanation method for individual decisions, and an inspection right for the government and any auditor it designates. You can protect a vendor's genuine trade secrets and still see what you need to see; what you cannot do is negotiate that access after a dispute has already started.
5. Monitoring and remediation: what happens after go-live
Models drift. The world changes and accuracy decays even when nothing about the system changes. Require ongoing monitoring, scheduled re-validation, notification when performance drops, and a defined fix-it obligation with a clock on it. This is the clause Denise was missing entirely, which is why her auditor's finding turned into a negotiation instead of a corrective action.
The Two Families the Pillars Do Not Cover
Alongside the five pillars sit two more requirement families that solicitations routinely forget until an assessment or an authorization stalls. The first is data requirements. Name the approved data sources the system may use, the quality standards and specific checks the data must pass, and the retention period for both input data and anything the system derives from it. An AI system that quietly retains constituent records past your retention schedule is a records problem, not just an engineering one.
The second is security requirements. State the cybersecurity standards that apply, the access controls governing who may modify the model, and the incident response obligation including the notification timeline. These are not AI-flavored nice-to-haves; they are the terms your security office and your authorizing official will ask for, and adding them by modification after award costs more than including them in the solicitation. Where the system will touch regulated data, name that in the requirement rather than assuming the vendor will infer it from context.
A Structure Every Requirement Should Follow
Consistency in structure is what makes a requirements document reviewable by people who did not write it. Give every requirement a number, a category, a statement, a success criterion, and a verification method. The last field is the one most often left blank, and it is the one that decides whether you have written a requirement or an aspiration. If you cannot describe how you will verify something, you have not specified it.
| Field | What it holds |
|---|---|
| Requirement Number | [REQ-AI-001] |
| Category | Functionality, Performance, Fairness, Data, Testing, or Monitoring |
| Statement | The clear, specific requirement |
| Success Criteria | How you will know it is met |
| Verification Method | How you will test it |
Worked through, a fairness requirement in that structure reads as follows. Requirement REQ-AI-001, category Fairness. Statement: the system shall achieve at least 92% accuracy across all demographic groups. Success criteria: accuracy tested separately for each demographic group defined in the data specification, with all groups achieving at least 92% accuracy. Verification method: the vendor provides fairness testing results comparing accuracy across demographic groups, and the government independently validates on a hold-out test set. Note that the verification field does two things at once. It commits the vendor to produce evidence, and it reserves the government's own independent check.
Numbering by family keeps a long document navigable and makes gaps visible at a glance. The bracketed prompts below are the agency's own values to fill in, not omissions; leave them bracketed in your template so a reviewer can see what has not yet been decided.
| Family | Example entries |
|---|---|
| Fairness (Fair) | Fair-1: System shall achieve within X% accuracy across all demographic groups; Fair-2: System shall provide explanations for individual decisions; Fair-3: System shall document known limitations |
| Data (DR) | DR-1: Data sources: [List approved sources]; DR-2: Data quality standards: [Specific checks]; DR-3: Data retention: [How long will data be kept?] |
| Testing (TR) | TR-1: Accuracy testing methodology: [Describe]; TR-2: Fairness testing: [Which metrics? Which groups?]; TR-3: Robustness testing: [Edge cases to test] |
| Monitoring (MR) | MR-1: Performance metrics to be reported: [List]; MR-2: Reporting frequency: [Monthly? Daily?]; MR-3: Alert thresholds: [When should vendor alert agency?] |
| Security (SEC) | SEC-1: Cybersecurity standards: [NIST, ISO, etc.]; SEC-2: Access controls: [Who can modify the model?]; SEC-3: Incident response: [Notification timeline?] |
From Vague to Specific: One Requirement, Three Drafts
The distance between a poor requirement and a good one is easiest to see side by side. Take a permit classification system. The poor version reads: "The system shall classify applications correctly." Every vendor reading that will price a different system, and the ones who price it honestly will lose to the ones who do not.
A good version names the categories and the threshold: "The system shall classify incoming permit applications into one of five categories: Building, Environmental, Safety, Zoning, Other. Classification shall be correct at least 95% of the time." That is testable. A reviewer can construct the test from the sentence alone, and two vendors would build recognizably similar systems from it.
The better version adds the operational reality around the classification. The system shall classify permit applications into Building Permits, Environmental Review, Safety Compliance, Zoning Variance, and Other; accept applications in PDF, Word, and plain text formats; extract key information including applicant name, property address, application type, and requested action; classify into the correct category with at least 95% accuracy based on content and metadata; provide confidence scores for each classification; route applications to the appropriate department based on classification; and support human override when a classification is incorrect. Seven clauses, each independently verifiable, and the last one is the clause that keeps a person in the loop.
The performance side of the same specification states each metric with its plain-language definition, so that nobody argues about vocabulary at acceptance: overall accuracy at or above 95%, meaning the percentage of correct classifications; precision for each category at or above 90%, meaning of the items the system predicted for that category, how many actually belong there; and recall for each category at or above 85%, meaning of the items that actually belong to that category, how many the system found. Defining the terms inside the document costs three sentences and prevents an entire class of dispute.
Sample Clause Library
Adapt these to your acquisition. Bracketed values are the levers you set with your program office and technical evaluators, and the figures shown are illustrative rather than recommended. Keep the numbers realistic for your use case and your risk level, and be prepared to defend each one, because a threshold you cannot justify is a threshold a vendor will negotiate away.
Capability clause
The Contractor shall deliver a system that [extracts the 14 data fields specified in Attachment B] from [PDF and TIFF documents up to 80 pages] and assigns each document to one of [six categories defined in Attachment C], producing for each output a confidence score and the source location of each extracted field.
Performance clause
System performance shall be measured on a government-controlled test set withheld from the Contractor during all training and tuning. The system shall achieve no less than [90%] precision and [90%] recall on this set. Acceptance is contingent on meeting these thresholds. The Government will repeat this measurement [quarterly]; sustained performance below threshold for [two consecutive quarters] constitutes a material deficiency.
Fairness clause
The Contractor shall report precision, recall, false-positive, and false-negative rates disaggregated across [the population groups identified in Attachment D]. The absolute difference in [false-positive rate] between any two groups shall not exceed [5 percentage points]. The Contractor shall deliver this analysis at acceptance and [quarterly], and shall remediate any breach of this threshold within [60 days].
Transparency clause
The Contractor shall deliver model documentation describing training data sources, intended use, known limitations, and evaluation results. The Contractor shall provide a method to explain individual decisions in plain language sufficient for an affected member of the public to understand the basis of an outcome. The Government and its designated independent auditor shall have the right to inspect the model, its documentation, and its performance data. Assertions of proprietary information shall not limit the Government's audit and oversight rights under this clause.
Monitoring and remediation clause
The Contractor shall monitor system performance continuously and notify the Contracting Officer within [5 business days] of any sustained drop below the thresholds in this SOW. The Contractor shall provide remediation at no additional cost for deficiencies attributable to the delivered system, with a corrective action plan due within [10 business days] and resolution within [45 days]. The Government may suspend system use pending remediation of any fairness or safety deficiency.
Note that the two clocks in the library are deliberately different. The fairness clause allows a longer window because remediating a disparity usually means retraining, while the monitoring clause carries a shorter one because a performance drop is generally a fix rather than a rebuild. Set both consciously. A single blanket remediation period applied to every kind of deficiency is a sign nobody thought about what remediation actually involves.
Scale the Requirements to the Risk
Applying identical requirements to every AI acquisition is its own failure. An internal tool that drafts meeting summaries and a system that determines eligibility for benefits do not warrant the same testing regime, the same monitoring cadence, or the same documentation burden. Over-specify the low-risk system and vendors over-invest, prices rise, and small firms drop out of your competition entirely. Under-specify the high-risk one and you have written Denise's first solicitation again.
Set the tier before you draft. Systems that affect rights or safety need more stringent testing, more frequent independent validation, tighter fairness thresholds, and explicit human-override and appeal paths. Lower-risk systems can carry a simpler specification and a lighter reporting rhythm. Making the tier explicit in the solicitation also tells bidders how to price the work honestly, which improves the proposals you receive rather than just the contract you sign.
Scoring Proposals and Dodging Traps
Requirements only protect you if your evaluation rewards meeting them. Build them into the proposal evaluation criteria so vendors compete on the things that matter rather than on presentation quality, and get the requirements right before the solicitation goes out. Changing a specification during performance is expensive and slow; consulting end users, compliance, and security beforehand is neither.
- Demand a demonstration on your data, not their demo. Hand finalists a sample of your real, de-identified documents and score live performance. Vendors who can only show a polished demo reveal themselves quickly.
- Make vendors explain how they would meet each requirement. This is the cheapest test of your own document. If capable bidders read a requirement three different ways, the requirement is ambiguous and you found out before award instead of after.
- Reward documentation quality. Make the model card and the explanation method scored evaluation factors rather than afterthoughts nobody reads.
- Price the whole lifecycle. A cheap build with no monitoring obligation is expensive. Evaluate total cost including re-validation, monitoring, and remediation across the full period of performance.
- Watch for "we will meet any metric you set." A vendor who agrees instantly to every threshold without discussing tradeoffs either does not understand the problem or plans to argue definitions later. Good vendors push back thoughtfully, and that pushback is information.
- Beware the accuracy headline. "98 percent accurate" is meaningless without the test set, the population breakdown, and the error types. Require all three before you credit the claim.
- Remember that vendors optimize what you measure. Whatever metric you name will become the metric the delivered system is tuned for, including at the expense of the ones you left out. That is a reason to specify fairness, latency, and monitoring explicitly, not a reason to specify less.
Anti-Patterns
- Requirements too vague. "The system shall classify documents accurately." One bidder prices 85 percent, another prices 95 percent, and the disagreement surfaces after award. Define accuracy by category, by group, and by use case, and define the test procedure alongside it.
- Requirements without verification methods. The vendor commits to something you have no way to test. For every requirement, write down whether you will test it independently, audit the vendor's testing, or accept a certification, and who does the work.
- Fairness requirements without specificity. "The system shall be fair" is not a requirement. Name the metrics, the groups, and the acceptable threshold, or acknowledge internally that you have not specified fairness at all.
- One-size-fits-all requirements. Identical stringency for an internal tool and a benefits-eligibility system wastes money on one and under-protects the other. Scale with risk deliberately.
- Requirements changed during performance. You specify accuracy, the vendor optimizes for accuracy and meets it, and only then do you discover you needed fairness and latency too. Finalize requirements before issuing the solicitation, after consulting end users, compliance, and security.
- Treating a written clause as an outcome. A clause is leverage, not a guarantee. Enforcing it still takes a contracting officer willing to act, documented evidence, and time. Write the clause and also plan who monitors it, how often, and what the first escalation step is.
- Treating a passed acceptance test as proof the system works. Hitting a threshold on a held-out test set is evidence about that test set on that day. It is not a guarantee of field performance, which is exactly why the monitoring and re-validation clauses exist.
- Letting "proprietary" settle the question after award. Trade-secret claims are much harder to overcome when your solicitation reserved no inspection right. Reserve audit and oversight rights in the document, before anyone has an incentive to resist them.
Practice Prompts
- Identify what actually matters. For an AI system you expect to buy, rank accuracy, fairness, speed, explainability, and security against each other. Which one would you trade away first, and could you defend that ranking to the people the system affects?
- Define the metrics. For each requirement you ranked, write the metric, the threshold, and the test set. Where you cannot name a test set, you have found the requirement most likely to fail.
- Specify the fairness requirement. Which population groups are relevant to your system, which fairness metric matters for the harm you are worried about, and what gap between groups would you treat as a material deficiency?
- Write the verification column. Take five requirements and fill in only the verification method for each. Anything you cannot verify goes back for rewriting or comes out of the document.
- Draft the template. Produce 5 to 10 requirements using the number, category, statement, success criteria, and verification structure, covering at least functionality, performance, fairness, data, testing, monitoring, and security.
- Stress-test with a reader. Give the draft to someone who was not in the room and ask them to describe the system it specifies. Every difference between their description and yours is a defect you can still fix for free.
Reflection
- Look at your most recent AI solicitation. How many of its requirements could a third party test without asking you a clarifying question?
- What would you monitor in production, and does anything in your current contract obligate the vendor to report it to you?
- How specific can you honestly be given real uncertainty about the problem, and where does specificity turn into guessing at a number you cannot defend?
- If a vendor delivered exactly what your document requires and nothing more, would you be satisfied? If not, the gap is in the document.
- Who in your agency, other than you, has read the requirements and would notice a missing verification method?
Glossary
- Requirement. A specific statement of what the system must do or achieve, written so that a third party can test it.
- Functional requirement. A requirement about what the system does.
- Performance requirement. A requirement about how well the system performs, stated as a metric, a threshold, and a test set.
- Fairness requirement. A requirement about equitable treatment across groups, expressed as disaggregated metrics and a limit on the gap between them.
- Success criteria. The specific measure that determines whether a requirement is satisfied.
- Verification method. The procedure for testing whether a requirement is met, and the party responsible for running it.
- Acceptance testing. Testing performed to verify that a system meets its requirements before delivery is accepted.
- Held-out test set. Data controlled by the government and withheld from the contractor during all training and tuning, used so that performance is not measured on data the system has already seen.
- Precision and recall. Of the items the system assigned to a category, how many belong there; and of the items that belong to that category, how many the system found.
- Model card. Model documentation describing training data sources, intended use, known limitations, and evaluation results.
- Drift. Degradation of model performance over time as the world changes, even when the system itself is unchanged.
Related Lessons
- Federal Acquisition of AI: FAR/DFARS covers the acquisition regulations and clause selection that sit underneath the requirements you draft here.
- OMB M-24-18 and AI Procurement Governance is the acquisition and procurement governance companion to the requirement-writing craft in this lesson.
- Requirements Gathering for AI covers the upstream work of learning what the program office actually needs before any of it becomes contract language.
- AI Vendor Evaluation Methodology turns these requirements into a scoring approach for competing proposals.
- Evaluating AI Vendor Claims goes deeper on the accuracy headline problem and on testing what a vendor asserts.
- Managing AI Vendor Performance picks up after award, where the monitoring and remediation clauses either work or sit unused.
- Rights-Impacting and Safety-Impacting AI Safeguards explains the risk tiering that determines how stringent your requirements should be.
Closing
Denise's second solicitation ran 40 pages where the first ran 12. The extra pages were the five pillars turned into clauses, the two additional families her security and records staff insisted on, and a verification method attached to every requirement. Three vendors dropped out at the live-data challenge. The winner came in at 92 percent precision on her held-out set, above the threshold she had set, delivered a model card, and accepted both the fairness threshold and the remediation clock. When an auditor flagged an issue six months later, Denise pointed at the clause and the vendor fixed it at their own cost.
The difference between her two procurements was not budget, and it was not technology. It was whether the words on the page could be tested and enforced. That is entirely within a contracting officer's control, it costs time rather than money, and it is the cheapest risk reduction available in an AI acquisition. Invest the time upfront: poor requirements produce poor systems, and excellent requirements guide vendors toward building the thing you actually needed.
Key Takeaways
- If you cannot test and enforce it, it is not a requirement. Replace adjectives like "accurate" and "fair" with a metric, a threshold, and a named test set.
- Cover all five pillars, plus data and security. Capability, performance, fairness, transparency, and monitoring carry the AI-specific ground; data sources, quality, retention, and security controls are the two families most often left out until an authorization stalls.
- Every requirement needs a verification method. Number, category, statement, success criteria, and verification. The last field is the one that separates a requirement from an aspiration.
- Never let the vendor grade their own homework. Measure on a government-controlled set withheld during all training and tuning, and reserve your own independent validation in writing.
- Make fairness numeric. Name the metrics, the groups, and a hard limit on the gap between them, with a remediation clock attached.
- Reserve oversight rights before award. Trade-secret claims are far harder to overcome afterward. Inspection and audit rights belong in the solicitation, not the negotiation.
- Contract for the whole lifecycle. Models drift, so require ongoing monitoring, scheduled re-validation, notification, and a defined, time-bound remediation obligation at no additional cost.
- Scale requirements with risk. Rights-impacting and safety-impacting systems warrant stringent testing and monitoring; low-risk internal tools do not, and over-specifying them just narrows your competition.
- Vendors optimize exactly what you specify. That is a reason to name fairness, latency, security, and monitoring explicitly, and a reason to finalize the set before the solicitation goes out.
- Test on your own data before award. A live challenge with your real, de-identified documents separates working systems from good slide decks faster than any written proposal can.
Frequently Asked Questions
How do I pick thresholds when I have never bought this kind of system before?
Start from the operational consequence rather than from a benchmark. Ask what error rate your current manual process produces, what an error costs the person on the other end, and what volume you handle. Then set a threshold you can defend from those three facts, and treat the figures in any template, including the ones in this lesson, as illustrative rather than recommended. If you genuinely cannot justify a number, say so in the solicitation and require bidders to propose one with their reasoning, which turns your uncertainty into an evaluation factor instead of a gap.
Should I cite M-24-10 or the NIST AI Risk Management Framework in the statement of work?
You do not have to, and citing a framework is not a substitute for specifying behavior. The frameworks tell you which questions to ask; the contract has to carry the testable answer. Translate the concern into a requirement with a threshold and a verification method. Where your agency's own policy requires a specific reference, follow that policy, and use Federal Acquisition of AI: FAR/DFARS and OMB M-24-18 and AI Procurement Governance for what belongs in the acquisition document itself.
Will a strong requirements document guarantee I get what I paid for?
No. It gives you leverage and evidence, which is considerably more than Denise had the first time, but enforcement still depends on someone monitoring performance, documenting the shortfall, and being willing to act on it. Write the clause and then name the person responsible for checking it, the reporting cadence, and the first escalation step. A monitoring clause that nobody reads is functionally the same as no clause at all.
The vendor says our fairness and transparency requirements are impossible. Now what?
Distinguish informed pushback from a sales objection. A vendor explaining that your fairness metric conflicts with your recall floor at your data volume is giving you useful information, and the right response is to revisit the specification with your technical evaluators. A vendor asserting that explanations are simply not possible is telling you about the product they want to sell, not about the state of the field. Either way, resolve it before award, because the same conversation after award happens on their terms.
How much detail is too much?
Specificity that you can verify is never wasted; specificity you invented to look thorough is a liability, because a threshold you cannot justify is one you will quietly waive. Aim for a document where every number traces to an operational fact, every requirement carries a verification method, and every bracketed placeholder is a decision your program office has consciously deferred rather than a blank nobody noticed.
Skill.re