AI Vendor Evaluation Methodology
Priya Nair, a contracting officer at a county Department of Health, had three AI vendors pitching the same product: a tool to triage incoming benefit applications. All three demos looked dazzling. All three said "92% accurate." All three promised the moon. Priya's program director wanted to pick the one with the slickest interface. Priya knew that was how counties end up with a $400,000 contract for a system that fails its first audit. She needed a way to compare three confident sales decks on something other than vibes. She needed a methodology.
Why AI Procurement Is Not Like Other Procurement
Evaluating an AI vendor is not like buying laptops. With laptops, the specs tell the truth. With AI, the specs are a marketing claim, the risks are invisible until deployment, and the worst problems show up months later in a fairness complaint or a procurement protest. This lesson gives you a repeatable evaluation method built on an eligibility-first hierarchy, three scoring dimensions, a weighted scorecard, independent piloting, and a reference-checking routine designed for the way AI vendors actually behave.
Government AI acquisition also happens inside a squeeze that vendors understand better than most buyers do. Budget cycles impose deadlines. Leadership wants innovation quickly. Congressional or inspector general attention creates urgency. Vendors read those constraints accurately and answer them with optimistic timelines and impressive-sounding capabilities. A systematic evaluation methodology is what separates a successful acquisition from an expensive mistake, because it holds a fixed shape while the schedule pressure pushes on everything else around it.
Three differences from commercial buying matter most. First, government decisions are public and defensible: if your methodology is flawed, Congress, the Government Accountability Office, inspectors general and the public all have standing to scrutinize your choices. Second, government systems affect vulnerable populations. A benefits system serves people in need; a hiring system shapes career prospects. The stakes reach into people's lives in a way a commercial software purchase rarely does. Third, agencies carry accountability obligations to statutes, regulations and executive orders that commercial companies do not face. Civil rights compliance, FISMA security requirements and FOIA obligations are not optional add-ons; they are mandatory elements of any evaluation.
Your evaluation methodology is the record of whether you did your due diligence. If something goes wrong after deployment, the question you will be asked is what evaluation process you followed. If you can describe a rigorous, transparent and comprehensive methodology, you have an answer. If you cannot, you are exposed, and no amount of after-the-fact reconstruction will fix that.
The Five Tensions Your Method Has to Navigate
Every evaluation framework balances five competing demands, and pretending they do not compete is how methodologies quietly collapse under deadline. Speed, set by the budget cycle and leadership expectations, pulls against rigor in accuracy, fairness and security. Simplicity, meaning an evaluation your team can actually execute, pulls against comprehensiveness in covering all material risks. Technical depth in data science assessment pulls against stakeholder accessibility for non-technical leadership. Innovation in exploring new approaches pulls against caution in protecting against downside risk. Cost focus in getting value for money pulls against mission focus in ensuring the system actually serves the public.
Navigate these explicitly rather than ignoring them. Write down, before the solicitation goes out, where you are choosing rigor over speed and where you are accepting a simpler test because your team can genuinely run it. That written trade-off is the thing you will point at when someone asks why you spent four weeks on fairness testing, and it is also what stops the trade-offs from being made silently by whoever is under the most pressure that week.
The Evaluation Hierarchy: Four Layers
Not all evaluation dimensions deserve equal weight, and not all of them belong in a scorecard at all. Some are filters. Structuring the evaluation in layers stops you from spending three weeks scoring the fairness of a vendor who was never eligible in the first place.
Layer 1, basic eligibility. These are hard filters. If a vendor fails any of them, stop the evaluation. Does the vendor understand government procurement, meaning working knowledge of FAR and DFARS? Can the vendor deliver by the required date? Does the vendor meet your minimum security and compliance baseline, whether FISMA, FedRAMP or an equivalent? Is the vendor financially viable, so they will still exist in three years? Does the system architecture fit within your infrastructure constraints?
Layer 2, core capability assessment. These dimensions decide whether the system can fulfil the mission at all: technical performance in accuracy, latency and throughput on your use cases; integration feasibility with your data and systems; fairness and bias risk across demographic groups; reliability and failure mode handling; and transparency and explainability.
Layer 3, implementation viability. These affect whether you can operate the thing after the award: data pipeline and infrastructure requirements, model update and retraining methodology, monitoring and performance measurement capability, documentation and knowledge transfer, and support and vendor responsiveness.
Layer 4, governance and risk. These establish whether the system can be governed responsibly: data ownership and security handling, audit and oversight capabilities, change management and version control, human and AI interaction design covering where humans override and how escalation works, and regulatory compliance and documentation requirements.
Score Three Dimensions, Never One
The trap Priya's director fell into is judging AI vendors on a single dimension, usually the demo. A sound evaluation scores three separate things, because a vendor can be brilliant at one and dangerous at another. The layers tell you what to look at; the dimensions tell you how to add it up.
Technical: can it actually do the job, on your data?
The "92% accurate" figure is meaningless until you ask three follow-up questions: accurate at what, measured on whose data, and accurate for which groups? A model trained on suburban applicants may be 92% accurate overall and badly wrong for rural or non-English-speaking applicants. Define the metrics that matter for your mission rather than accepting the vendor's preferred metrics, and insist that scoring rests on pilot results using your data, not vendor claims about someone else's.
- Performance metrics. What does accuracy mean here, what is the baseline, and how are errors distributed? Precision and recall both matter, and different use cases prioritize them differently. Latency, meaning real-time versus batch decisions. Throughput at production scale, in items per day. Graceful degradation when confidence is low or data is incomplete.
- Data pipeline quality. Test integration against your actual data infrastructure. Does the system handle your formats, encodings and volumes? How many data quality problems can it tolerate? Can it extract relevant features from your raw data? Does it flag anomalies, or accept bad data silently? Can you backtest on historical data for validation?
- Integration and infrastructure. API design, whether REST, message queue or direct database connection. Authentication and authorization against your identity management. Data residency, meaning where data is stored and whether that meets your geographic requirements. Scalability tested at your peak load, with a known breaking point. Monitoring hooks that let your operations team see inside the system.
- Reliability and error handling. What the system does with data it was not trained on. Robustness to data drift as the input distribution changes. Catastrophic failure recovery back to a baseline. Partial failure handling when a share of inference requests fail. Rollback to the previous version, quickly, when you need it.
- Security posture, including whether the service meets FedRAMP authorization if it is a cloud service handling federal data.
Do not let vendors minimize integration complexity during evaluation. Integration is the dimension most often underestimated in vendor timelines, and it is the one that turns a technically excellent system into a shelf-ware line item on next year's budget.
Ethical: will it treat the public fairly and lawfully?
This is the dimension commercial buyers skip and government buyers cannot. Because your agency makes decisions about people's rights and benefits, an unfair model is not just a quality problem; it is a legal and constitutional one. It is also the dimension most often rushed or dropped when the schedule tightens, which is precisely why it needs its own protected time in the evaluation plan.
- Has the vendor tested for disparate performance across protected groups, and will they show you the results?
- Can the system explain individual decisions well enough to support an appeal?
- Is the tool accessible to people with disabilities, meeting Section 508 requirements?
- What data was the model trained on, and was it lawfully obtained?
- Will the vendor support an algorithmic impact assessment before go-live?
Underneath those questions sit four concrete analyses. Selection-rate analysis asks whether the system makes decisions at similar rates across demographic groups: disaggregate performance metrics by protected classes such as race, gender, age and disability status, and look for disparate impact, the pattern where a system accepts applications from one group at 95% and from another at 75%. Then work out the root cause, because a gap can reflect system bias or legitimate operational factors, and the two demand different responses. Define what disparities you consider defensible in advance and record that decision. You decide it, not the vendor.
Error-rate analysis asks whether errors fall equally on everyone. False positive rates can diverge sharply, with one group flagged incorrectly at 5% and another at 12%. False negative rates can diverge too, with one group approved incorrectly at 2% and another at 8%. Calibration asks whether a stated 80% confidence really means 80% for every group. Threshold analysis asks whether the system would need different decision thresholds for different groups, which is a question to put to counsel as well as to your data team, because setting decision thresholds by protected group carries its own legal exposure and is not a purely technical adjustment.
Interpretability and transparency asks whether decision-makers can explain what the system decided and why: feature importance showing which inputs drove the decision, counterfactual explanations showing whether a different input would have changed the outcome, consistency across similar cases, and explainability aimed at the affected party, so a resident can understand why they were denied. Mitigation capability asks what happens if bias is discovered after deployment: can the system be retrained on adjusted data, can fairness constraints be added, what is the remediation timeline, and does the vendor support ongoing fairness monitoring? A vendor with no answer to the last question is offering you a system you cannot fix.
Operational: can your agency live with it for five years?
The cheapest model can be the most expensive system. Operational evaluation looks past the demo to the long marriage, and it is where implementation viability and governance risk turn into daily reality.
- Total cost including implementation, training, monitoring, and retraining, not just the license
- What happens when you want to leave: can you export your data and models, or are you locked in?
- How the vendor handles model updates that change behavior in production
- Support responsiveness and whether a real human answers when the system misfires
- Vendor financial stability, so the company is still there in year four
Human and AI interaction design deserves its own scrutiny. Which cases does the system handle and which route to humans? Can humans make a decision different from the recommendation, and is that override genuinely available rather than nominally available? What are the escalation pathways to a supervisor? Do human corrections feed back into the model? Is there an audit trail complete enough to reconstruct why a decision was made months later, when someone asks?
Monitoring and governance asks whether you can tell that the system is still working: performance tracking to see whether metrics are stable or degrading, drift detection for changes in the input distribution, fairness monitoring for emerging demographic disparities, explainability audits of recent decisions, and a change management process that says who approves an update. Vendor accountability is the contractual half: performance guarantees with a defined consequence if accuracy drops below the threshold you set, service level agreements and uptime commitments, audit rights that let your agency inspect the vendor's system and processes, data ownership and portability so you can leave with your data, and support escalation with a defined response time.
Audit rights deserve a specific note. When a vendor resists them, do not stop at the refusal and do not treat it as a confession. Probe the specific objection: is it about protecting source code from disclosure, about trade secrets appearing in a FOIA-able record, about the cost of supporting an audit, or about something in the system they would rather you not see? Each of those has a different remedy, from a protective order to a scoped audit of outputs rather than internals. What does not change is the requirement itself. An agency making rights-impacting decisions needs the contractual right to audit, and a vendor who cannot reach any workable version of it is a vendor you cannot govern.
Weighted Scoring: Make the Trade-Offs Explicit
Three dimensions are not equal for every purchase. A rights-impacting eligibility tool should weight ethical factors heavily; an internal transcription tool can weight operational cost more. Weighting forces the program office to decide what matters before seeing the demos, which is exactly when judgment is clearest. Here is a scorecard structure Priya can hand to her evaluation panel.
Criterion- the specific question being scoredDimension- technical, ethical, or operationalWeight- percentage of total, set in advanceScore 1 to 5- evidence-based rating per vendorWeighted score- score multiplied by weightEvidence note- the document or test that justifies the score
For a rights-impacting triage tool, Priya's panel set weights at 35% technical, 40% ethical, 25% operational. That single decision changed everything. Vendor A had the slick interface the director loved and scored highest on operational, but the vendor refused to share group-level fairness results, scoring a 2 on the heaviest-weighted ethical criterion. Vendor C had a plainer interface but produced a full fairness report and supported appeals explanations. On raw demo appeal, A won. On the weighted score, C won by a clear margin. The methodology overruled the vibe, which is its job.
A second scheme is worth knowing because it splits the same territory four ways rather than three, separating implementation from governance. It is a useful reference template when your acquisition is large enough that those two concerns need their own owners and their own scores. The weights and sub-weights below are the template's, not a rule; whichever cut you use, set the numbers before the responses arrive and record why you chose them.
| Dimension | Weight | Sub-components | How it is scored | Minimum to pass |
|---|---|---|---|---|
| Technical performance | 35% | Accuracy 12%, integration complexity 12%, reliability 11% | Pilot results on your data, not vendor claims | 70%, non-negotiable minimum |
| Fairness and bias | 30% | Selection-rate parity 10%, error-rate parity 10%, transparency 10% | Independent bias audit, red team testing, reference checks | 75%, civil rights non-negotiable |
| Implementation viability | 20% | Integration feasibility 7%, data pipeline quality 7%, timeline realism 6% | Technical architecture review, reference implementation timelines | 65% |
| Governance and risk | 15% | Monitoring capability 5%, vendor accountability 5%, security and compliance 5% | Contractual terms, security assessment, regulatory review | 70% |
The composite is the weighted sum: technical multiplied by 0.35, plus fairness by 0.30, plus implementation by 0.20, plus governance by 0.15. The template's recommendation bands are 75 and above for a strong accept, from 65 up to 75 for a conditional accept with remediation, and below 65 for rejection. Check the sub-weights when you adapt this: each dimension's components must sum to that dimension's weight, and the four dimensions must sum to 100, or the composite is quietly reporting something other than what you agreed.
The most important part of the framework is the override rule, because scoring informs the decision without determining it. Technical excellence does not override fairness concerns, and a fairness score below the minimum is a rejection whatever else the vendor achieved. Cost does not override security and compliance requirements. Vendor enthusiasm does not override integration complexity. Schedule pressure does not override evaluation completeness. Averaging is what lets a serious gap in one dimension disappear behind strength in another, and the minimum thresholds exist specifically to stop that trade.
Pilots: Mandatory, and Yours to Run
Accept no pilot result you did not independently verify. The vendor has every incentive to show favorable results, and when the vendor controls the environment, the data, the monitoring and the success criteria, a good result tells you about the pilot rather than about the system. Your team runs the pilot, or at minimum co-runs it with the vendor in a supporting role. You define the test data, the success criteria and the monitoring. You evaluate the results.
Assume good faith and verify anyway. Vendors optimize presentations for winning contracts, which is their job, and your methodology should require independent validation of any material claim. Design the pilot so a failure is visible: define in advance what result would stop it, not just what result would pass it. A pilot with no defined failure criterion will pass, because someone will always find a reading of the numbers that lets the schedule continue.
Two Worked Scenarios
Scenario one: a competitive permitting acquisition
A city permitting office needs to evaluate AI systems for reviewing building permit applications. Four vendors submitted proposals, with an eight-week evaluation window. Weeks 1 and 2 covered basic eligibility screening: vendor financial health, government experience, and the baseline security and compliance posture. Weeks 2 and 3 covered technical evaluation, requesting system documentation, reviewing architecture, and examining preliminary accuracy claims. Weeks 3 and 4 went to the fairness and bias deep dive, requesting disaggregated metrics and conducting an independent bias audit on provided test data. Weeks 4 through 6 ran pilots, with each remaining vendor running a one-week pilot on 500 actual permit applications while the city team monitored. Weeks 6 and 7 covered implementation and governance: detailed contract review, integration architecture against city systems, and monitoring design. Week 8 produced final scoring and the recommendation.
The outcome inverted the first impression. Vendor A scored highest on technical performance at 88% but only medium on fairness at 72%, citing "diverse training data" without evidence, which the panel flagged for follow-up. Vendor B scored slightly lower on technical at 85% but highest on fairness at 89%, with detailed disaggregated metrics, an independent bias audit, and open discussion of the system's limitations, on a realistic implementation timeline. Vendor C claimed the highest accuracy of all at 95% but refused to share methodology or fairness data, which the panel read as red flags across the board and recommended for rejection. Vendor D withdrew during evaluation after discovering it could not modify its architecture to meet the city's data residency requirement. The recommendation went to Vendor B, with enhanced monitoring requirements around demographic fairness and quarterly re-validation attached to the award.
Scenario two: a sole response
A federal agency received exactly one response to an RFP for a case routing system. With no competitive alternative, the evaluation cannot rest on comparison, so it has to be more thorough rather than less. The agency raised the bar at every layer. Layer 1 eligibility screening ran as standard and the vendor passed. Layer 2 capability required a four-week pilot rather than the usual two, with 1,000 test cases rather than 500, deliberately including edge cases. Layer 3 implementation required two reference sites with legacy infrastructure similar to the agency's own, not simply two reference sites. Layer 4 governance called for enhanced contract terms precisely because competitive pressure was absent: a longer commitment, stronger performance service level agreements, and more flexible exit clauses.
The pilot earned its cost. The system underperformed badly on multilingual cases, reaching 40% accuracy against a claimed 95%. That gap was material because the agency serves a population that is 35% non-English-primary. The negotiated resolution reduced scope to English-primary cases, with humans handling multilingual matters, adjusted contract pricing to reflect the reduced scope, and imposed quarterly monitoring on bilingual case accuracy with improvement milestones. If no improvement arrived by month nine, the vendor would provide alternative solutions or refund a portion of fees. Note what made this possible: the agency ran its own pilot on its own cases. A vendor-run pilot on curated English-language samples would have shown 95% and the agency would have found the 40% in production, one denied applicant at a time.
Reference Checking for AI Vendors
Vendors give you references who will say nice things. Useful reference checks get past the script by asking specific, awkward questions. Priya called two references for Vendor C and asked:
- "Tell me about a time the system got a decision wrong. What happened next?" A vendor whose references cannot recall any failure has references who are not really using it.
- "How long did implementation actually take versus the estimate?" AI projects routinely run double.
- "Have you had any fairness complaints, audits, or records requests about this system? How did the vendor support you?"
- "If you were buying again, what would you negotiate differently?"
- "What does the vendor do when a model update changes the results you were getting?"
One reference let slip that an update had silently changed the scoring threshold, and several borderline applicants were suddenly denied before anyone noticed. That single anecdote told Priya more about the vendor's change-management discipline than the entire sales process did. Always check at least two references, and always ask about failure, because the demo only ever shows success. Where you can, choose references whose environment resembles yours: two references from large private companies tell you little about how a system behaves inside a legacy government infrastructure with a public records obligation attached to it.
Putting It Together as a Procurement-Ready Packet
Priya's final recommendation was not "Vendor C looks best." It was a defensible packet: the weighted scorecard with evidence notes, the two reference summaries, the vendor's fairness report, and a one-page rationale. When a losing vendor protested the award, that packet was what the county had to show. A documented methodology does not just pick a better tool; it makes the reasoning reviewable when the protest, the audit, or the public records request arrives.
Be careful about what documentation actually buys you. It makes your decision explainable, which is necessary and not sufficient. A thoroughly documented evaluation that skipped fairness testing is a thoroughly documented evaluation that skipped fairness testing, and writing it down neatly does not repair it. The packet is evidence of the judgment you exercised; it is not a substitute for having exercised any. Build the scorecard and weights before you sit through a single demo. Score on evidence, not impressions. Run your own pilot. Check references about failure. Document everything. That sequence turns a confident sales pitch into a comparison you can defend to an auditor, a judge, and the public.
Anti-Patterns
- Evaluating toward a predetermined choice. A vendor is the safe choice, or someone in leadership likes them, and you unconsciously weight the criteria that favor them and downweight the ones where they are weak. It comes apart when objective parties, such as the civil rights office or the technical lead, notice the evaluation was structured to support a conclusion. Establish criteria before vendor responses arrive, use independent evaluators, document how decisions were made, and when someone raises a concern, address it substantively rather than defensively.
- Scoring only what is easy to measure. Performance is the easiest dimension to quantify, fairness is the hardest, and time pressure pushes attention toward whatever produces a number fastest. The result is a high-performance system that is biased or unmonitorable, pulled down after deployment. Make fairness, security and governance non-negotiable dimensions, allocate real timeline to them, and bring civil rights and security teams in from the start rather than at the review.
- Accepting a vendor-run pilot. When the vendor controls the environment, data, monitoring and success criteria, the pilot measures the vendor's staging setup. You discover the difference on real data with real infrastructure. Run it yourself, or co-run it with the vendor supporting, and define the test data, success criteria and monitoring yourself.
- Compressing evaluation phases under schedule pressure. Budget cycles, leadership expectations and vendor availability manufacture urgency, and fairness and implementation testing are the first things cut. The problems you would have caught surface after award, when you cannot undo the acquisition and remediation costs far more. Build a realistic timeline; if you cannot evaluate thoroughly in the time available, delay the acquisition or reduce the scope. A delayed good decision beats a rushed bad one.
- Letting a strong average hide a fatal gap. Averaging permits high scores in one dimension to compensate for a serious deficiency in another, and the dimension you traded away becomes the critical problem after deployment. Set minimum thresholds for the critical dimensions of fairness, security and compliance that cannot be traded off. A system that is 95% good but 40% fair is unacceptable, period.
- Treating the paperwork as the diligence. A complete scorecard, neat evidence notes and a tidy rationale describe the evaluation you ran. They do not improve it. If the fairness column was scored from a vendor assertion rather than a test, the documentation records that fact rather than curing it, and an auditor reading carefully will see exactly that.
- Reading a refusal as a verdict, or as the end of the conversation. When a vendor resists audit rights or declines to share fairness data, neither of the easy responses is right: dropping the requirement abandons your ability to govern the system, and treating the refusal as proof of guilt stops you learning anything. Ask what specifically they object to, price the objection, and negotiate a form of the right that still gives you inspection. Then score what actually happened.
Practice Prompts
- Design an evaluation framework. Your agency is about to acquire AI for benefits determination. What are your Layer 1 basic eligibility criteria, meaning what would disqualify a vendor immediately? What are your Layer 2 core capability criteria? Which fairness metrics matter most for your mission, and how will you measure them? How will you weight the dimensions, and why those weights? What is your minimum acceptable score in each critical dimension?
- Triage five red flags. A vendor proposes an AI hiring recommendation system. During evaluation these emerge: the system achieves 92% accuracy overall but only 78% for candidates with disability-related work history gaps; integration requires modifications to your legacy HR system that the vendor calls "possible but expensive"; the vendor can offer only two reference customers, both large private companies rather than government; the support agreement commits to "reasonable effort" to respond within two business days; and the vendor refuses audit rights, citing proprietary system architecture. For each: is it a red flag, how serious, what follow-up would you do, and would you reject or negotiate remediation?
- Build a scoring matrix for a document classification system. Include the dimensions and sub-components, the scoring scale, the weighting for each dimension, the pass and fail criteria, and worked examples of how you would score specific dimensions.
- Design a rigorous pilot for a system you are considering. Specify duration and phases, test data selection and how representative it is of operational reality, success metrics, the failure criteria that would stop the pilot, who participates from which teams, how you would test for fairness explicitly, and the post-pilot evaluation process.
- Compare two proposals. Vendor A offers strong technical performance, good fairness metrics, excellent support terms, a 12-week implementation, at $500K annual. Vendor B offers exceptional technical performance, incomplete fairness data, minimal support, a 6-week implementation, at $250K annual. What follow-up would you do for each? What would cause you to select A despite higher cost and longer timeline? What would cause you to select B despite incomplete fairness evaluation? What is non-negotiable and what can you trade?
Reflection
Think about an AI acquisition your agency is planning or might undertake. What evaluation framework would you design for it? Which dimensions are most critical to that mission, how would you weight them, and why those weights? What red flags would cause you to reject a vendor despite genuine strengths elsewhere? Which fairness metrics matter most for this specific use case, and who in your agency is qualified to interpret them?
Then turn it around. Imagine the system has been running for eighteen months and something has gone wrong. An inspector general asks what evaluation process you followed. Walk through the answer you would give today. Where does it get thin? The thin part is where your methodology needs work now, while the acquisition is still ahead of you and the fix is a paragraph in a solicitation rather than a remediation plan.
Glossary
- Disaggregated metrics. Performance reported separately for different demographic groups or data subsets. Essential for fairness evaluation, because aggregate metrics alone hide group-specific disparities.
- Fairness. The principle that AI systems do not systematically disadvantage protected groups. Multiple definitions exist, and the one you are using must be defined explicitly for your context.
- False positive rate. Of all negative cases, the percentage the system incorrectly classified as positive. Important in domains where false positives are costly.
- False negative rate. Of all positive cases, the percentage the system failed to identify. Important in domains where missing positives is costly.
- Integration. The process of connecting a vendor system to existing data pipelines, decision workflows and compliance reporting. Routinely underestimated in vendor timelines.
- Pilot program. A limited trial deployment to validate vendor claims and test integration before full acquisition. Essential for government AI procurement, and only meaningful when you control it.
- Service level agreement (SLA). A contractual commitment on system availability, response times and performance metrics. Should include consequences for non-compliance.
- Weighted scoring. A methodology that assigns importance weights to evaluation dimensions and calculates a composite score reflecting those priorities.
Related Lessons
- Federal Acquisition of AI: FAR/DFARS covers the procurement law framework this methodology operates inside.
- Writing AI Requirements in RFPs and SOWs is where your evaluation criteria become solicitation language before responses arrive.
- Evaluating AI Vendor Claims goes deeper on interrogating a single accuracy claim than this lesson's scorecard has room for.
- Algorithmic Impact Assessments covers the assessment you should be requiring from the vendor as a contract condition.
- OMB M-24-18 and AI Procurement Governance connects vendor evaluation to the wider procurement governance expectations.
Closing
Effective AI vendor evaluation balances competing demands: comprehensiveness against speed, technical depth against stakeholder accessibility, cost focus against mission focus. The framework in this lesson gives you structure for that navigation, and it is a template rather than a rule. Define your Layer 1 hard filters. Establish Layer 2 core capability requirements. Assess Layer 3 implementation viability honestly. Evaluate Layer 4 governance and risk. Apply weighted scoring, set decision rules that stop critical dimensions from being traded away, require pilots with independent verification, and document the methodology so the decision can be reviewed.
AI vendor evaluation is one of the highest-leverage decisions a government procurement team makes. A good evaluation prevents costly mistakes, protects public trust, and gives the mission a system that actually serves it. A poor one produces an acquisition that Congress, auditors and residents will all eventually examine. Priya did not out-argue three sales teams. She simply decided what mattered before they walked in, and then held to it.
Key Takeaways
- Filter before you score. Layer 1 eligibility criteria on procurement literacy, delivery date, security baseline, financial viability and architecture fit are hard filters. Failing one ends the evaluation rather than costing points.
- Score three dimensions separately. Technical capability, ethical fairness and lawfulness, and operational sustainability are different risks; a vendor can excel at one and fail another.
- Interrogate the accuracy claim. "92% accurate" means nothing until you know the metric, whose data it was measured on, and how it performs across different groups of people.
- Set weights before the demos. Deciding what matters in advance keeps the flashiest interface from overriding the factors that actually carry legal and public risk, and it is the only time your judgment is uncontaminated by the pitch.
- Weight ethics heavily for rights-impacting tools. When a system affects benefits, rights, or opportunities, fairness and explainability deserve more weight than cost or interface polish, and minimum thresholds should stop a strong average from hiding a fatal gap.
- Run your own pilot. A vendor-controlled pilot measures the vendor's staging environment. Define the test data, success criteria, monitoring and failure conditions yourself, and verify every material claim independently.
- Reference-check for failure. Ask references about wrong decisions, real implementation timelines, and silent model updates; vendors whose references recall no failures are not really being used.
- Watch for lock-in and update risk. Confirm you can export your data and that the vendor disciplines model updates, because silent changes can quietly deny real applicants.
- Keep audit rights, and probe the objection. If a vendor resists, ask what specifically they object to and negotiate a workable form. The right to inspect is what makes the system governable after the award.
- Document the evaluation as a defensible packet, and know its limits. A weighted scorecard with evidence notes and reference summaries makes your reasoning reviewable when the protest or the records request arrives; it records the diligence you did, and cannot supply diligence you skipped.
Frequently Asked Questions
The vendors all quote a single accuracy number. What do I ask for instead?
Ask for the metric behind the number, the dataset it was measured on, and the same number broken out by group. Precision and recall usually matter more than a blended accuracy figure, and which one dominates depends on whether a false positive or a false negative does more damage in your program. Then ask for the error distribution, because a system that is uniformly slightly wrong and a system that is badly wrong for one population can report the same headline accuracy.
We only got one bid. Can we shorten the evaluation?
The opposite. With no competitive alternative, comparison cannot do any of the work for you, so every dimension has to be tested on its own terms. Lengthen the pilot, enlarge the test set, insist on reference sites that resemble your infrastructure rather than any reference at all, and strengthen the contract terms specifically because competitive pressure is not available to discipline the vendor's behavior after award.
What if the vendor will not share fairness data?
Score it as what it is, which is an absence of evidence on your heaviest-weighted criterion for a rights-impacting system, and find out what the objection actually is before you conclude anything about the reason. Some vendors genuinely cannot release a client's disaggregated results; some can produce results on your test data under the pilot instead. If no version of the disclosure is achievable, you are being asked to deploy a system whose fairness you cannot assess, which is a decision you should make explicitly and in writing rather than by default.
How do I keep fairness testing from being cut when the schedule slips?
Give it a minimum threshold rather than a weight alone. A weight can be absorbed; a minimum cannot, because a score below it is a rejection regardless of the composite. Put the threshold in the evaluation plan before responses arrive, name the civil rights office as a required participant, and treat schedule relief as coming out of scope rather than out of the fairness phase.
Is a weighted scorecard enough on its own to defend an award?
No. The scorecard makes your reasoning legible; what defends the award is that the reasoning was sound and the evidence behind each score was real. Keep the evidence note for every score tied to a specific document, test result or reference conversation. An evaluator reviewing the packet is checking whether the scores rest on anything, and a column of confident ratings with no evidence behind them is worse than no scorecard at all because it looks like rigor.
How much should past performance with other governments count?
A great deal, if the environment resembles yours. A vendor with two references from large private companies has demonstrated something, but not the thing you need to know: how the system behaves inside legacy infrastructure, under public records obligations, with an appeals process attached. Ask references specifically about audits, complaints and records requests, because those are the pressures your deployment will face and a commercial reference has never met them.
Skill.re