←
AI for Government
Proficient · M28 · lesson 28 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Evaluating AI Vendor Claims
📖
now learning

Evaluating AI Vendor Claims

15 min

Lena Whitfield, a program director at a state transportation department, sat through a vendor pitch for an AI system that would "predict road failures with 97 percent accuracy and cut maintenance costs by 40 percent." The slides were gorgeous. The demo was flawless. Six months and $1.8 million later, the system flagged a quiet residential street as critical while missing a bridge that needed repair. The 97 percent number, it turned out, came from the vendor's own lab on the vendor's own cleaned data, measured in a way that made a useless model look brilliant. Lena did not get fooled by a lie. She got fooled by a true number that answered the wrong question.

Learning to ask the right questions is the entire job of evaluating vendor claims, and in government it is more than good practice. Every AI procurement decision is a policy decision. The vendor an agency selects shapes how citizens experience public services, how staff roles change, and how risk is allocated between the agency and a company. This lesson gives you a structured way to pressure-test AI vendor claims before they become contract commitments. The goal is not cynicism. Good vendors exist and good tools work. The goal is to separate what the vendor has actually proven from what the marketing implies, and to leave a written record of which is which.

The frameworks you already work under supply the lens. OMB Memorandum M-24-10, Advancing Governance, Innovation, and Risk Management for Agency Use of Artificial Intelligence, issued in 2024, sets minimum practices for rights-impacting and safety-impacting AI, requiring agencies to verify pre-deployment testing, ongoing monitoring, human oversight, impact assessment and public notice. The NIST AI Risk Management Framework, which is voluntary guidance rather than a binding rule, provides the governance and technical scaffolding. The GAO AI Accountability Framework supplies audit-ready components covering governance, data, performance and monitoring. The Federal Acquisition Regulation supplies the procurement structure. Together they tell you what evidence a vendor should be able to produce.

The anatomy of a misleading metric

Most AI vendor claims are technically true and practically meaningless, because the number was measured under conditions you will never operate in. That is what makes them hard to challenge. A false claim can be contradicted; a true claim measured against the wrong question survives every objection except the right one, which is what the number is a measurement of. Lena's "97 percent accuracy" is the classic example, and three hidden questions unravel it. They unravel almost every headline number you will ever be shown, which is why they are worth memorizing rather than looking up.

  • Accurate at what? If 95 percent of roads are fine, a model that says "everyone's fine" scores 95 percent accuracy and catches zero real problems. The headline number hides whether it finds the rare cases that matter.
  • Measured on whose data? A model tested on the vendor's clean, curated dataset behaves very differently on your messy, real-world data. Performance almost always drops when it meets reality.
  • Compared to what? "40 percent cost reduction" against what baseline, over what period, at what scale? Without the comparison, the percentage is decoration.

A vendor metric without its measurement conditions is a headline without a story, and your job is to demand the story. Notice what this does to the negotiation: you are no longer arguing about whether the number is true, which the vendor will always win, and instead asking what the number is a measurement of, which the vendor can only answer with evidence or evasion. Both answers are useful to you.

Cutting through the marketing

AI marketing leans on a small set of impressive-sounding phrases that mean little until you define them. Translate them out loud in the room, and write down the answer you get. Doing this in the meeting rather than afterwards matters, because a vague phrase left unchallenged tends to reappear in the proposal, then in the evaluation summary, and eventually in the award justification, gaining apparent authority at every step without anyone ever having tested it. The translation column below is the whole technique: replace the adjective with the question that would make it checkable.

Marketing phraseWhat to ask to make it real
"99 percent accurate"Accurate on what task, what data, and how does it do on the rare cases that actually matter?
"Used by leading agencies"Which ones, for what use, and may I speak to a reference who deployed it at our scale?
"Powered by advanced AI"What kind of model, and does the architecture even matter for our outcome?
"Reduces costs by X percent"Against what baseline, measured over what period, including the cost of running and overseeing it?
"Continuously learning"Does it change behavior in production, and how will we be notified and re-verify it?
"Bias-free / fair"Tested how, on which groups, with what results you can show me?
"NIST AI RMF aligned"Which GOVERN, MAP, MEASURE and MANAGE artifacts exist, and may I see them by name?
"FedRAMP ready"Ready, authorized or equivalent? At what impact level, and what is inside the boundary?

The verification gates: authorization and framework alignment

Authorization status is usually the first gate, and it is the one vendors blur most often. Many claim to be FedRAMP ready, FedRAMP authorized or FedRAMP equivalent. These are not the same thing. FedRAMP Ready means an independent third-party assessment organization confirmed the vendor met some pre-authorization criteria. FedRAMP Authorized means the vendor holds a current Authority to Operate issued by an agency or the joint authorization body. FedRAMP Equivalent is a limited concept used in defense contexts with a specific meaning of its own. Verify current status on the FedRAMP Marketplace rather than in the vendor's deck.

Then read past the badge. Check the authorization boundary, because an authorization covers a defined set of components and not everything the vendor sells. Confirm that the impact level, Low, Moderate or High, matches the classification of the data you intend to put in it. Review the continuous monitoring reports, which tell you whether the posture is being maintained rather than merely achieved once. Document any inherited controls, so you know which protections you are relying on someone else to run. And be exact about what clearing this gate means: an authorization speaks to the security posture of an assessed environment. It says nothing whatsoever about whether the model is accurate, fair or fit for your decision.

Framework alignment is the second gate. A vendor claiming alignment with the NIST AI Risk Management Framework should be able to produce artifacts mapped to each function: a governance policy for GOVERN, context and risk framing for MAP, evaluation evidence for MEASURE, and ongoing monitoring and incident response for MANAGE. Many vendors produce glossy decks that name-drop the framework without a single documented artifact behind it. The remedy is unglamorous and effective: ask for each artifact by name, in writing, and treat the absence of one as a finding rather than an oversight.

Independent evaluation evidence

Independent evaluation is where most vendor claims fail, and it is where a procurement officer earns their keep. A vendor that cites a single aggregate accuracy number is not providing evidence; they are providing a summary statistic whose construction you cannot see. Real evaluation evidence has parts, and each part answers a question the aggregate deliberately or accidentally hides. Ask for all of them, ask what was measured, by whom, on what data, and note that a vendor who has done this work will produce it quickly because it already exists. Long delays usually mean the evidence is being created for you rather than retrieved.

  • Representative test sets. Was the model tested on data resembling the population and conditions you operate in, or on the clean subset that survived the vendor's own preparation?
  • Subgroup stratification. How does performance break down across the groups your program actually serves? An aggregate number can stay flat while performance for one group collapses.
  • Calibration. When the system reports high confidence, is it actually right that often? A miscalibrated model misleads every human reviewer downstream.
  • Uncertainty. Does the system express how sure it is, and is that expression meaningful rather than decorative?
  • Adversarial robustness. What happens under inputs designed to break it, or simply under the strange inputs your intake process produces on a bad day?

Where an independent benchmark exists for the capability in question, ask for results against it. Face recognition, for example, has long-running NIST evaluation programs, and standardized safety benchmarks exist for generative systems. Academic partners can run evaluations too. When a vendor cannot produce evidence of this kind, the agency should require it as a contract deliverable or decline to proceed. That is a strong position, and it is the right one: a system whose performance nobody outside the vendor has measured is a system whose performance nobody has measured.

Testing the demonstration

A vendor demo is a performance, rehearsed on inputs the vendor chose. The single most powerful move you have is to take control of the inputs. Insist on a live demonstration with your data and your edge cases, not their script. Bring a small set of cases you already know the answers to, including the hard ones: the ambiguous record, the unusual situation, the case that broke your last system. Watch what the tool does with them. A confident wrong answer on a case you know cold tells you more than any slide.

If the vendor resists testing on your data, that resistance is itself the answer. But hold the win loosely when you get it. A tool that performs well on the twenty cases you brought has told you something about those twenty cases. It has not told you how it behaves across the full mix of work your program sees in a year, on the days when intake quality drops, or two years from now when the population shifts. A successful pilot is evidence that justifies proceeding with monitoring in place. It is not a certification, and writing it up as one is how agencies talk themselves past a gate they meant to keep.

Validating benchmarks

When a vendor cites a benchmark, treat it as a claim to verify, not a fact to accept. Ask three things. Is the benchmark independent? A score the vendor measured itself, on a test it designed, is a marketing artifact until somebody else reproduces it. A result from an independent body carries more weight, though you should still ask who commissioned and paid for it. Does the benchmark match your use? Strong performance on a generic task does not predict performance on your specific decision, your data, or your population.

Can you reproduce it? Ask for the conditions in enough detail that your team, or a pilot, could actually check it: the dataset, the split, the metric definition, the preprocessing, the version of the model. A vendor confident in its numbers will share those conditions, and a vendor that only shares the headline is telling you something. Be careful with the inference, though. Sharing the conditions is not the same as the number being reproduced. Until you or an independent party rerun the measurement on representative data, a shared methodology is a more detailed claim, not a verified one. Write it in the file as "unverified, conditions disclosed" and move on.

Civil rights and equity obligations follow the tool

Buying an AI system does not outsource your legal obligations to the vendor. Title VI of the Civil Rights Act, Title VII in employment contexts, and the Americans with Disabilities Act all apply to what your agency does with the tool. EEOC guidance addresses AI in hiring, and federal contractors carry affirmative action obligations administered through OFCCP. The Department of Justice Civil Rights Division has publicly announced investigations into AI discrimination under the ADA and Title VI. The Blueprint for an AI Bill of Rights, the non-binding White House Office of Science and Technology Policy document, sets out an algorithmic discrimination protections principle that reinforces the same expectations.

What this means in an evaluation is concrete. Vendors should produce subgroup performance evidence, and they should commit contractually to ongoing disparate impact monitoring rather than a one-time test at acceptance. Ask what groups were evaluated and why those groups, since a vendor's demographic breakdown may not match the population your program serves. Ask what the vendor does when a disparity is detected, and whether you would learn about it from them or from a complaint. A "bias-free" claim with no disaggregated numbers behind it is not a finding of fairness; it is a phrase.

The enforcement record: unsupported AI claims are a legal risk

Vendor marketing is rarely a reliable source of truth, and regulators have said so with cases rather than guidance. AI vendor claims are subject to deceptive practices law under Section 5 of the FTC Act and comparable state laws. The Federal Trade Commission's 2023 action against Rite Aid concerned facial recognition that misidentified customers, particularly people of color, and led to a five-year ban on the company's use of facial recognition. The Commission's action against DoNotPay alleged deceptive claims that the tool was the world's first robot lawyer, and resulted in a settlement. Its action against Automators AI concerned earnings claims. Its action involving WeightWatchers Kurbo required deletion of improperly collected data, a remedy worth understanding because deletion obligations can reach models as well as datasets.

The Commission's Operation AI Comply sweep reached multiple deceptive AI marketers, including claims built on AI-generated fake endorsements. Beyond the FTC, the Department of Justice Civil Rights Division has opened AI discrimination investigations, and state attorneys general have brought their own enforcement actions. The point for a procurement officer is not that vendors are dishonest. It is that a claim which cannot survive your evaluation may also be a claim that cannot survive a regulator's, and that your agency does not want to be the deployment cited in the complaint.

What the public record shows

Five procurement histories are worth knowing because each one failed at a different point in the evaluation chain, and because in each the vendor's claims were accepted before they were tested at scale. Read them as diagnostics rather than as cautionary tales about particular companies. The products involved have changed, and some no longer exist in the form described; what has not changed is the shape of the mistake, which recurs in new procurements with new vendors every year. Each entry below names the point in the chain where the check should have happened.

  • IRS identity verification. Vendor claims about accuracy and accessibility did not withstand scrutiny at scale, and the arrangement was reversed by Treasury in 2022. The failure point was assuming that a system working for the typical applicant works for the atypical one.
  • Michigan MIDAS. Vendor performance claims for automated unemployment fraud detection did not withstand operational reality, and the wrongly accused bore the cost. The failure point was the absence of an independent accuracy measurement on live cases.
  • COMPAS. Vendor claims of accuracy masked racial disparities that ProPublica later documented. The failure point was an aggregate metric accepted without subgroup stratification, which is exactly the gap the evaluation evidence checklist closes.
  • Houston HISD teacher evaluation. Vendor claims about an evaluation model did not withstand due process scrutiny. The failure point was procuring a consequential decision system whose reasoning could not be explained to the people it judged.
  • The Dutch childcare benefits scandal. Mass data analytics produced cascading harm across thousands of families. The failure point was scale without monitoring, where an error rate that looks small in a slide becomes a national scandal at population volume.

Allegheny County stands as a counterpoint worth studying for what went right: external validation before scaling, community advisory input, and transparent ongoing monitoring. None of those are technical measures. They are procurement and governance choices, made before deployment, by people who assumed the vendor's evidence was incomplete until someone independent had looked at it.

Turning verified claims into contract terms

A claim you have verified is worth very little unless the contract obliges the vendor to keep it true. This is where evaluation hands off to negotiation, and where the questions you asked become clauses somebody can enforce. The Federal Acquisition Regulation and your agency's procurement policies shape which vehicles are available; the substance below belongs in the contract regardless of which one you use.

  • Audit rights. The ability to verify performance and compliance yourself, on a schedule and after significant changes, rather than receiving a summary the vendor prepared.
  • Exit rights. The ability to leave when performance or ethics problems emerge, which is only real if your data is portable and a transition is defined.
  • No-training clauses. A restriction preventing agency data from entering the vendor's training pipelines without authorization, with the covered services named.
  • Evaluation as a deliverable. Where the vendor could not produce independent evidence during evaluation, require it as a contract deliverable with a date and an acceptance standard.
  • Indemnification. Explicit allocation of risk between agency and vendor for the harms this system could plausibly cause.
  • Liability caps. Negotiated in proportion to the rights impact of the use, not accepted at the vendor's standard figure.
  • Service levels with remedies. Performance commitments with consequences attached, which is what aligns vendor incentives with the outcome your program needs.

Building a due diligence template you reuse

Everything above becomes portable the moment you write it down as a template rather than performing it from memory each time. Evaluations run under schedule pressure, usually by people covering several procurements at once, and anything that depends on remembering nine questions in the right order will be skipped in the week it matters most. A workable AI vendor due diligence template has five sections and fits on a few pages, which is short enough that a colleague can complete it without training and structured enough that two evaluators reach comparable conclusions.

Section one is authorization and boundary: status verified on the marketplace, impact level, what sits inside the boundary, continuous monitoring current, inherited controls listed. Section two is framework artifacts: each NIST AI Risk Management Framework function with the specific document that evidences it, or the word "none." Section three is evaluation evidence: representative test sets, subgroup stratification, calibration, uncertainty, adversarial robustness, each marked as provided, partial or absent, with who produced it. Section four is claim verdicts: every significant claim run through the worksheet, scored proven, unverified or unsupported. Section five is contract carry-forward: which verified claims become clauses, and which absent artifacts become deliverables with dates.

Fill it out for every AI procurement, keep the completed templates, and compare across vendors and across years. The comparison is where the value compounds. After three or four procurements you will know which artifacts vendors in your market can actually produce, which requests reliably go unanswered, and roughly what the evidence floor is for the category you buy in. That knowledge is worth more than any individual evaluation, and it is the thing a peer agency will most want from you.

Red-flag identification

Some patterns reliably signal trouble. None of them proves a vendor is unsuitable, and treating them as automatic disqualifiers would eliminate several good products. Treat each instead as a reason to slow down and dig harder, and to record what you found. The pattern that matters most is not any single flag but a cluster of them appearing in the same conversation, because the underlying condition they usually indicate is a product whose real performance has never been measured by anyone with an incentive to measure it honestly.

  • Refusal to test on your data. The strongest red flag. Good tools survive contact with reality; vendors who know this welcome the test.
  • Round, unconditioned numbers. "99 percent" with no definition of the task or data.
  • "Proprietary, can't disclose." Fine for the secret sauce; not fine for whether the thing works and how it was measured.
  • No discussion of failure modes. Every real AI system fails somewhere. A vendor who claims theirs does not is either naive or selling.
  • Vague data rights. If they dodge who owns your data and the model trained on it, that becomes your lock-in later.
  • Pressure to skip the pilot. "This offer expires Friday" is a sales tactic, not a technical merit.
  • Authorization language that shifts. "Ready," "authorized" and "equivalent" used interchangeably across the deck, the proposal and the conversation.
  • No incident history. A vendor with real deployments has had real failures. A clean history usually means a short one or an undisclosed one.

Structuring the vendor interview

Prepare questions that require evidence rather than adjectives, and ask them in an order that makes evasion visible. Ask for documentation, test reports and references you may contact. Ask how the vendor handles incident disclosure, rollback and remediation, and what the last incident was. Ask who will own the data, for how long, and what happens to it at termination. Ask whether the vendor's sub-processors carry their own authorizations, because your boundary is only as good as theirs. Ask about known failures and how they were addressed.

Request the specific artifacts by name: model cards, datasheets for the training data, red team reports, evaluation results with subgroup breakdowns, and the incident history. Then read what arrives with a clear head. A vendor that answers these questions with documented evidence has cleared a bar most cannot, and that is a genuine signal. It is not a verdict. Documentation describes what the vendor did; it does not establish that the system will perform in your environment, and a model card is a description rather than a test result. The verdict comes from your own testing plus the monitoring you build into the contract.

A vendor-claim validation worksheet

Run every significant claim through this before it influences your selection. Document the answers; they become evidence in your evaluation file, and that file is what protects the decision later.

  1. Restate the claim precisely. Write down the exact metric and number as the vendor stated it.
  2. Conditions. On what task, what data, what baseline, measured by whom, and when?
  3. Relevance. Does this measurement reflect how we will actually use it, on our population?
  4. Independent evidence. Is there proof beyond the vendor's own marketing, and who paid for it?
  5. Our-data test result. What happened when we ran our own cases through it, including the hard ones?
  6. Verdict. Proven, plausible but unverified, or unsupported. Score accordingly.

Replay Lena's pitch through this worksheet and the "97 percent" claim lands as unsupported: vendor's own data, undefined task, no test on real roads, no independent evidence. That single line in an evaluation file would have changed the decision, and saved $1.8 million and a missed bridge. Skepticism here is not obstruction. It is the craft.

Anti-Patterns

  • Treating a vendor benchmark as a measurement. A number the vendor produced, on a test the vendor designed, using data the vendor prepared, is a marketing artifact until an independent party reproduces it. This holds even when the methodology is disclosed in full, because disclosure enables reproduction rather than constituting it. Record such claims as "unverified, conditions disclosed," and either reproduce them during the pilot or require reproduction as a contract deliverable before the number is allowed to influence a selection score.
  • Reading a security authorization as a quality signal. A FedRAMP authorization tells you an assessed environment met security controls within a defined boundary at a stated impact level. It says nothing about model accuracy, subgroup performance, calibration or fitness for your decision, and it does not extend to components outside the boundary. Agencies routinely let a compliance badge do the work of an evaluation, and vendors have no incentive to correct them.
  • Accepting "ready," "authorized" and "equivalent" as one thing. They mean three different things, established by three different processes, with three different levels of assurance. Verify status and boundary on the marketplace rather than in the deck, and note in the file which term the vendor actually used in writing, because that is the one that will matter if the claim is later disputed.
  • Letting a successful pilot become a certification. A pilot tells you how the system behaved on the cases you happened to test, in the conditions that happened to hold. It does not bound behavior across your full case mix, across seasonal variation, or after the population shifts. Proceed on a good pilot, but proceed with monitoring, thresholds and a defined response, not with the pilot written up as proof.
  • Mistaking documentation for evidence. A model card, a datasheet and a policy mapped to a framework all describe intentions and design. They are necessary and their absence is disqualifying, but their presence establishes that the vendor wrote documents. Test results with subgroup breakdowns, produced by someone with no revenue at stake, are the thing you are actually looking for.
  • Accepting a fairness claim scoped to the vendor's groups. Subgroup testing is only meaningful against the population your program serves. A vendor's demographic breakdown reflects their existing customers and their available labels, which may not include the groups where your exposure lives. Ask which groups were evaluated, why those, and what the vendor does when a gap appears.
  • Evaluating the claim and never contracting for it. Verification at selection is a snapshot of a system that will change. Without audit rights, ongoing disparate impact monitoring, incident disclosure and evaluation deliverables written into the contract, the diligence you performed expires at signature and you have no mechanism to detect that it has.
  • Treating a third-party assessment as neutral by definition. An independent assessor is independent of your agency, not necessarily of the vendor who engaged and paid them, and the assessment covers a scope the vendor helped define. Ask who commissioned it, what was in scope, what was excluded, and whether anything material changed in the product since.

Practice Prompts

  1. Deconstruct a live claim. Take the strongest performance number from a vendor deck currently in front of your agency and run it through the six-step worksheet in writing. Then write the single sentence you would put in the evaluation file recording its verdict, and note what evidence would be required to move it up one level.
  2. Build the artifact request. Draft the written request you would send a shortlisted vendor asking for each NIST AI Risk Management Framework artifact by name, plus model cards, datasheets, red team reports, evaluation results with subgroup breakdowns and the incident history. Include a deadline and state what a non-response means for scoring.
  3. Design the our-data test. Assemble ten to twenty real cases from your own program where you already know the correct answer, deliberately including the ambiguous record, the unusual situation and the case that broke your last system. Define in advance what result would count as a pass, so the standard is set before the demo rather than after it.
  4. Check an authorization claim end to end. For one product your agency uses or is considering, establish from the FedRAMP Marketplace whether it is ready, authorized or equivalent, at what impact level, what sits inside the authorization boundary, and which controls you would be inheriting. Compare what you find to what the vendor's materials say.
  5. Convert diligence into clauses. Take three claims you verified during an evaluation and write the contract language that would keep each one true through the period of performance, including how it is measured, how often, by whom, and what happens when it is not met.

Reflection

Think about the last technology your agency bought on the strength of a demonstration. What did you actually verify, and what did you accept because it was said confidently by someone who seemed to know? If a vendor told you today that their system is 97 percent accurate, what would be your next three questions, and would you ask them in the room or afterwards in an email nobody answers? Which of the artifacts in this lesson could your current vendors produce within a week if you asked? When a claim you relied on turns out to have been unverified, who in your agency finds out, and how? And what would have to be true for you to walk away from a tool your leadership already wants?

Glossary

  • Accuracy. The share of predictions a system gets right overall. Nearly useless on its own when the outcome of interest is rare, because always predicting the common case scores well while catching nothing.
  • Subgroup stratification. Reporting performance separately for each group of interest rather than as a single aggregate, which is the only way a collapse affecting one population becomes visible.
  • Calibration. Whether a system's stated confidence matches its actual hit rate. A model that says "90 percent sure" and is right half the time misleads every reviewer downstream.
  • Adversarial robustness. How a system behaves under inputs designed to break it, and by extension under the malformed and unusual inputs normal operations produce.
  • Model card. A vendor-produced document describing a model's intended use, training data, evaluation results and known limitations. A description, not an independent test result.
  • Datasheet. The equivalent document for a dataset: where it came from, how it was collected and labelled, what it covers and what it omits.
  • Red team report. The record of a structured attempt to make a system fail or misbehave, and what was found. Its absence is more informative than its contents.
  • FedRAMP Ready. A status indicating an independent third-party assessment organization confirmed a vendor met certain pre-authorization criteria. Not an authorization.
  • FedRAMP Authorized. A status indicating the vendor holds a current Authority to Operate issued by an agency or the joint authorization body, for a defined boundary at a stated impact level.
  • Authorization boundary. The specific set of components an authorization covers. Everything the vendor sells outside that boundary is unassessed for this purpose.
  • Continuous monitoring. The ongoing reporting that shows whether an authorized security posture is being maintained, rather than whether it was achieved once at assessment.
  • Inherited controls. Protections your system relies on but does not implement, provided instead by an underlying authorized service. You still own the risk if they fail.
  • No-training clause. A contract restriction preventing agency data from entering a vendor's model training pipelines without authorization, effective only for the services it names.
  • Disparate impact monitoring. Ongoing measurement of whether a system produces materially different outcomes across groups, continued through operation rather than performed once at acceptance.

This lesson sits inside the procurement chapter. Federal Acquisition of AI: FAR/DFARS establishes the acquisition framework, Writing AI Requirements in RFPs and SOWs covers writing the requirements that make claims testable in the first place, and AI Vendor Evaluation Methodology is the wider scoring and selection process this evaluation feeds. The verified claims here become terms in AI Contract Negotiation and obligations you enforce in Managing AI Vendor Performance. For the technical depth behind the evidence you are demanding, see Testing and Validating AI Systems and Understanding AI Bias; for the authorization gate, FedRAMP and AI Cloud Authorization; and for the frameworks supplying the lens, NIST AI RMF: MAP, MEASURE, MANAGE and GAO AI Accountability: Four Principles in Practice. Third-Party AI Risk Management and Vendor Lock-In Prevention extend the supply chain and portability questions raised here.

Closing

Evaluating vendor claims is a craft, not a personality trait. It does not require distrust of vendors and it does not require a technical background. It requires a habit: for every number that will influence a decision, write down what was measured, on whose data, against what baseline, verified by whom, and what your own testing showed. Do that consistently and the difference between a proven claim and a persuasive one stops being a judgement call and becomes a line in a file.

Lena Whitfield lost $1.8 million and, more seriously, missed a bridge, because a true number answered a question nobody had asked. The number was never the problem. The absence of a structure that would have exposed what the number measured was the problem, and that structure costs a few hours per procurement. Bring it to your next vendor conversation, ask for the artifacts by name, and be willing to record "unsupported" next to a claim that leadership finds compelling. Skepticism, applied with a worksheet and a written record, is professionalism.

Key Takeaways

  • True numbers can mislead. Most bad AI buys come not from lies but from real metrics measured under conditions you will never operate in.
  • Demand the conditions. A metric without its task, data and baseline is a headline with no story. Always ask "accurate at what, on whose data, compared to what."
  • A vendor benchmark is marketing until reproduced. Disclosed methodology is a more detailed claim, not a verified one. Reproduce it, or contract for reproduction, before it scores.
  • Know your authorization vocabulary. Ready, authorized and equivalent are three different things. Verify status, boundary and impact level on the marketplace, not in the deck.
  • An authorization is not a quality signal. It speaks to security posture inside a boundary. It says nothing about accuracy, fairness or fitness for your decision.
  • Ask for evidence with parts. Representative test sets, subgroup stratification, calibration, uncertainty and adversarial robustness. An aggregate accuracy number is not evaluation evidence.
  • Test on your data, then hold the win loosely. Your own cases are the strongest move you have, and resistance to them is the loudest red flag. A good pilot justifies proceeding with monitoring, not a certificate.
  • Civil rights obligations do not transfer to the vendor. Require subgroup evidence and a contractual commitment to ongoing disparate impact monitoring, scoped to the population you serve.
  • The enforcement record is real. Unsupported AI claims have drawn FTC action under Section 5, DOJ civil rights investigations and state attorney general enforcement.
  • Document every claim's verdict. Proven, unverified or unsupported, in the evaluation file, so the decision rests on evidence rather than slides.

Frequently Asked Questions

The vendor says their accuracy number comes from a third-party evaluation. Is that enough? It is better than a self-reported figure and it is still not the end of the inquiry. Ask who commissioned and paid for the evaluation, what was in scope and what was excluded, which version of the product was assessed, and whether anything material has changed since. An assessor independent of your agency may have been engaged by the vendor, against a scope the vendor helped define. Where a genuinely independent benchmark exists for the capability, such as the long-running NIST evaluation programs for face recognition, ask for results against that instead.

We do not have the technical staff to reproduce a benchmark. What can we realistically do? More than you think. Running your own known-answer cases through the system requires program knowledge rather than data science, and it is the single most informative test available to you. Beyond that, three moves are open to any agency: require the evaluation as a contract deliverable with an acceptance standard, ask a peer agency or an academic partner whether they have evaluated the same product, and write audit rights into the contract so you can commission an assessment later without renegotiating. Record the claim as unverified in the meantime rather than letting it score as proven.

Does a FedRAMP authorization mean the AI itself has been vetted? No, and this is the most common substitution in AI procurement. An authorization concerns the security posture of an assessed environment within a defined boundary at a stated impact level. Nobody in that process evaluated whether the model is accurate, whether it performs equitably across groups, whether it is calibrated, or whether it suits your decision. Those are separate questions requiring separate evidence. Check the boundary too, since vendors often hold an authorization covering part of what they are selling you.

What if the vendor refuses to let us test on our own data, citing security or confidentiality? Take the reason seriously and then solve it rather than waiving the test. Synthetic or de-identified versions of your hard cases, testing inside their authorized environment, a mutual non-disclosure agreement, or an on-site session where nothing leaves the room all address genuine concerns. What none of them require is abandoning the test. If every accommodation is refused, you have learned something important, and you should record it in the evaluation file in those words rather than as a scheduling difficulty.

How do I raise all this without poisoning the relationship with a vendor we may work with for years? Frame it as process rather than suspicion, and apply it identically to everyone. Send the same written artifact request to every shortlisted vendor, use the same worksheet on every claim, and tell vendors up front what evidence will be required and how it will be scored. Strong vendors generally welcome this, because rigorous evaluation is how they beat competitors who are selling adjectives. The vendors who find it hostile are usually telling you why.

A claim we verified during evaluation has stopped being true in production. What should have prevented that? Verification at selection is a snapshot, and AI systems change through retraining, model updates and shifts in the population they see. The protections are contractual: audit rights exercisable on a schedule and after any significant change, evaluation reports as recurring deliverables rather than one-time submissions, ongoing disparate impact monitoring, incident disclosure obligations, and service levels with remedies attached. If none of those are in the contract, the diligence expired at signature, which is a procurement design problem rather than a vendor betrayal.