The Questions That Make a Vendor Sweat
Forty minutes into the sales meeting, the demo has been flawless and the room is warm. The account executive is describing the roadmap with the easy confidence of someone who has given this exact performance a hundred times, because he has. Then the operations lead, a quiet woman who has said almost nothing, slides a single printed page out of her folder and asks the first question on it: "Which of your customers left last year, and can we speak to one of them?" Watch what happens next, because it is the most information-dense moment of the entire procurement cycle. The AE glances at the sales engineer. The sales engineer glances at his laptop. Somebody says the word "churn" the way people say the name of a disease. Whatever answer eventually arrives, the three seconds before it arrive first, and those three seconds are why this lesson exists. This is the capstone of Chapter 1.4: the complete 20-question due-diligence sheet, five families of four questions each, with what a good answer sounds like and what a bad one sounds like. It is the artifact you will carry into every vendor conversation for years.
Borrowed Experience: The Sheet as an Asymmetry-Remover
Every vendor meeting you will ever attend has a structural imbalance that no amount of intelligence on your side can fix in the room. The vendor has run this sale a hundred times. They know which questions buyers ask, which objections dissolve under which stories, and exactly how long a silence can hang before someone fills it with a concession. You, meanwhile, have bought enterprise AI perhaps twice, possibly never. In any repeated game, the side with the reps wins the improvisation. The only counter is to stop improvising.
The sheet is borrowed experience. Each of its twenty questions is a scar from a deal that went wrong for someone else: the company that discovered its data had quietly become training material, the team that watched output quality lurch sideways after a silent model swap, the buyer who learned at exit that "your data is always yours" meant a folder of PDF reports. You do not need to have lived those failures to be protected by them. You need the questions they generated, printed, in order, in your folder.
Two things happen when you run the sheet, and both work in your favor. First, good vendors respect it. A vendor with real answers recognizes the sheet as the signature of a serious buyer, and serious buyers close faster, negotiate more honestly, and churn less. More than one sales engineer will visibly relax when the questions get specific, because specifics are where a good product wins. The sheet does not lengthen good sales cycles; it shortens them, by clearing away the theater both sides otherwise have to perform.
Second, evasive answers are answers. You are not only collecting information; you are running twenty small experiments on how this company behaves when a truthful answer would be expensive. A deflection on the churned-customer question is not a gap in your data. It is your data. Chapter 1.4 has already given you the demo-conditions checklist, the benchmark deflation sheet, and the reprice worksheet; all three taught the same reflex, which is to move the conversation from the vendor's prepared ground to yours. The sheet is that reflex, systematized.
An evasive answer is not a missing data point. It is the data point.
One rule of use before the families: send the sheet ahead of the meeting. This is not softness; it is rigor. You are not trying to ambush a salesperson into stumbling, which proves nothing. You are giving the organization behind the salesperson every chance to produce its real documentation, so that when the answer is still vague a week later, you know the vagueness is structural.
Family One: Data Handling
The scene: a procurement call, second meeting. Your infosec analyst asks whether customer complaint text will be used to train models. The AE says "great question" and then, for eleven full seconds, nobody at the vendor answers it.
- 1. Where does our data go, and is any of it used to train or improve models? This is the question that separates a processor from a harvester. "We use your data to improve our services" is a training clause wearing a customer-service costume. Good: a data-flow diagram, named hosting regions, and a contractual no-training clause offered before you ask for it, with any improvement use strictly opt-in and separately documented. Bad: "your data is completely safe with us" (a security answer to a usage question), or terms described as "flexible," which means flexible in one direction.
- 2. What are the retention and deletion terms, and how is deletion verified? Data you cannot provably delete is a liability with your name on it, and "deleted" often excludes backups, logs, and derived artifacts. Good: retention windows stated per data class, deletion within a specific number of days of request, backup purge cycles named, and a deletion certificate or audit evidence you can file. Bad: "data is deleted upon request" with no timeline, no mention of backups, and no verification mechanism beyond trust.
- 3. Who can access our data: which vendor roles, and which subprocessors? Your data's real perimeter is the vendor's org chart plus every company they subcontract. Good: role-based access with logging you can inspect, a current published subprocessor list, and contractual notification before that list changes. Bad: "only authorized personnel," which describes every breach in history, or a refusal to name subprocessors for "commercial reasons."
- 4. What happens to our data at contract end? Exit terms negotiated at exit are dictated, not negotiated; the only time you have leverage on this clause is now. Good: export formats, export scope, and a deletion timeline written into the contract you are about to sign. Bad: "we can work that out at renewal," or data retained indefinitely "for compliance purposes" the vendor cannot cite.
Family Two: Failure Modes
The scene: the demo has just ended, flawlessly, as demos do. You ask for the measured error rate on customers like you, and the sales engineer begins scrolling for a slide that you can tell, from his scrolling, does not exist.
- 5. What is your measured error rate on customers like us, and in what units? As the benchmarks lesson put it: an accuracy number without units, a denominator, and conditions is decoration, not evidence. Good: "on mid-size utilities processing correspondence like yours, 6 to 8 errors per 1,000 documents, where an error is defined as X, measured over the last two quarters." Bad: "99 percent accuracy" with no per-what, no error definition, and a measurement set the vendor curated.
- 6. Show me your worst production incident and what changed after it. Every vendor at scale has had a bad day; you are testing whether they metabolize failure or bury it. A vendor with a postmortem culture will almost enjoy this question. Good: a specific incident, its timeline and blast radius, the root cause, and the named control or process that exists now because of it. Bad: "nothing major comes to mind," which means either they do not monitor or they do not disclose, and both are disqualifying.
- 7. What does the product do when it is uncertain? Confident wrongness is the signature AI failure mode; a system with no uncertainty behavior delivers its worst outputs with the same fluency as its best. Good: confidence thresholds, abstention, explicit flagging, and routing to a human queue, all configurable by you. Bad: "the model is extremely accurate," which answers a different question, or the quiet admission that it always produces an answer no matter what.
- 8. Who handles exceptions, and what does that queue look like at our volume? Exception handling is where pilots that demo well go to die in production, because a 4 percent exception rate is invisible at 50 items a day and a staffing problem at 5,000. Good: exception rates from comparable customers, multiplied out at your volume, with the human review hours stated plainly. Bad: "exceptions are rare," with no rate, no denominator, and no answer to who does the work.
Family Three: Model and Change Management
The scene: you ask which models actually run underneath the product. The account executive says "our proprietary AI engine" at the precise moment the sales engineer says the name of a well-known foundation model. They look at each other. You write down both answers.
- 9. Which models run underneath, and what happens when you swap them? Most AI products are orchestration built on someone else's model, which is fine; hiding it is not, and swaps change behavior. Good: named model providers and versions, plus a documented swap process with advance notice and regression evidence. Bad: "proprietary" as a full answer, or the cheerful "we always use the latest and greatest," which translates to: your outputs change whenever ours do.
- 10. How do you test releases before they reach us, and do outputs change silently? In deterministic software, an unannounced patch fixes bugs; in AI systems, an unannounced update can change the substance of what the product says on your behalf. Good: a regression suite run against a golden set before every release, published release notes, and a changelog of behavioral changes. Bad: "you won't even notice our updates," delivered as a selling point, when it is precisely the problem.
- 11. Can we pin or stage versions? If output changes silently and you cannot hold a version, every vendor release day is an unplanned change event inside your process, unmanaged by your change management. Good: version pinning, or at minimum a staged rollout with a sandbox window where you test the new version on your data before it goes live. Bad: a single multi-tenant version that ships when it ships, with no customer-side gate of any kind.
- 12. What is your evaluation suite, and can we see results on our sample set? This is the family's closing move: not their benchmark, your data. A vendor confident in their product treats a structured trial on your sample as a chance to win. Bad: "our benchmarks are industry-standard," aggregate metrics under NDA only, or a pilot offer that mysteriously requires a signed annual contract first. Good: a defined evaluation on 200 to 500 of your real (appropriately handled) documents, methodology shared, results in writing.
Family Four: Integration and Operations
The scene: "What does integration actually require from our side?" The AE says "it just plugs in." Three weeks later the statement of work arrives and it says, in effect, 400 hours of your people's time. You are asking now so the number arrives before the enthusiasm does.
- 13. What does integration actually require from us: systems, hours, roles? The 10-20-70 rule from Level 1 applies to vendors too: the algorithm is the easy part, and the integration and process work lands mostly on your side of the table. Good: a named list of your systems to be touched, hour estimates by role, and real implementation timelines from two or three comparable customers. Bad: "seamless," "plug and play," or "our team handles everything," which no vendor team in history has done.
- 14. What is the admin and oversight workload per week at steady state? Every AI system carries a permanent operating tax: configuration upkeep, exception review, user management, output spot-checks. The reprice worksheet from earlier in this chapter needs this number, and vendors know it, which is why they round it to zero. Good: an honest figure, such as "customers your size typically spend 5 to 8 hours a week," broken down by activity. Bad: "it basically runs itself."
- 15. What logs and audit trails do we get, and are they exportable? When a customer disputes an AI-touched decision, or a regulator asks how one was made, the answer lives in logs you either have or do not have. Good: per-decision records of inputs, outputs, model version, and reviewer, exportable via API in a standard format you own. Bad: a dashboard you can look at but not export, or "logs are available on request," meaning their request queue, their timeline.
- 16. What are your uptime history and support SLAs, actual rather than promised? A target is a hope; a history is a fact. Good: a public status page with 12 or more months of real history, the actual figure quoted unprompted, service credits with teeth, and support tiers with measured (not aspirational) response times. Bad: "we target 99.9 percent" with no history offered, and support described entirely through anecdotes about heroics.
Family Five: Lock-In and Exit
The scene: the question about customers who left has just landed, and the room's temperature has dropped two degrees. The AE offers to "connect you with our customer success team on that." You write down that this, too, is an answer.
- 17. What exactly is exportable on exit, and in what format? Enumerate it: raw data, configurations, prompt libraries, workflow definitions, fine-tuning artifacts or at least the training pairs behind them. Two years of accumulated configuration is two years of your operational knowledge, and it either leaves with you or it does not. Good: a written export scope in open formats (CSV, JSON) with a tested export path. Bad: "you can always export your reports," which on inspection means PDFs, which means nothing machine-readable leaves the building.
- 18. What is the real switching cost after two years? An honest vendor knows exactly what accumulates: integrations, tuned configurations, staff habits, workflow dependencies. Asking them to price your exit is a character test with a useful by-product. Good: a serious answer naming what would be hard to move and what the vendor does to keep it portable. Bad: "there's no lock-in at all," delivered quickly, which is either untrue or means the product is shallow.
- 19. Which customers left you last year, and why? May we speak to one? The demo lesson made the point about references: the customers a vendor hands you are marketing; the customers who left are evidence. Every vendor at scale has churn, so the honest answer exists. Good: a real number, real reasons ("two outgrew our volume tier; one needed on-premise we don't offer"), and at least a genuine attempt at a churned reference. Bad: "we have 98 percent retention" as a deflection, followed by a reference call that keeps almost getting scheduled and never does.
- 20. What happens to pricing at renewal, and what is your history of increases? AI vendor pricing in a moving market drifts, and it drifts against whoever has the higher switching cost, which after two years is you. Good: a stated increase history, a contractual renewal cap (say, CPI plus a fixed percentage), or multi-year price protection offered in writing. Bad: "we've always been fair with our customers," plus a refusal to cap anything, which tells you exactly what fairness will cost.
The Sheet, and the Sheet at Work
The Artifact: The 20-Question Due-Diligence Sheet
Here is the chapter's capstone artifact in its carry-in form: family, question, and the red-flag answer that should stop you. Print it. The detailed sections above are the training; this table is the field instrument. Score each answer as clean, partial, or red flag, and log the vendor's words verbatim, because verbatim answers are what you will compare across finalists and what you will negotiate into the contract.
| Family | Question | Red-flag answer |
|---|---|---|
| Data handling | 1. Where does our data go; is it used for training? | "We use it to improve our services" / terms are "flexible" |
| 2. Retention and deletion terms, verified how? | "Deleted on request," no timeline, silence on backups | |
| 3. Who can access it (staff, subprocessors)? | "Only authorized personnel"; subprocessors unnamed | |
| 4. What happens to our data at contract end? | "We can discuss that at renewal" | |
| Failure modes | 5. Measured error rate on customers like us, in what units? | "99 percent accurate," no denominator, no error definition |
| 6. Worst production incident and what changed after? | "Nothing major comes to mind" | |
| 7. What does the product do when uncertain? | "The model is extremely accurate"; it always answers | |
| 8. Who handles exceptions; the queue at our volume? | "Exceptions are rare," no rate, no staffing answer | |
| Model and change management | 9. Which models underneath; what happens on swaps? | "Proprietary engine" / "always the latest and greatest" |
| 10. How are releases tested; do outputs change silently? | "You won't even notice our updates" | |
| 11. Can we pin or stage versions? | One multi-tenant version, no customer-side gate | |
| 12. Evaluation suite; results on OUR sample set? | "Industry-standard benchmarks"; pilot only after signing | |
| Integration and operations | 13. What does integration require from us (systems, hours, roles)? | "It just plugs in" / "our team handles everything" |
| 14. Weekly admin and oversight workload at steady state? | "It basically runs itself" | |
| 15. What logs and audit trails do we get; exportable? | Dashboard only; "logs available on request" | |
| 16. Uptime history and support SLAs, actual not promised? | Targets without history; support by anecdote | |
| Lock-in and exit | 17. What is exportable on exit, in what format? | "You can export your reports" (PDFs) |
| 18. Real switching cost after two years? | "There's no lock-in at all" | |
| 19. Which customers left last year and why; churned reference? | Retention statistic as deflection; reference never schedules | |
| 20. Renewal pricing; history of increases? | "We've always been fair"; no cap offered |
A scoring note from practice: no vendor worth buying scores 20 clean. A vendor who scores 20 clean is either extraordinary or rehearsed, and the sheet's verbatim log will tell you which. What you are looking for is the pattern: honest partials on hard questions are the mark of a company you can operate with; smooth answers on easy questions paired with fog on questions 1, 17, and 19 are the mark of a company you will meet again in an exit negotiation you did not schedule.
This sheet completes a toolkit. It compounds with the demo-conditions checklist, the benchmark deflation sheet, the reprice worksheet, and the agent autopsy card from the four lessons before it: together they cover the demo, the evidence, the economics, the architecture, and now the relationship, which is the complete anatomy of an AI purchase read skeptically.
The Harlow Test: Two Finalists, One Sheet
A composite, hypothetical case to see the sheet decide a real decision. Harlow Municipal Utilities (fictional), 380 employees, is choosing between two finalists for a customer-correspondence AI to help draft responses to roughly 70,000 letters and emails a year. Budget: $120,000 annually. Both demos were excellent; Vendor B's was, if anything, better. Harlow's readiness lead sends the 20-question sheet to both vendors a week before the final sessions and scores the answers live.
Vendor A answers 17 of 20 cleanly. Data handling: no-training clause offered unprompted, deletion within 30 days with certificate, subprocessor list published. Failure modes: 7 errors per 1,000 documents on two comparable utilities, error definition included; a 2025 incident (a template bug that misdated 1,100 payment notices) described in detail along with the golden-set regression test built because of it; uncertainty routes to a review queue at a threshold Harlow controls; at 70,000 items a year, an estimated 4 percent exception rate, about 2,800 items, roughly half an FTE of review, stated without flinching. Then the stumble, and it is an honest one: on question 11, the sales engineer says plainly, "We can't pin model versions; we run one platform version for all customers." Instead of fog, an offer: 21 days advance notice of every model change plus a parallel test window where Harlow can run the new version against the current one on its own sample before cutover. Not a perfect answer. A real one. Admin workload: "6 to 9 hours a week at your volume." Renewal: capped at CPI plus 3 percent in writing.
Vendor B is smoother on features and slipperier on the sheet. Question 1 produces "our data terms are flexible and we can tailor them," which, pressed, resolves to a default license to use customer data for "service improvement." Question 19 produces a retention statistic, then a promise to arrange a call with "a customer who evaluated leaving," and then three weeks of scheduling emails that never converge on a meeting. Question 17 produces "all your reports are fully exportable," which under one follow-up means PDF reports: prompts, configurations, and routing rules stay in the platform. Eleven clean answers, and fog precisely on the three questions where fog is most expensive.
Harlow signs with Vendor A, and the decision memo cites the sheet, not the demo: 17 clean answers versus 11, one honest limitation with a negotiated compensating control (the test window went into the contract), against three red flags on data usage, churn, and exit. Total decision cost: two meetings and a printed page. Six months later, a story circulates at an industry roundtable (hypothetical, like everything in this example): a utility two states over is nine months into leaving Vendor B, paying an integrator roughly $60,000 to reconstruct routing logic and response templates from, of course, PDF reports. Harlow's readiness lead does not gloat. She updates the sheet's margin note on question 17: this one pays for the whole page.
What to Do Monday Morning
The sheet only becomes yours the first time it leaves your folder in a live meeting. Here is the sequence.
- Print the 20-question sheet from the table above, one page, and put it in the folder or notebook you actually carry. An artifact on a shared drive is a good intention; an artifact in the room is an instrument.
- Attach it to the next vendor conversation on the calendar, whatever stage it is at. If no meeting exists, run it retroactively against your organization's most recent AI purchase and score what was never asked.
- Send five questions ahead: numbers 1, 5, 11, 17, and 19, one from each family. This gives the vendor's organization a fair week to produce real documentation and makes any remaining vagueness structural rather than situational.
- Assign roles before the meeting: one person asks, one person logs answers verbatim, including the deflections. The verbatim log is what turns three seconds of hesitation into a citable finding in the decision memo.
- Score clean, partial, or red flag within an hour of the meeting, while the hesitations are still audible in memory, and mark which answers must convert into contract language (data terms, export scope, notice periods, renewal caps).
- File the scored sheet in your readiness portfolio next to the demo-conditions checklist, the deflation sheet, the reprice worksheet, and the autopsy card. Five artifacts now; Chapter 1.5 is about the career that owning this toolkit quietly builds.
Key Takeaways
- Treat the vendor meeting as a repeated game you are playing for the second time against an opponent playing it for the hundredth: the 20-question sheet is borrowed experience, and it removes the asymmetry that improvisation cannot.
- Run all five families every time: data handling, failure modes, model and change management, integration and operations, lock-in and exit; four questions each, covering the purchase from first data flow to final export.
- Read evasions as findings, not gaps: "flexible terms," "nothing major comes to mind," and the churned reference that never schedules are answers, and they belong verbatim in the decision memo.
- Demand units with every performance claim: an error rate per 1,000 documents on customers like you, with the error defined, or it is decoration; the benchmarks lesson's rule applies in the room, not just in the datasheet.
- Send five questions ahead of the meeting so a vague answer a week later is proof of structure, not an ambushed salesperson; you are testing the organization, not the person.
- Prefer honest partials to smooth perfection: a vendor who admits it cannot pin versions and offers a test window is operable; a vendor with fog on data usage, exit formats, and churn is an exit negotiation on a delay timer.
- Convert good answers into contract language before signing: no-training clauses, export scope and formats, change-notice periods, and renewal caps are worth nothing as meeting notes and everything as clauses.
- File the scored sheet with the demo-conditions checklist, deflation sheet, reprice worksheet, and autopsy card: the complete skeptic's toolkit for reading AI claims, and the foundation of the readiness role Chapter 1.5 introduces.
Skill.re