←
AI Readiness & Process Transformation
Aware · M9 · lesson 9 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

How to Read a Vendor Demo Without Getting Played

15 min

The solutions engineer has done this demo two hundred times, and it shows. He drags a contract into the window and the clauses light up like a pinball machine: payment terms, auto-renewal, liability cap, each extracted and summarized in seconds. He does it again with a second contract, chatting easily through the four seconds of processing so you never quite register them as waiting. Someone on your side of the table whispers "wow." Here is what you cannot see from your chair: those two contracts were chosen from thousands, the sequence was rehearsed this morning, and the one output that came back weird during rehearsal was quietly regenerated before you walked in. None of this is fraud. It is theater, professionally produced, and every demo you will ever watch is a performance engineered under conditions the vendor controls. This lesson teaches you the counter-move: converting the performance into a structured test under conditions you control, using a checklist you can fill in without saying a single hostile word.

The Anatomy of Demo Engineering

Start by respecting the craft, because you cannot counter what you refuse to understand. A vendor demo is engineered the way a stage magic act is engineered: nothing shown is fake, and everything shown is selected. Knowing the six techniques changes how you watch.

Curated inputs. The documents, tickets, or records in the demo were chosen because the product handles them beautifully. This is rational: no vendor opens with their hardest case. But it means the demo tells you how the product performs on the vendor's ninetieth percentile, while your production reality will be lived in your own tenth percentile: the scanned fax, the handwritten note, the contract with the amendment stapled backwards.

The rehearsed sequence. The clicks come in an order that was practiced. Features that shine get sequenced early, when attention is highest; features that wobble are "something we can show in a follow-up." Watch for the demo that moves briskly past a screen you wanted to linger on. The pace is not accidental.

The happy path. Every workflow shown completes. No lookup fails, no document is rejected, no confidence score dips below the threshold. Production is made of unhappy paths: the exception, the timeout, the ambiguous match. A demo that contains zero failures is not evidence the product does not fail; it is evidence the failures were edited out.

Latency laundering. Processing time is hidden inside conversation. The engineer starts a job, then asks about your quarter, and by the time you finish answering, the result is on screen. You never experience the wait your team will experience four hundred times a day. Multiply an unnoticed eight seconds by your real volume and you have found a full-time employee's worth of waiting.

The engineer drives. The person operating the product is the person who knows exactly which buttons never to press. Your operators will not have that map. A product that is impressive in expert hands and baffling in normal hands demos identically to a product that is good, as long as the expert is driving.

Off-screen regeneration. Generative products produce variable output. The polished summary you are shown may be the best of three attempts, with the other two discarded before the meeting. You saw the maximum, and you will buy the average.

Chapter 1 gave you the concept: the pilot that demos well and dies in production is the signature arc of the 95 percent, because the demo and production are two different environments. That lesson explained the cliff. This one hands you the meeting-level instrument that keeps you from stepping off it.

The Five Conditions That Decide What a Demo Proves

A demo is not good or bad. It is informative or uninformative, and five conditions decide which. Together they form this lesson's artifact, the Demo-Conditions Checklist, which you will fill in during the meeting itself.

Condition one: whose data

A demo on the vendor's samples proves the product works somewhere. A demo on your typical samples proves it works here. A demo on your worst samples proves it will survive here. That is the full hierarchy, and every step up transfers more information. The move is simple and should be arranged before the meeting: send or bring twenty real examples from your own operation, chosen by the person who handles your exceptions, not by anyone motivated to be polite. Five of them should be the ugly ones: the formats that break your current process today. Pass: the vendor runs your samples live, including the ugly ones. Flag: your samples are accepted "for a follow-up analysis" that arrives as a slide deck. Fail: the demo runs only on their data, and your samples are deflected.

Condition two: who drives

Ask, politely and mid-meeting, for the mouse: "Can our coordinator run the next one?" Watch what changes. If a reasonably competent member of your team can operate the core loop with two minutes of guidance, the product's usability is real. If the request produces visible discomfort, or a detour into "we'd normally do enablement first," you have learned that the product's smoothness lives in the engineer's fingers. Pass: your person drives a full cycle unassisted. Flag: your person drives with continuous coaching. Fail: the mouse never crosses the table.

Condition three: what volume

Three samples is an anecdote. Your Tuesday is a queue. The question that converts one into the other: "What does this look like processing an hour of our real intake, and can we watch?" Volume exposes what single samples hide: latency at load, rate limits, the exception percentage, what the queue screen actually looks like when eleven items need human review. If a live volume run is impractical in the meeting, the pass condition becomes a scheduled one: a structured trial at production volume with logging on. Pass: volume behavior demonstrated or contractually scheduled with metrics defined. Flag: "it scales" asserted with a reference to another customer. Fail: the conversation stays at three samples.

Condition four: which edge cases

Before the meeting, write down the three input types that break your current process. You know them by heart: the supplier who sends photographed paper, the request that spans two categories, the document in a second language. In the meeting, name them and ask to see each one handled. This is the single highest-information move available to you, because the vendor cannot have rehearsed your specific pathologies. Pass: each edge case is attempted live, with honest results, including failures gracefully handled. Flag: edge cases are acknowledged as "roadmap." Fail: edge cases are reframed as "not typical usage."

Condition five: what failure looks like

This is the condition buyers skip and regret. Say, in these words: "Show me what it looks like when the product gets one wrong, and walk me through the recovery." A mature vendor has a genuine answer: here is the confidence flag, here is the exception queue, here is the correction flow, here is what the correction does or does not teach the system (you know from the learning-gap lesson why that last part decides the verification tax). A vendor who cannot show failure handling has not handled failure, and you are being offered the happy path as the whole product. Pass: a real failure shown, with a designed recovery path. Flag: failure handling described verbally but never shown. Fail: "it really doesn't make mistakes" or any sentence within earshot of that one.

A demo without a failure in it is not a demonstration of the product. It is a demonstration of the editing.

The Three Closing Asks

The checklist's five conditions structure the body of the meeting. Three short asks close it, each designed to extract information the performance format suppresses.

"Run five of our samples cold." Cold means now, unseen, no preprocessing. If your twenty samples were sent ahead, assume the vendor rehearsed on them (you would). Keep five back and produce them in the meeting. The delta between performance on the fifteen they had and the five they did not is one of the most honest numbers you will ever extract from a sales cycle.

"Regenerate that answer twice." For any generative output, ask to see the same input run again, twice. You are not being difficult; you are measuring variance, which is a real property of the product your team will live with. Three near-identical outputs tell you the behavior is stable. Three meaningfully different totals, tones, or extractions tell you that "the answer" you were shown was a sample from a distribution, and your verification process must be designed for the distribution, not the sample.

"Which customers stopped using this, and can we speak to one?" Reference customers are curated exactly like demo data. The churned customer is the uncurated sample. A vendor who engages with this question honestly, even partially ("I can't share names, but I can tell you the two patterns behind most of our churn"), is telling you something valuable about both the product and themselves. A vendor for whom the question produces a scheduling delay that never resolves has also answered it. You will meet this ask again in this chapter's final lesson, where it anchors the full 20-question due-diligence sheet.

The Etiquette of the Prepared Buyer

Everything above can be executed warmly, and should be. You are not trying to humiliate a salesperson; you are trying to buy the real thing at the real price, which is also what the good vendors want to sell you. This is worth internalizing, because hesitation about "being difficult" is the main reason smart buyers sit through theater politely and sign anyway.

Here is the secret the sales side already knows: good vendors love prepared buyers. A buyer who brings real samples, asks for failure handling, and defines trial criteria is a buyer who will actually reach a decision, deploy successfully, renew, and refer. The alternative buyer, the one who claps at the demo and goes silent for six weeks, is the expensive one on both sides. When you produce the checklist, a strong product's team relaxes, because specifics are where strong products win. It is the weak product's team that needs the performance protected.

Two practical courtesies keep the temperature right. First, send the structure ahead: tell the vendor you will bring samples, that you will want a team member to drive part of the session, and that you will ask about failure handling. Ambush proves nothing except that unrehearsed humans stumble, which you already knew. Second, frame every ask operationally, not adversarially: "we're mapping what our verification workload would be" lands better than "prove it," and means the same thing.

Kestrel Parts: The Same Demo, Twice

Here is the checklist's value measured in one composite scenario with realistic numbers. Kestrel Parts is a fictional 600-person industrial distributor drowning in supplier contracts and purchase orders. A vendor demos a contract-analysis product. We will run the meeting twice.

The first telling: the polite audience

The demo is excellent. Two clean contracts, clauses extracted, summaries crisp. The room is impressed; the head of procurement asks about pricing; nobody asks who chose the contracts. Kestrel signs a $120,000 annual pilot. Reality arrives on schedule: 40 percent of Kestrel's supplier paperwork is scanned or handwritten purchase orders, a category the demo never touched, and the product's miss rate on them turns out to be high enough that two analysts now spend their mornings correcting extractions. The verification tax nobody measured eats the time savings nobody baselined. By month seven the pilot is "under review," which you know from Chapter 1 is where pilots go to not-die. Nothing about this required anyone to lie. It only required nobody to test.

The second telling: the checklist in the room

Rewind. Same vendor, same product, same opening act, but this time Kestrel's operations lead arrives with the checklist and twenty real documents, five held back. Condition one: the vendor gamely runs the samples, and the five cold ones include two photographed purchase orders; the product misses about a third of the fields on them, live, in the room. Not a scandal, a data point. Condition five: asked to show failure handling, the engineer displays a confidence threshold and an exception queue, which is genuinely decent design, but the answer to "what does a correction teach the system" is "corrections are logged for our quarterly model review," which the operations lead correctly hears as the learning gap wearing a process costume. The regenerate-twice ask produces three different contract-value totals on one messy amendment, so verification will need to cover numbers, not just clauses.

Nobody storms out. Instead, the deal changes shape. Kestrel proposes what the evidence supports: a four-week paid trial at $8,000, run on 500 of Kestrel's own documents at production mix, with logging on, a baseline already captured (cycle time and error rate for the current manual process), and exit criteria agreed in writing: field-level accuracy above an agreed bar on the full mix, including the scanned tail, and a measured verification workload below two analyst-hours a day. The vendor, notably, agrees; the good ones do. Whatever the trial concludes, Kestrel's decision will now be made by evidence at $8,000 instead of by theater at $120,000. The checklist did not kill the deal. It moved the decision from the vendor's conditions to Kestrel's.

The Artifact: The Demo-Conditions Checklist

Reproduce this on one page and bring it to every demo. Fill it in as the meeting runs; the blank cells at the end are your follow-up agenda.

  • 1. Whose data. Their samples / our samples / our worst samples. Pass: our samples run live, ugly ones included. Flag: taken away for follow-up. Fail: their data only. Notes: ____
  • 2. Who drives. Their engineer / our operator assisted / our operator solo. Pass: our person completes a cycle. Flag: continuous coaching required. Fail: mouse never crosses the table. Notes: ____
  • 3. What volume. Samples / batch / an hour of our real queue. Pass: volume shown or trial-scheduled with metrics. Flag: "it scales" plus a logo. Fail: three samples and a slide. Notes: ____
  • 4. Which edge cases. Our three named breakers, attempted live. Pass: attempted honestly, failures included. Flag: "roadmap." Fail: "not typical usage." Notes: ____
  • 5. What failure looks like. Confidence flags, exception queue, correction flow, what corrections teach. Pass: failure shown with designed recovery. Flag: described, not shown. Fail: "it doesn't really make mistakes." Notes: ____
  • Closing asks. Five cold samples run: Y/N, result ____. Regenerated twice: stable / variable ____. Churned-customer question: engaged / deflected ____.

Scoring is deliberately blunt: any fail, or three or more flags, means no purchase decision leaves this meeting, and the next step is a structured, baselined, paid trial on your data or nothing. One page, ten minutes of preparation, and the single most expensive hour of theater in enterprise software becomes an hour of evidence.

What to Do Monday Morning

The checklist becomes yours the first time it enters a real meeting. The sequence:

  1. Build the one-pager from the artifact section above. Fifteen minutes in a document template, reusable for years.
  2. Assemble your sample pack: twenty real documents or cases from the process you are shopping for, chosen with your exceptions handler, five held back for cold runs. Refresh it quarterly.
  3. Write down your three breakers: the input types that break your current process. These are your condition-four questions, and only you know them.
  4. Send the structure ahead of your next vendor meeting: samples coming, operator will drive a segment, failure handling on the agenda. Watch how the vendor responds to the email; that is data too.
  5. Run the meeting off the checklist and score it before anyone debriefs. First impressions drift within hours; ink does not.
  6. Convert enthusiasm into trial design: if the room still wants the product, the next artifact is a paid trial with your baseline, your mix, and written exit criteria, never a signature. Chapter 1's five questions and this checklist together make that trial nearly impossible to design badly.

Key Takeaways

  • Treat every demo as a performance engineered under vendor-controlled conditions: curated inputs, rehearsed sequences, happy paths, laundered latency, expert drivers, and off-screen regeneration are standard production technique, not deception.
  • Convert the performance into a test by controlling five conditions: whose data runs, who drives, what volume is shown, which edge cases are attempted, and whether failure handling is demonstrated rather than described.
  • Bring twenty real samples chosen by your exceptions handler and hold five back for cold runs; the gap between rehearsed and cold performance is the most honest number a sales cycle will give you.
  • Ask to see the same generative output regenerated twice, because your team will live with the distribution, not the best-of-three sample the demo showed.
  • Demand a live failure and its recovery path; a vendor who cannot show failure handling has never handled failure, and a demo with zero failures is evidence only of editing.
  • Execute all of it warmly and announce it in advance: good vendors love prepared buyers, and discomfort with specifics is itself a finding about the product.
  • Score the checklist in ink before the debrief, and let any fail or three flags convert the next step into a paid, baselined trial with written exit criteria instead of a signature.
  • Remember Kestrel's arithmetic: the same product, tested instead of watched, moved the decision from $120,000 of theater to $8,000 of evidence, and the checklist is what moved it.