←
AI Readiness & Process Transformation
Aware · M5 · lesson 5 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Benchmarks, Case Studies, and Other Fiction
📖
now learning

Benchmarks, Case Studies, and Other Fiction

15 min

The follow-up email arrives the morning after the demo, and it is friendly and fatal. Two attachments. The first is a one-page benchmark sheet: "98.6% extraction accuracy, independently validated." The second is a glossy case study: a logo you recognize, a quote from a VP of Operations, and a headline number, "invoice processing time cut 62%." Your steering committee will read both documents in under four minutes, and unless someone in the room knows how these artifacts are manufactured, both will be entered into evidence as facts about your future. They are not facts about your future. They are facts about someone else's past, produced under conditions you were never shown, and the previous lesson's rule applies with new force here: the demo was something you watched, but benchmarks and case studies are something you read, and reading them is a separate skill with its own instrument. This lesson builds that instrument.

Two Documents, One Trick

Start with what these two artifacts have in common, because it is the same structural trick wearing two costumes. A benchmark number and a case study are both evidence produced under conditions selected by the seller. The benchmark measures the product against a test set the vendor chose, scored by a method the vendor chose, summarized in a unit the vendor chose. The case study describes a customer the vendor chose, over a period the vendor chose, telling the parts of the story the customer's legal and marketing teams approved. Neither document is a lie, usually. Both are curations, always.

That matters because of what you are actually trying to predict. You do not need to know whether the product performed well somewhere. You need to know whether it will perform well on your documents, your exception rate, your staff, your systems, and your volumes. Transfer is the entire question. And the two most common evidence artifacts in AI sales are systematically weak on transfer for reasons you can name, out loud, in the meeting, without any technical background. This lesson gives you those names.

Keep the program's standing rule in view the whole way through: treat every vendor and research performance figure as a benchmark to verify, never a guarantee. That rule is not cynicism. MIT's GenAI Divide work found that roughly 95 percent of enterprise GenAI pilots produced no measurable P&L return, and most of those pilots were purchased on exactly the kind of evidence you are holding: a strong number, a strong logo, and no test under the buyer's own conditions. The 95 percent read these documents too. They just read them as promises.

Why Benchmark Numbers Do Not Travel

A benchmark is a score on a fixed test set. That sentence contains the whole problem, so take it slowly.

The test set is fixed, which means it was assembled once, from some population of documents or tasks, at some point in time. Your work is not fixed. It is a live distribution with a long, ugly tail: the fax-scanned delivery note, the handwritten correction in the margin, the supplier who puts the PO number in the footer, the invoice photographed on a phone in a warehouse at dusk. When a vendor says "99.2 percent accuracy," the number is true of the overlap between their test set and their product. It says nothing, mathematically nothing, about the part of your distribution their test set never contained. The question that opens the gap is always the same: accurate on whose documents? If the test set was 5,000 clean, typed, English-language invoices and your inbound mix is one-third scans, the number was measured on a different job than the one you are hiring for.

Then comes the second axis: measured how? "Accuracy" is not one measurement. Extraction accuracy can be scored as exact match (the extracted value must be character-perfect) or fuzzy match (close counts). It can be scored per field, per line item, or per document. It can include or exclude documents the system refused to process, and that exclusion is a beautiful place to hide failure: a system that declines the hardest 10 percent of documents and scores 99 percent on the rest is not a 99 percent system, it is a 89-percent-with-good-manners system. None of these choices is visible in the headline number. All of them are visible the moment you ask for the measurement protocol, which reputable vendors have and will share, and which the other kind will describe as proprietary.

And the third axis: counting what as a unit? Their test set is their denominator. If it underrepresents your hard cases, the score overrepresents your future. A vendor whose test set is 92 percent typed documents is honestly reporting a number that will honestly collapse on a customer whose mix is 61 percent typed. Nobody lied. The number just did not travel, because benchmark numbers are properties of a pairing (this product, that test set), not properties of the product alone.

The Units-of-Accuracy Shell Game

One of these measurement choices is so consequential, and so routinely exploited, that it needs its own section: the unit. Character accuracy, field accuracy, and document accuracy sound like siblings. They can differ by more than 30 points on the same system, on the same documents, on the same day.

Here is the arithmetic, and it is the single most useful piece of math in this chapter. Suppose a system extracts fields from invoices with 98 percent field-level accuracy, and suppose a typical invoice has 15 fields you care about: supplier, date, PO number, currency, line totals, tax, and so on. For the document to be usable straight through, all 15 fields must be right. If errors were spread evenly, the chance of a fully correct document is 0.98 multiplied by itself 15 times, which is about 74 percent. A number that sounded like near-perfection describes a process where one invoice in four needs a human. Drop field accuracy to 95 percent and the straight-through rate falls to about 46 percent: a coin flip wearing a tuxedo.

Field-level accuracyFields per documentFully correct documents (approx.)
99%1586%
98%1574%
95%1546%

The real world is slightly kinder than this arithmetic (errors cluster on bad documents rather than spreading evenly), but the direction and the magnitude of the gap are exactly right, and the operational point survives every refinement: your staffing plan runs on document accuracy, because a document with one wrong field still needs a person. Your exception queue, your rework loop, your cycle time: all of them are document-level realities. When a vendor quotes field accuracy, or better yet character accuracy (which is higher still, since most characters in a wrong field are correct), they are quoting the unit that flatters the product, and your cost model will be built in the unit that staffs the queue. Asking "is that character, field, or document accuracy, and how many fields per document?" is a fifteen-second question that has repriced entire deals.

Why Case Studies Do Not Travel

The case study fails to transfer for four different structural reasons, and each has a name you can use in the meeting.

Survivor selection

A case study is, by construction, drawn from the customers who succeeded and agreed to talk about it. You will never be handed the case study of the churned customer, and the vendor's marketing team is not obligated to mention that the featured deployment is one of the two, out of eleven attempts, that reached production. This is the same selection logic the program flagged in the agent-washing lesson, where Gartner found that of thousands of vendors claiming agentic products, only about 130 were judged real: the market shows you its survivors and its costumes, never its casualty list. A case study is a survivor telling a survivor's story. Useful, but as a sample of one from a population you cannot see.

The unnamed baseline

"40 percent faster" is a fraction with a hidden denominator. Faster than what, measured when, by whom? If the baseline was the customer's old process at its worst, the week before go-live, measured by the vendor's own success team, the 40 percent is real and nearly meaningless. Most case-study numbers name no baseline period, no measurement method, and no measurer. In your own organization you would never accept a KPI improvement claim without its baseline; extend the same courtesy of suspicion to a PDF.

The confound

Read case studies for what else changed, because something else always changed. The lighthouse customer did not just install software: they hired three process analysts, standardized their intake templates, renegotiated supplier document formats, and rebuilt their exception workflow, usually with the vendor's best implementation team living in their office. Then the tool got the headline. This is not fraud; it is the ordinary confounding of any uncontrolled before-and-after story, and it is the 10-20-70 rule from earlier in this program appearing in evidence form: if 70 percent of results come from people and process change, then most of a case study's headline number was purchased with work the case study barely mentions, and that you would have to repeat.

The scale and attention mismatch

The lighthouse ran 200,000 documents a month and got a dedicated vendor engineering pod, a named executive sponsor on the vendor side, and quarterly roadmap influence. You will run 3,000 a month and get a shared success manager and a ticket queue. The product may be identical; the deployment is not, and the case study's results were produced by the deployment, not the product. Ask directly: "what did this customer get from you that a customer our size does not?" The honest answer is always instructive.

A benchmark is their test, in their units. A case study is their best customer, minus the parts that would help you. Evidence starts when your data enters the room.

The Honest Uses, and the Evidence Ladder

None of this means benchmarks and case studies are worthless. It means they are rungs near the bottom of a ladder, and the failure mode is treating them as the top.

The honest use of a benchmark is shortlisting. If one category of tool scores 90-plus on public document-extraction tests and another category scores 60, the benchmark has told you something real about categories, and it has saved you from evaluating the wrong species of product. What it cannot do is separate the top three vendors on your shortlist for your workload, because at that resolution the test-set differences dominate the product differences.

The honest use of a case study is pattern mining, not proof. Read it for integration architecture: what systems did they connect, in what order, with what team? Read it for sequencing: what did they pilot first? And above all, use it as a directory of phone numbers. A reference call is a case study with the marketing layer removed, if you ask questions the PDF cannot answer: "What did the first three months actually look like?" "Who owned exceptions, and how many people is that today?" "What did you change about your own process to make this work?" "What almost made you quit?" A reference who has no answer to "what almost made you quit" has been media-trained, and that is data too.

Rank every piece of vendor evidence on this ladder, top rung strongest:

  1. A trial on your data: your document mix, your edge cases, scored in document-level units against your baseline, with exit criteria written before it starts.
  2. Structured reference calls: two or three customers your size, your industry if possible, interviewed with a fixed question set including the failure questions.
  3. The case study: mined for integration patterns, confounds, and phone numbers, never for its headline number.
  4. The benchmark: used to shortlist categories, then retired from the conversation.

Every rung down is a discount on evidential weight. A steering committee about to sign on rungs three and four should be told, in these words, that it is buying on the two weakest forms of evidence available, when rung one costs a few weeks and is routinely granted to buyers who ask.

The Artifact: The Claim-Deflation Sheet

This lesson's artifact is a one-page table you fill in during procurement, one row per vendor claim. Four columns: the claim as stated, what it was measured on and by whom (which the vendor must answer, and "unknown" is an answer you record), the question that deflates it, and the your-conditions test that replaces it. The discipline of the sheet is that no claim graduates from column one to your business case until column four has been run. Here it is, filled in for the six claim types you will meet most often.

The claim as statedMeasured on what, by whom?The deflating questionThe your-conditions test
"98.6% extraction accuracy"Vendor's own test set; composition and scoring unit undisclosed until asked"Character, field, or document accuracy? On what document mix? Were refused documents excluded?"200-plus document sample from your real mix, scored document-level, exceptions counted
"Cuts processing time 62%"One customer's before/after, baseline period and measurer unnamed"62% against what baseline, measured when, by whom, and what else changed at the same time?"Time-and-motion on your own pilot lane against your pre-pilot baseline
"7x ROI in year one"Vendor ROI calculator with vendor-chosen assumptions"Show the model: which costs are in, is verification labor in, what utilization does it assume?"Rebuild the ROI model with your fully loaded costs (next lesson gives you the method)
"92% user adoption"Logins or seat activations, usually in week two"Adoption measured as logins or as work completed through the tool, and at month six?"Your pilot's value metrics: hours redeployed, volume through the tool, sustained at month three
"It learns your business"Usually nothing measurable; a roadmap verb"When my team corrects an output, what specifically improves, and after how many corrections?"Correct the same error 20 times in the trial; document whether recurrence falls
Named-logo case studyOne surviving customer, marketing-approved narrative"Can we call them? What did they change besides installing you? What did they get that we won't?"Two structured reference calls with the failure questions, customers your size

Notice what the sheet does to a meeting. It does not accuse anyone of lying, which keeps the relationship workable. It converts every claim into a measurement question and a test, which is a register vendors with good products handle easily and vendors with good slideware handle badly. Like the Demo-Conditions Checklist from the previous lesson, its diagnostic value is double: the filled sheet tells you about the product, and the vendor's reaction to the sheet tells you about the vendor.

Tolvane Systems: A Full Deflation Walk

Here is the sheet in action, in a composite, hypothetical scenario with realistic numbers. Every company in it is fictional.

A mid-market logistics firm processes 3,000 supplier invoices a month through a five-person accounts-payable team. Tolvane Systems, an AI document-processing vendor, bids with two artifacts: a benchmark, "98.6 percent extraction accuracy," and a case study, "Acme Freight cut invoice processing time 62 percent." The AP manager, who has taken this course, opens a Claim-Deflation Sheet instead of a bottle of champagne.

Deflating the benchmark. Column two: measured on what? Tolvane's answer, obtained politely on the second ask, is that 98.6 is field-level accuracy on their standard test set. Composition of the test set: 92 percent typed, machine-generated invoices. The firm samples its own inbound month: 61 percent typed, 28 percent scanned, 11 percent photographed or faxed. Column four: a trial on 200 invoices drawn proportionally from the firm's real mix, scored at document level (every required field correct or the document counts as an exception). Result: 81 percent document-level accuracy. Not a scandal, and honestly not a bad system. Just a different number than 98.6, and the difference is where the staffing lives.

Deflating the case study. The firm asks for a reference call with Acme Freight and asks the failure questions. Three facts surface that the PDF omitted. First, Acme standardized invoice templates with its top 20 carriers before go-live, which converted a third of its volume to near-perfect input; the reference estimates the process change did "at least half" of the 62 percent. Second, Acme staffed a two-person exception desk that still exists, three years in. Third, Acme processes 200,000 invoices a month and had a Tolvane implementation pod on site for a quarter; the firm, at 3,000 a month, will get a shared success manager. The 62 percent was real. It was also purchased with work and attention the firm would have to fund itself, and the tool alone did not produce it.

Repricing the deal. Now the arithmetic that changes the contract. At 81 percent document accuracy on 3,000 invoices, 570 invoices a month become exceptions. At roughly 12 minutes per exception (retrieve, correct, revalidate), that is about 114 hours a month, which against roughly 140 productive hours per person is 0.8 of a full-time employee, permanently, a cost line absent from Tolvane's ROI slide. The firm does not walk away; the trial showed a workable system. It renegotiates: pricing gated to measured document-level accuracy on the firm's own mix, reviewed quarterly, with a rate step-down if accuracy stays below 90 percent and an exit clause below 75. Tolvane accepts, which tells you Tolvane believed its own product under real conditions. That acceptance is itself the final test result.

LineVendor's evidence saidYour-conditions test said
Accuracy98.6% (field-level, 92% typed test set)81% document-level on your 61%-typed mix
Exceptions at 3,000/monthNot mentioned570/month, about 114 hours, 0.8 FTE
Case-study driver"The tool cut time 62%"Template standardization plus a 2-person exception desk did roughly half
ContractFlat licenseAccuracy-gated pricing, quarterly measurement, exit below 75%

Total elapsed time for the deflation walk: about four weeks. Total cost: one 200-invoice trial and two phone calls. What it bought: a real number to staff against, a contract that pays for performance delivered rather than performance advertised, and an exception desk that was planned instead of discovered. The alternative timeline, the one where the committee signed on the PDF, discovers the 0.8 FTE in month five and calls it a surprise. It was never a surprise. It was arithmetic nobody ran.

What to Do Monday Morning

  1. Pull the evidence pack from any live vendor conversation (or the last one you lost sleep over) and list every quantitative claim in it: accuracy, time saved, ROI, adoption. That list is column one of your Claim-Deflation Sheet.
  2. Fill column two by asking, in writing: what was this measured on, in what unit, by whom, and when? Record "declined to say" verbatim where it occurs; it is the most informative cell on the sheet.
  3. Run the units question on any accuracy claim: character, field, or document, and how many fields per document? Then do the multiplication yourself and write the document-level estimate next to the vendor's number.
  4. Ask for two reference customers your size, and run the failure questions: what did the first three months look like, who owns exceptions today, what did you change about your own process, what almost made you quit?
  5. Propose the rung-one test: a 200-sample trial on your real document mix, scored document-level, with the acceptance thresholds written down before it starts. Attach the Demo-Conditions Checklist from the previous lesson as the trial's terms of reference.
  6. Bring the filled sheet to the steering committee and put the vendor's number and your-conditions number side by side. The gap between them, priced in FTEs, is the agenda.

Key Takeaways

  • Treat benchmarks and case studies as evidence produced under seller-selected conditions: true somewhere, curated always, and weak on the only question that matters, which is transfer to your data, volumes, and team.
  • Interrogate every benchmark on three axes: whose documents (their fixed test set versus your live distribution and its scanned, faxed tail), measured how (exact versus fuzzy, refusals excluded or counted), and in what unit.
  • Run the units arithmetic yourself: 98 percent field accuracy on 15-field documents is roughly 74 percent document accuracy, and your staffing, queues, and cycle times live at document level, not field level.
  • Name the four transfer failures of case studies out loud: survivor selection, the unnamed baseline, the confound (the customer also changed process and people, then the tool got the headline), and the scale and attention mismatch.
  • Use benchmarks only to shortlist categories and case studies only to mine integration patterns and reference phone numbers; call references and ask what the first three months looked like, who owns exceptions, and what almost made them quit.
  • Climb the evidence ladder before money moves: a trial on your data beats structured reference calls, which beat the case study, which beats the benchmark, and each rung down is a discount on evidential weight.
  • Fill a Claim-Deflation Sheet for every material claim: the claim as stated, measured on what and by whom, the deflating question, and the your-conditions test that replaces it; no claim enters the business case untested.
  • Convert your-conditions results into contract terms, the way the Tolvane walk did: accuracy-gated pricing, quarterly measurement on your mix, and exception staffing planned up front instead of discovered in month five; the next lesson does the same surgery on the ROI model itself.