The Pilot That Demos Well and Dies in Production
The demo happens on a Thursday afternoon, and it is genuinely impressive. The vendor's solutions engineer drags a supplier invoice onto the screen and twelve fields fill themselves in about four seconds: supplier name, PO number, line items, tax, total. He does it again with a second invoice, then a third. Somebody in the room actually applauds. The head of shared services, who has spent six years watching a five-person team key invoices by hand, feels something she has not felt about this process in a decade: hope. The contract is signed within three weeks. Fourteen months later, the tool processes nine percent of invoice volume, one analyst spends most of her week correcting what it gets wrong, and the project has not been mentioned in a steering meeting since spring. Nothing about that arc was bad luck. MIT's GenAI Divide research found that only about 5 percent of custom enterprise AI tools ever cross from pilot into production, and the distance between the Thursday demo and the month-six silence is not a mystery. It is a measurable gap between two different environments, and this lesson teaches you to measure it before anyone signs anything.
The Demo Is a Designed Environment
Start with the correction that reframes everything else in this lesson: a demo is not a small version of production. It is a different environment, built for a different purpose, and it should be evaluated the way you would evaluate any other environment built for a purpose. Production exists to produce work. A demo exists to produce a decision, specifically a yes. Every design choice inside a demo serves that goal, and once you list the choices, you will never watch a demo the same way again.
- Curated data. The documents, tickets, or records in the demo were selected, and often cleaned, in advance. If the vendor asked you for samples beforehand, someone on your side sent the good ones; if the vendor supplied their own, they have run those exact files a hundred times. The data that will actually hit the system, the scanned fax, the handwritten margin note, the supplier who invoices in a spreadsheet, is nowhere in the room.
- The happy path. The click sequence you watch has been rehearsed to avoid every branch where the product struggles. Demos do not wander; they follow a route the way a stage play follows a script, and the route was chosen precisely because nothing interesting goes wrong on it.
- The expert at the keyboard. The person driving is a solutions engineer who knows exactly what to type, what not to click, and how to phrase the prompt so the tool shines. In production, the driver is a distracted employee with eleven browser tabs open and a queue behind them.
- No volume. You watch three documents. Production is eight thousand a month. Every cost that is invisible at three, the per-transaction fee, the thirty seconds of human checking, the occasional timeout, multiplies by thousands the moment the pilot meets real throughput.
- No integration. The demo runs in a standalone screen. Nothing has to read from your ERP, write to your document archive, respect your access controls, or survive your network. Integration is where the majority of implementation pain lives, and it is precisely the part a demo cannot show.
- Charming errors. When the tool does stumble in a demo, the error is laughed off as endearing ("it's still learning!"). In production the same error is a mis-paid invoice, a compliance exposure, or forty minutes of someone's afternoon, and nobody laughs.
An analogy from a different industry makes the pattern intuitive. Walk through a builder's model home and notice the staging: the furniture is slightly undersized so the rooms read larger, nobody has ever cooked in the kitchen, the closets hold six beautiful garments instead of a family's actual belongings, and the light is always afternoon light. None of this is fraud. Staging is what a rational seller does. But no sane buyer prices the house off the staging; they check the plumbing, the neighborhood at night, the commute at rush hour. A vendor demo is staging, and it is exactly as honest and exactly as incomplete. Your job is not to resent it. Your job is to refuse to price the pilot from it.
Week-Two Magic, Month-Six Silence
When a pilot is launched on the strength of demo conditions, its decline follows an arc so consistent that you can date the milestones in advance. Learning the arc matters because each stage produces signals, and the signals arrive early enough to act on if you know what they mean.
Weeks one and two: magic. Onboarding runs on friendly data, often the same curated samples from the sales cycle. Early users are enthusiasts who forgive everything. Screenshots go to the steering committee. Somebody uses the word "game-changer" in an email that will be quietly embarrassing later. This is the demo environment persisting a little longer, not production beginning.
Month two: first contact. The real data stream arrives and the first exceptions appear: the format the model has never seen, the supplier who does things differently, the request that spans two systems. Individually each one looks minor, and the team invents the first workaround: "just route those around the tool for now." Write that sentence down when you hear it, because it is the sound of the pilot beginning to shrink.
Months three and four: the quiet inversion. The exception queue becomes a standing pile of work assigned to someone whose time was never in the business case. Users have been personally burned by a specific bad output, and trust does not recover on the schedule enthusiasm was built on. Usage narrows to the easy cases, which are precisely the cases that had the least value in automating. The tool is now doing the work that was cheap and avoiding the work that was expensive, which is the exact inverse of the pitch.
Month six: silence. Nobody presents the project anymore. Its status line reads "still in pilot," which is the project-portfolio equivalent of a hospice. The renewal invoice arrives and forces the only conversation anyone has had about it in a quarter. The pilot is not killed, because killing it would require someone to narrate this whole arc out loud; it persists as what the previous lesson taught you to call a zombie line item.
You already know the aggregate statistics from the previous lesson, so a single line of context is enough here: 95 percent of enterprise GenAI pilots showed no measurable P&L return, and inside that finding sits this lesson's sharper number, MIT's observation that only around 5 percent of custom tools ever cross from pilot to production. Gartner, watching the same cliff from a different angle, predicted that 30 percent of GenAI projects would be abandoned after proof of concept by the end of 2025. The cliff is not an anomaly. It is the default outcome whenever an organization prices a production commitment using demo evidence.
The demo shows you the tool at its best, on their data, at no volume. Production bills you for the tool at its worst, on your data, at full volume, forever.
What Production Does That the Demo Never Did
The arc has mechanics underneath it. Six specific forces show up in production that were structurally absent from the demo, and each one is worth recognizing on sight because each one generates its own line of hidden cost.
- Edge cases. Every real process has a long tail of weird: the credit note that references three invoices, the customer with two conflicting records, the document in the wrong language. The tail is small by count and enormous by effort; a process where 5 percent of cases consume half the team's time is normal, not exceptional. Edge cases are usually the reason the process needed skilled humans in the first place, which means they are exactly where an automation earns or loses its keep, and exactly what the demo removed.
- Data variance. Demo data was clean and frozen. Production data has hundreds of source formats, degrades through scanning and forwarding, and drifts over time as suppliers, customers, and upstream systems change things without telling you. A model tuned to last quarter's inputs meets next quarter's inputs on no particular schedule.
- Volume economics. Per-unit costs that round to zero at demo scale become budget lines at production scale. A thirty-second human check is nothing on three documents and roughly 67 hours a month on eight thousand. Per-page processing fees, retry costs, and storage do the same multiplication. Nothing about the demo required anyone to do this arithmetic, so in most pilots nobody did.
- Exception handling. Automation splits your work stream in two: the cases the tool handles and the exceptions it hands back. Here is the trap almost every business case misses: an exception is usually more expensive per unit than the old manual process, because the human must first discover what the tool got wrong, then investigate, then correct, then re-enter. If the exception rate is high enough, the pilot raises your cost per transaction while the dashboard celebrates automation.
- Integration under load. A copy-paste bridge that works for ten documents a day collapses at eight thousand a month, a pattern you met in the previous lesson. But even real integrations behave differently under load: authentication expires mid-batch, APIs throttle, timeouts cascade, two systems disagree about which record is current. The demo ran on a laptop; production runs on your infrastructure on your worst day.
- The verification burden. This is the deepest one. AI errors do not announce themselves; a wrong invoice total looks exactly like a right one. At scale you must choose a verification posture: check everything (and pay for it in labor), check a sample (and pay for it in escaped errors), or check nothing (and pay for it in incidents). The demo had a fourth option, audience trust, and that option does not exist in production.
Because these six forces recur in every AI initiative you will ever assess, compress them into a reusable mental model: two columns, demo conditions and production conditions. Whenever anyone shows you a result, your first classification question is which column the result came from.
| Dimension | Demo conditions | Production conditions |
|---|---|---|
| Data | Curated, cleaned, frozen samples | Full variance, degrading quality, drifting formats |
| Volume | A handful of cases, no time pressure | Thousands per month, backlog dynamics |
| Operator | Vendor expert who knows the safe path | Busy employees with no rehearsal |
| Edge cases | Removed in advance | Arrive daily and consume disproportionate effort |
| Integration | Standalone screen, nothing connected | Must read and write systems of record under load |
| Errors | Rare, charming, cost nothing | Silent, look like correct output, cost real money |
| Verification | Audience trust | A designed control someone staffs and pays for |
| Economics | Free; costs invisible at tiny scale | Per-transaction costs multiplied by everything |
Keep the table. It is not a one-lesson prop; it is the lens you will use in every vendor meeting for the rest of your career. Any claim, any metric, any glowing anecdote gets sorted into a column before it gets believed.
The Arithmetic of the Cliff: A Worked Example
Now run the numbers on one composite pilot. The scenario is hypothetical and the figures are chosen to be round and realistic, but the pattern is assembled directly from the documented failure modes, and the arithmetic is the part you should steal.
A regional insurance services firm processes 8,000 supplier and claims-related documents a month through a five-person team. Manual handling averages six minutes per document, about 800 hours of work a month, with each processor costing roughly $32 an hour fully loaded (salary plus benefits plus overhead), around $61,000 a year. A document-extraction vendor demos on 20 samples. The samples came from the firm itself: when the vendor asked for examples, the operations manager sent twenty clean, recent, single-page PDFs from the largest supplier, because those were easy to find and nothing confidential was in them. The demo extracts with 96 percent accuracy. The business case writes itself: automate extraction, redeploy three of the five processors, save roughly $180,000 a year against an $84,000 annual license. Signed.
Production month two: the real stream turns out to contain over 300 distinct supplier formats, plus scans of scans, multi-page attachments, credit notes, and the occasional handwritten correction. A sampled audit puts real-world performance at 84 percent: sixteen of every hundred documents come back with at least one field wrong or missing. Nobody lied. The tool is performing exactly as well as it performed in the demo, on documents like the demo's. The firm simply never tested it on documents like its own.
Now the arithmetic, and notice that every step is multiplication a spreadsheet could have done before signature.
- Exception volume. 16 percent of 8,000 documents is 1,280 exceptions a month.
- Exception unit cost. Handling an exception takes about 20 minutes: discover which field is wrong, pull the source document, investigate, correct, re-enter. At $32 an hour that is $10.67 per exception, versus $3.20 for the old six-minute manual entry. For those 1,280 documents, the firm now pays more than triple its pre-AI unit cost.
- Exception labor. 1,280 exceptions at 20 minutes is about 427 hours a month, roughly 2.7 full-time equivalents. The exception desk was in nobody's business case, and it is now the largest single line in the pilot's true cost.
- Verification. Because errors do not flag themselves, the firm institutes a 30-second check of three critical fields on every document: 8,000 checks a month is about 67 hours, another 0.4 FTE.
- Process ownership. Someone must monitor accuracy, file vendor tickets, and retrain templates: a quarter of one analyst, 0.25 FTE.
Total labor retained: roughly 3.3 of the original 5 FTE. Actual redeployment: about 1.7 FTE, worth around $104,000 a year, against the $84,000 license. Net saving: about $20,000 a year, on a business case that promised $180,000. Then month five delivers the drift event: the firm's largest supplier redesigns its invoice template, accuracy on that supplier's 900 monthly documents quietly drops toward 60 percent, and three weeks pass before anyone traces the swelling exception queue back to the cause, because drift detection was never anyone's job. By month six the pilot is not a scandal; it is something worse for its sponsors, a rounding error. It gets "put under review," and the review never ends.
One more cut of the same numbers, because it is the frame you will use in the five questions below: express the error rate as dollars per thousand transactions. At the demo's 96 percent, 40 exceptions per thousand documents at $10.67 each is about $427 per thousand. At production's 84 percent, 160 exceptions per thousand is about $1,707. The twelve accuracy points between the demo and reality are worth roughly $1,280 per thousand documents, which at 96,000 documents a year is about $123,000. Accuracy points are dollar figures wearing disguises, and the exchange rate is knowable in advance.
The Artifact: The Five Questions That Predict the Cliff
Here is this lesson's deliverable for your readiness portfolio: five questions that convert the demo-versus-production gap from a surprise into a measurement. They are pre-commitment questions. Their entire power comes from being asked, answered in writing, and priced before the contract is signed or the pilot is chartered, because afterward every answer becomes a negotiation with sunk costs. A vendor or internal team that cannot answer them has told you the pilot is running on demo evidence; a vendor that refuses the first one has answered it.
Question 1: Will you run it on our worst 50 cases?
Not the vendor's samples. Not your twenty cleanest PDFs. The fifty ugliest, most representative cases your process actually produced, pulled from the exception folder, the escalation queue, and the complaints file. Structure it as a scored trial: same cases, defined fields, accuracy counted per document, results in writing. Mini example: the insurance firm's operations manager pulls fifty documents from last quarter's manual escalations: faded scans, multi-page claims, the spreadsheet-invoice supplier. The tool that demoed at 96 percent scores 71 percent on this set. That number, not the demo's, goes into the business case, and suddenly the business case tells the truth.
Question 2: Who handles the exceptions, and at what cost each?
Every automation produces an exception stream. Demand a name, not a shrug: which role receives the exceptions, through what queue, and how long does one take end to end? Then multiply. Mini example: time one exception with a stopwatch: 20 minutes at $32 an hour is about $11. At a 12 percent exception rate on 8,000 monthly documents, that is 960 exceptions and roughly $10,200 a month, about $122,000 a year. If that figure is not in the business case, the business case is fiction, and if the named handler's manager has not agreed to it, the exception desk will be staffed by resentment.
Question 3: What happens at ten times the pilot volume?
Pilots run at boutique scale; production does not. Ask for the arithmetic at 10x: per-transaction fees, latency and throttling limits, queue behavior on the Monday after a long weekend, and above all human review capacity. Mini example: the pilot verifies 800 documents a month with one reviewer spending 30 seconds each, under 7 hours, invisible. At production's 8,000 documents that same check is 67 hours a month, and at a 10x growth scenario it is a full-time job and a half that nobody has hired. If the scaling math has never been written down, the pilot's economics are a photograph of a moment, not a forecast.
Question 4: What breaks when the input format drifts?
Production inputs change without notice: suppliers redesign templates, upstream teams add a field, a regulator alters a form. Ask three sub-questions: how would we detect that accuracy has dropped, how fast, and what is the blast radius until we do? Mini example: the firm's biggest supplier changes its invoice layout in month five. If detection is "the monthly accuracy audit," a drift event costs up to a month of degraded output, which at 900 affected documents and a 40-point accuracy drop is hundreds of silent errors. The mature answer names a monitoring metric, an alert threshold, and an owner. The immature answer is "the vendor handles that," which means nobody does.
Question 5: What does the error rate cost per thousand transactions?
This question forces two disciplines at once. First, pin down the unit: 96 percent of what? Field-level accuracy and document-level accuracy are wildly different claims; a twelve-field document at 96 percent per field is fully correct only about 61 percent of the time (0.96 multiplied by itself twelve times). Second, convert the honest rate into money. Mini example: at 84 percent document accuracy, 160 exceptions per thousand at $10.67 each is about $1,707 per thousand documents, before counting the errors that escape verification and surface downstream as mis-payments and supplier calls. Run the same formula at the vendor's claimed rate and at your worst-50 rate, and the gap between the two lines is the risk you are being asked to sign for.
The Same Pilot, Run Right
To prove the five questions are a design tool and not just a veto, replay the insurance firm's pilot with the questions asked first. The worst-50 trial scores 71 percent, so the firm does not conclude "reject the tool"; it concludes "reject this scope." The trial data shows accuracy is excellent on the three highest-volume suppliers, whose documents are structured and consistent: 97 percent document-level. So the pilot charter is rewritten: automate those three suppliers only, about 3,100 documents a month, with a named exception handler budgeted at the measured unit cost, a weekly accuracy sample as the drift alarm, and expansion to the next supplier tier only when document-level accuracy on a worst-50 sample of that tier clears a pre-committed threshold. The promise shrinks from $180,000 to about $55,000 a year. The delivery, for once, matches the promise, and eighteen months later the automated share has grown supplier by supplier to over half of volume, each expansion earned with evidence. Smaller promise, real number, compounding trust: that is what the right side of the cliff looks like.
Two systemic notes to file with the artifact. First, BCG's 10-20-70 rule (10 percent of the effort in algorithms, 20 percent in technology and data, 70 percent in people and process) is the cliff's deep explanation: everything that killed the first version of this pilot lived in the 70, exception staffing, verification design, drift ownership, and none of it appears in a demo. Second, MIT found externally partnered deployments succeeding about twice as often as internal builds, but read that finding correctly: partnering does not exempt you from the five questions, it changes who has to answer them. A good partner answers in writing. A staging-only vendor changes the subject, and the subject change is your data point.
The questions work best when they stop being personal heroics and become a stage gate: a checkpoint in your organization's pilot process where written answers to all five, with numbers, are entry criteria for signature. The previous lesson gave you the seven-question autopsy checklist for any AI initiative; these five are its pre-contract twin, aimed at the specific moment when demo enthusiasm is about to be converted into production commitment. That moment is the cheapest place in the entire lifecycle to be right.
What to Do Monday Morning
Turn the lesson into motion while the demo-versus-production lens is fresh.
- Pick one live initiative, a pilot in flight or a proposal on the table, and find out exactly what data its impressive results were produced on. If nobody can tell you, you have learned the most important fact available.
- Build your worst-50 file. Spend an hour pulling the fifty ugliest real cases from the process in question: exceptions, escalations, complaints, the formats everyone dreads. This file is reusable ammunition for every vendor conversation this year.
- Time one exception end to end and price it at a fully loaded rate. One stopwatch session turns "there will be some exceptions" into a unit cost you can multiply.
- Run the per-thousand arithmetic twice, once at the claimed accuracy and once at your honest estimate, and put both lines on one slide. The gap between them is the conversation.
- Install the five questions as entry criteria in the next pilot charter or steering-committee template: written answers, with numbers, before signature. You are not adding bureaucracy; you are moving the autopsy to before the death.
- File the answers in your readiness portfolio next to the failure-mode checklist from the previous lesson; the two instruments together cover an initiative from proposal to post-mortem.
Key Takeaways
- Treat the demo as a designed environment whose product is a yes: curated data, a rehearsed happy path, an expert driver, no volume, no integration, and errors that cost nothing, none of which survives contact with production.
- Recognize the arc in advance: week-two magic, month-two workarounds ("route those around the tool for now"), the months-three-and-four inversion where usage narrows to the lowest-value cases, and month-six silence.
- Anchor the stakes with the record: only about 5 percent of custom enterprise AI tools cross from pilot to production per MIT, and Gartner predicted 30 percent of GenAI projects abandoned after proof of concept by end of 2025.
- Name the six production forces the demo structurally excludes: edge cases, data variance, volume economics, exception handling, integration under load, and the verification burden, and sort every claim you hear into the demo-conditions or production-conditions column.
- Do the cliff arithmetic before signature: exception rate times volume times exception unit cost, plus verification labor, plus drift exposure; in the worked example twelve accuracy points were worth about $123,000 a year.
- Ask the five questions before commitment: run it on our worst 50 cases; who handles exceptions and at what cost each; what happens at 10x volume; what breaks when input formats drift; what does the error rate cost per thousand transactions.
- Pin down accuracy claims to their unit, because 96 percent per field on a twelve-field document means only about 61 percent of documents are fully correct, and vendors rarely volunteer which number they quoted.
- Use the questions to redesign scope, not just to veto: the honest version of the worked pilot shrank its promise to the three suppliers where accuracy held, delivered a real number, and expanded on evidence, which is how the 5 percent behave.
Skill.re