Pilot Design: Scope, Baseline, Success Criteria Before Tools
Pull the file on any AI pilot in your organization, living or dead, and compare two dates. The first is the date the pilot's scope and success criteria were written down. The second is the date of the first vendor demo. If the second date comes before the first, you can predict the rest of the file without reading it: a scope that looks suspiciously like the demo, metrics that look suspiciously like the vendor's dashboard, a success story nobody can wire to a financial statement, and a renewal decision made by inertia. MIT found that 95 percent of enterprise generative AI pilots deliver no measurable return, and the program you are in has spent two levels dissecting why. Here is the shortest version of the answer you will ever get, and it is a sentence about sequencing: the 5 percent decide scope, baseline, and success criteria before any vendor is in the room. The 95 percent let the demo go first. This lesson teaches the design discipline that keeps the dates in the right order.
Tool Gravity: The Force That Designs Your Pilot for You
Every tool that enters a conversation exerts a pull on the experiment being designed around it. Call this force tool gravity, and treat it as a law of nature rather than a moral failing, because it operates through completely rational behavior on both sides of the table.
Start with the vendor's side. A vendor's demo is not a random sample of the product's behavior; it is the most carefully engineered forty-five minutes in the company. The demo dataset is curated, the use case is chosen because the product shines on it, and the flow is rehearsed until every step lands. This is not deception, it is marketing doing its job: the vendor optimizes the demo for what the product does best, exactly the way you would. The problem is what happens next, on your side of the table, if no pilot design exists yet. The demo's use case quietly becomes the pilot's scope, because it is the most vivid version of the product anyone has seen. The vendor's dashboard becomes the pilot's metrics, because it already exists and produces charts on day one. And the product's strengths become the success criteria, because those are the dimensions the demo taught the room to care about. Three design decisions, the three most important decisions in the entire pilot, get made by osmosis, and none of them were made by you.
Now name what that produces. A pilot designed after the demo is an experiment that tests the vendor's story: can this product do the thing the demo showed, on inputs resembling the demo's, measured on dimensions the product is built to score well on? The answer is almost always yes, which is why so many pilots "succeed" and still change nothing, and why McKinsey keeps finding that 88 percent of organizations use AI while only about 39 percent can attribute any earnings impact to it. A pilot designed before the demo is a different instrument entirely: it tests your process's hypothesis. Can median cycle time on your exception queue fall below two days without the rework rate rising? The vendor's product is merely one candidate answer to that question, and it is allowed to fail.
A pilot designed after the demo tests the vendor's story. A pilot designed before the demo tests your hypothesis. The difference is decided by which document exists first.
You already hold the document that governs the pilot: back in Level 2 you built the pilot charter, the signed handoff with its seven clauses, and this program will not re-teach it here. What Level 2 could not yet give you, because the redesign did not exist yet, is the engineering discipline underneath three of those clauses: how to actually choose a scope, how to actually wire criteria to measurement, and how to actually sequence design ahead of procurement now that Chapters 3.1 and 3.2 have given you a redesigned process, an output schema, a verification architecture, and a Measurement Plan full of counters. That discipline lives in this lesson's artifact, the Pilot Design Brief: a short pre-vendor document with four parts, the hypothesis, the scope with its exclusions, the criteria wired to named counters, and the evidence plan. If the charter is the signed contract, the design brief is the set of engineering drawings behind it, and its defining property is a date: it is finished before the first vendor call is booked. The brief is built from three load-bearing decisions, and each one gets its own section.
Decision One: Scope as Hypothesis, Not Ambition
The first decision sounds like drawing a boundary and is actually writing a sentence. A pilot is an experiment, and an experiment tests one thing. Before you can bound the scope, you must be able to state, in a single falsifiable sentence, what the pilot believes. Here is the form, rendered on the running example this chapter has been building:
"We believe AI-assisted categorization of the three structured exception types will cut median cycle time from 3.2 days to under 2.0 days without raising the rework rate, because the queue wait, not reading speed, is the constraint."
Read that sentence slowly, because every clause is load-bearing. It is falsifiable: it names a number that either happens or does not, and there is a world in which the pilot loses. It is mechanism-stated: the "because" clause commits to a theory of why the improvement should occur, which means a pilot that hits the number for the wrong reason gets caught, and a pilot that misses it teaches you which belief was wrong. And it is segment-bounded: it claims the three structured exception types, not "invoice processing," not "finance operations." Test your own draft hypothesis by reading it aloud and asking what evidence would prove it false. If the honest answer is "nothing, really," you have written an ambition, and ambitions cannot be piloted, only funded.
Once the hypothesis exists, four scope rules follow from experimental logic rather than from appetite.
Narrow enough to attribute
One process, one lane, one site. The reason is attribution: when the number moves, you need to be able to say what moved it. Run the pilot across three sites simultaneously and every site difference (staffing, work mix, local workarounds, the enthusiasm of one team lead) becomes a confound you cannot untangle. A multi-site "pilot" is a rollout wearing a lab coat: it has the vocabulary of an experiment and the evidentiary power of none, because it can attribute nothing. The invoice-exception pilot runs on the routine lane at one shared-services site, full stop. If it wins there, scaling is a separate, later decision made on the evidence the narrow pilot produced.
Exclusions are controls, not confessions
The charter carried an OUT list, and in Level 2 the OUT list was mostly a defense against scope creep. Now upgrade its justification: in a designed experiment, exclusions are controls. The free-text exception lane, the messy 22 percent, stays out of this pilot, and the reason is not that it is hard (it is) but that including it would blur the hypothesis. The hypothesis claims a mechanism on structured exceptions; mixing in unstructured ones would smear two different processes into one aggregate number, and whatever that number did, you could not tell which lane did it. Write each exclusion in the brief with its experimental reason attached. An OUT list justified by logic survives the steering-committee member who asks "why so timid?"; an OUT list justified by difficulty invites the scope expansion that kills attribution.
Volume big enough to reach a verdict in the window
Here the sampling arithmetic from the verification lesson gets reused for a new purpose: sizing the experiment. The routine lane runs about 268 items per week, so a six-week pilot yields roughly 1,600 observations. Is that enough? Run the mental check: the error budget cares about movements around a 2 percent verified error rate, and at 1,600 observations a 2 percent rate produces around 32 error events, enough to see the rate move meaningfully rather than jitter. The same logic in reverse is the warning: a process handling 10 items a week produces 60 observations in six weeks, and at those volumes a single unlucky item swings the error rate by nearly 2 points on its own. Say this plainly in your brief and to your sponsor: small-volume processes get slower verdicts. They need a longer window, or a different evidence standard (deeper per-item review instead of rates), or a different pilot candidate. What they cannot have is a six-week pilot with statistical confidence, no matter how much the calendar wants one.
Duration set by the process's natural cycle
Finally, duration is not a round number chosen for the steering calendar; it is dictated by the process's rhythm. The invoice-exception queue spikes at month-end, when accruals and closing pressure change both the volume and the mix of exceptions. A pilot window that dodges month-end has quietly selected the easy weeks, and you already know that trap by name from the Measurement Plan: mix shift. So the rule is structural: the window must span at least one full occurrence of whatever cycle stresses the process (month-end, quarter-end, seasonal peak, Monday floods). This is mix-shift control designed in before launch rather than apologized for after, and it is the difference between a result and an asterisk.
Decision Two: Criteria Wired to Counters Before Commitment
The second decision turns the charter's promises into checkable ones. Here is the wiring standard: every success criterion, and the kill condition, must name four things. Its counter from the Measurement Plan you built in the last lesson. Its threshold. Its measurement window. And its verification layer, meaning which part of the verification architecture produces the number. A criterion missing any of the four is not yet a criterion; it is a sentiment with a percentage attached.
Watch the standard work on the criterion this chapter has carried since Level 2: "85 percent verified accuracy." That phrase is only meaningful because the word "verified" points somewhere specific: at the layer-two-plus-layer-three measurement, the human gate's catches plus the audit layer's 100-document weekly random sample, not at the tool's self-reported confidence score. Unwire it, and watch what happens in month four when the number comes in at 81: someone proposes, quite reasonably, that "verified" should really mean the tool's own scoring, which happens to read 88. If the wiring was never written down, that conversation is a negotiation, and negotiations under pressure go the way of whoever wants the pilot to survive. If the wiring was written down and signed before launch, the conversation lasts one sentence. That is the whole difference between wired and unwired criteria: unwired criteria are renegotiable, wired ones are checkable. One is a contract; the other is a vibe.
The criteria hierarchy: primary, guardrails, kill
Not all criteria do the same job, and a brief that lists eight coequal metrics has designed a committee argument. The hierarchy has three tiers.
One primary criterion. The hypothesis's own number, and only that. For the invoice pilot: median cycle time under 2.0 days on the in-scope lane. The primary is what the pilot is for; everything else exists to make sure the primary is not achieved by cheating.
Two or three guardrails. A guardrail is a metric that must not get worse, and its job is to catch the pilot that hits its target by cannibalizing something unmeasured. Here is the canonical example, and it is worth internalizing because some version of it happens in a large share of "successful" pilots: cycle time improves beautifully, and it improves because the system started auto-approving marginal cases that a human would have questioned. The speed is real; it was purchased with silent errors that will surface downstream as credit notes and supplier disputes in month five. The rework-rate guardrail exists precisely to catch that purchase while the pilot is still running. The invoice brief carries three guardrails: rework rate held at or below 12 percent (the baseline pack measured 11, so the guardrail tolerates noise but not deterioration), the human gate's queue latency under four hours (so speed is not achieved by starving the control you designed), and clerk overtime flat (so the "saving" is not being manufactured by unpaid evenings). Choose your own guardrails by asking one question: if this pilot wanted to fake its primary number, what would it quietly spend? Guard that.
The kill condition as the pre-agreed floor. The charter already carries it: verified accuracy below 85 percent at week six, or any critical escape, and the pilot stops. This lesson's job is to revisit why pre-commitment is the only version of a kill condition that works. A kill condition is signed when it is hypothetical and executed when it is painful; the entire mechanism depends on the signature predating the pain. By week six there are sunk costs, a vendor relationship, a sponsor's credibility, and a team that has worked nights: no floor agreed at that point will be low enough to trigger. Gartner's projection that over 40 percent of agentic AI projects will be canceled by the end of 2027 describes exactly what the alternative looks like: cancellations that arrive late, expensive, and acrimonious, precisely because nothing was pre-committed and every project had to be killed by force after value failed to appear, instead of by a clause everyone signed while it was still cheap to sign. A documented kill on a pre-agreed floor is a win; this program has said so since Level 1, and the design brief is where the floor gets poured.
Decision Three: The Evidence Plan Before the Tool
The third decision is the one that looks most like administration and does the most zombie-prevention per sentence. The evidence plan answers, in writing, before launch: who measures, what gets reported weekly, what the committee sees and when, and what meeting decides the pilot's fate.
Who measures. The charter's role clause already established measurement independence: whoever produces the numbers has no compensation, bonus, or reported performance tied to the pilot's outcome. The design brief names the person and connects them to the Measurement Plan's weekly tally ritual. The independence is not a slur on anyone's honesty; it is the same reason the accountant who cuts the checks does not also reconcile the bank statement.
What the weekly reports. The one-page weekly you specified in the Measurement Plan: volumes with the conservation line, verified error rate against budget, cycle time against baseline by segment, escapes, and ladder status. The brief adds only one thing: the criteria hierarchy printed at the top of the page, primary, guardrails, kill floor, so every weekly is read against the promises rather than against last week's mood.
Checkpoints calendared now. The committee sees the pilot at week three and week six, and both meetings go on the calendar before launch, with the week-three checkpoint framed explicitly as a health check (is the instrumentation producing? are guardrails stable?) rather than an early verdict, because week-three numbers still contain ramp. The reason to book them now is a law of organizational physics you have observed your whole career: a checkpoint scheduled after launch is a checkpoint that slips. Everyone is busy; the pilot is "going fine"; week three becomes week five becomes a quarterly update, and drift has no appointment at which to be noticed.
The decision meeting, booked before the pilot starts. This is the strongest single zombie-prevention device in the entire program, and it costs one calendar invite. Before launch, the scale, iterate, or kill meeting goes on the steering committee's calendar for a named date, with the criteria hierarchy attached to the invite itself. Now the pilot has a property the 95 percent almost never have: it ends at a known time, and something is decided. There is no ambient state in which it can drift, no month nine in which nobody remembers what success was supposed to look like, because success is literally in the meeting invite everyone has been able to read since before day one. The next lessons in this chapter run the pilot's operating rhythm and then that decision meeting itself; the design brief's job is to make the meeting inevitable.
Then, and Only Then, the Vendor Enters
With the brief signed, the sequence completes and procurement begins, and notice how completely the geometry of the vendor conversation has inverted. You are no longer attending the vendor's demo; the vendor is responding to your brief. They receive your scope (the routine lane, three exception types, one site), your output schema from earlier in this chapter, your acceptance samples (the deliberately hard, real documents you assembled when you built the verification architecture: the skewed scan, the multi-currency invoice, the vendor with the ambiguous line items), and your criteria hierarchy. The demo you then ask for is not "show us the product"; it is "run our acceptance samples through your product against our schema, and let us score the output with our own double-check protocol." A demo has just become a structured test.
This is also where two levels of this program click together, and the continuity is worth naming. In Level 1 you built the vendor due-diligence sheet, the discipline of interrogating claims, references, and benchmark conditions. That sheet told you which vendors were worth a conversation. The design brief now governs what the conversation is. Due-diligence sheet plus design brief equals a procurement process in which tool gravity has nothing to grip: the scope is signed, the metrics are wired to your counters, the success criteria predate the demo, and the vendor's product gets evaluated on your process's hardest inputs rather than its own happiest path. Vendors, incidentally, sort themselves fast under this regime. The strong ones engage with the brief, ask sharp questions about your schema, and negotiate honestly about which acceptance samples they will struggle with. The weak ones try to renegotiate the test back toward their demo. That reaction is itself due-diligence data, free of charge.
One more thing the sequence buys you, and it is the deepest one: it keeps the pilot pointed at the right object. This chapter is titled "The Pilot That Survives Contact," and what the pilot tests is the thing you built in Chapters 3.1 and 3.2: the redesigned process, with its rewritten SOP (standard operating procedure), its schemas, its human gate, its verification layers, its counters. McKinsey's finding runs underneath this whole level: the roughly 6 percent of high performers are about three times more likely to have fundamentally redesigned workflows, and workflow redesign is among the strongest drivers of impact anyone can measure. The pilot exists to test a redesign, with a tool inside it, not to test a tool with your process as the demo environment. Design before procurement is what keeps that true.
The Worked Example and the Mirror
Time to render the artifact whole. Everything that follows is hypothetical, the running illustration of this program, not research data.
The invoice-exception Pilot Design Brief, one page
| Section | Content |
|---|---|
| Hypothesis | AI-assisted categorization of the three structured exception types will cut median cycle time from 3.2 days to under 2.0 days without raising the rework rate, because the queue wait, not reading speed, is the constraint. |
| Scope | Routine lane only (the 78 percent: missing PO, price mismatch, quantity mismatch), one shared-services site, six weeks spanning the August month-end. Expected volume: about 268 items per week, roughly 1,600 observations. OUT, with reasons: free-text lane (would blur the hypothesis; different mechanism), second site (attribution), policy-exception approvals (human-only per the gate design). |
| Primary criterion | Median cycle time under 2.0 days on in-scope items; counter: event-log released-minus-received, identical definition to baseline pack; window: weeks two through six (week one labeled ramp). |
| Guardrails | Rework rate at or below 12 percent (baseline 11; counter: reopened rows plus ERP reversals); gate queue latency under 4 hours (counter: gate timestamps); clerk overtime flat versus baseline (counter: time system). |
| Kill condition | Verified accuracy below 85 percent at week six (counter: layer-two gate catches plus layer-three 100-document weekly audit sample; tool self-scores excluded), or any critical escape (wrong-vendor payment, ledger misposting) at any time. |
| Evidence plan | Independent assessor runs the weekly tally ritual; one-page weekly with criteria printed on top; committee checkpoints week three (health check) and week six (full read); decision meeting booked September 30, criteria hierarchy attached to the invite. |
One page, and every line traceable: the hypothesis to the bottleneck analysis of Chapter 3.1, the counters to the Measurement Plan of the last lesson, the kill floor to the charter signed in Level 2. Nothing in it required a vendor's opinion, and nothing in it can now be quietly rewritten by one.
Two vendors, one brief
The brief goes out to two shortlisted vendors from the Level 1 due-diligence pass, with the schema and twelve acceptance samples attached. Vendor A returns structured output against the schema on ten of the twelve samples, flags the two multi-currency documents honestly as a known weakness, and proposes a handling rule for them. Vendor B requests a live demo instead, shows a dazzling forty minutes on its own dataset, and, when finally pressed to run the acceptance samples, produces confident, fluent extractions in which the skewed-scan invoice has an invented purchase-order number and the ambiguous line items have been silently merged. In a demo-first world, Vendor B wins on charisma; the room never sees the samples because there is no scope forcing the test. In a brief-first world, the collapse happens in procurement, before a contract exists, at a cost of zero. That is the Level 2 skepticism arc (verify every AI-touched claim; the confident fabrication is the failure mode) paying its dividend in purchasing, where it is cheapest.
The mirror: the pilot that could not fail, and therefore could not end
Now the same movie with the dates reversed, assembled from the standard patterns with illustrative figures. A marketing-operations team at a mid-size firm sees a content-generation tool demoed at a conference: on-brand blog drafts in ninety seconds, the room audibly gasps. The pilot is designed the following week, which is to say the demo is transcribed into a project plan. Scope: "content acceleration," meaning whatever the tool did well on stage. Metric: outputs produced per week, straight from the vendor's dashboard, and if that phrase raises an old alarm, it should: it is an activity metric, the usage trap from Level 1, drafts generated rather than anything a P&L can feel. No hypothesis, no baseline of current content cost or revenue effect, no guardrails, no kill floor, no decision date.
The pilot runs a quarter and, by its own lights, triumphs: draft output up four times. Also true: no measured effect on pipeline, traffic, conversion, or time-to-publish (the editing bottleneck absorbed the flood; senior reviewers now spend their week triaging machine drafts), and the license costs $54,000 a year. At the end of the quarter, the team meets to decide, and discovers there is nothing decidable: no number was promised, so no number can have failed. This is the shape to fear, and it is worth being precise about why. It is not a disaster; disasters get killed. It is something worse, an unfalsifiable success: a pilot that produces impressive activity, cannot be proven valuable, cannot be proven worthless, and therefore renews by default, every year, as a budget annuity. MIT's 95 percent is substantially made of exactly this material, and S&P Global's finding that 42 percent of companies scrapped most of their AI initiatives in 2025 is what it looks like when the annuities finally get audited in bulk. Notice, finally, who designed this pilot. No one, and everyone: tool gravity ran the entire arc, from a conference stage to a renewal line item, and nobody caught it because everybody was busy, and the tool was, in its way, working. Evidence or it didn't happen; here, nothing could have happened, by design, except the design was nobody's.
The invoice-exception pilot cannot end that way, and now you can say precisely why: its hypothesis can lose, its guardrails price every shortcut, its counters were wired before commitment, and its ending is a calendar invite with the criteria attached. The design is done. What remains is to run it: the weekly cadence, the logs, the drift-watch that keeps six weeks of contact with reality from bending the experiment. That is the next lesson.
What to Do Monday Morning
The Pilot Design Brief is one page and roughly two days of work, all of it before any vendor call. Here is the sequence.
- Write the hypothesis sentence and read it aloud for falsifiability. One sentence: we believe [intervention] on [bounded segment] will move [number] from [baseline] to [target] without worsening [guardrail], because [mechanism]. Then ask the killer question: what result would prove this false? If nothing would, rewrite until something would.
- Draw the scope with its OUT list and volume arithmetic. One process, one lane, one site. Write each exclusion with its experimental reason (control, not confession). Multiply weekly volume by candidate duration and check the observation count against the movement you need to detect; if the arithmetic fails, lengthen the window or change the evidence standard, in writing. Confirm the window spans the process's natural stress cycle.
- Wire every criterion to a named counter. For the primary, each guardrail, and the kill condition: counter name from the Measurement Plan, threshold, window, verification layer. Any criterion that cannot name its counter goes back to the last lesson's charter-mapping exercise until it can.
- Book the decision meeting with the criteria in the invite. Scale, iterate, or kill, on a named date after the pilot window, on the steering committee's calendar this week, with the primary, guardrails, and kill floor pasted into the invitation body. Add the week-three and week-six checkpoints while the calendar is open.
- Only then, schedule vendor conversations, with the brief attached. Send shortlisted vendors the scope, the schema, the acceptance samples, and the criteria, and ask each to run your samples against your schema as the demo. Watch which vendors engage with the test and which try to renegotiate it back toward their slideware; write down what you observe, because that reaction belongs in the due-diligence file.
Key Takeaways
- Sequence design before procurement without exception: scope, baseline, and success criteria are decided before any vendor is in the room, because the pilots in MIT's 5 percent test their own hypothesis while the 95 percent test the vendor's story.
- Name and manage tool gravity: the moment a tool enters the conversation, its demo becomes the scope, its dashboard becomes the metrics, and its strengths become the success criteria, through rational behavior on both sides and to the total destruction of evidence.
- Build the Pilot Design Brief before the first vendor call: hypothesis, scope with justified exclusions, criteria wired to counters, and evidence plan, the engineering drawings behind the charter you signed in Level 2.
- Write scope as a falsifiable, mechanism-stated, segment-bounded hypothesis, then bound it by experimental logic: one process, one lane, one site (a multi-site pilot is a rollout in a lab coat), exclusions as controls, volume arithmetic that reaches a verdict in the window, and a duration spanning the process's natural cycle.
- Wire every criterion to a named counter, threshold, window, and verification layer before commitment, because unwired criteria are renegotiable and wired ones are checkable: the difference between a contract and a vibe.
- Structure criteria as a hierarchy: one primary (the hypothesis's number), two or three guardrails that catch a target hit by cannibalizing something unmeasured, and a kill condition signed while hypothetical, because Gartner's 40 percent agentic-cancellation wave is what un-pre-committed endings look like: late, expensive, acrimonious.
- Book the decision meeting before the pilot starts, with the criteria attached to the invite: a pilot that ends at a known time with something decided cannot become a zombie, and an unfalsifiable success renewed by default is the budget annuity the 95 percent is made of.
- Turn procurement into a structured test: vendors respond to your brief, run your acceptance samples against your schema, and get scored by your verification protocol, so the demo that would have designed your pilot instead has to pass it.
Skill.re