←
AI Readiness & Process Transformation
Visionary · M13 · lesson 13 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Running Experiments Without Burning Trust or Budget

15 min

The idea arrives in the wrong meeting. A contracts manager, two years into a governed AI program, says she thinks a grounded assistant could flag the clause deviations her team keeps missing, and the room does what a well-run program has trained it to do: it asks for the business case. What is the baseline? What is the expected hours saved? What is the payback period? She has none of it, because nobody in the company has ever tried this on these contracts, and she genuinely does not know whether it will work at all. So she does one of two things. She drops it, and the company stays excellent at improving what it already understands. Or she invents the numbers, and eleven months later there is a pilot nobody can kill, defended by a spreadsheet that was fiction on the day it was written. Both outcomes are failures of instrument, not of judgment. This lesson hands you the missing instrument.

The Work Your Discipline Cannot Handle

Everything you have built across this program is an apparatus for estimating value before you spend money on it. Baselines, so you can prove movement. Business cases, so a finance partner can sign the arithmetic. Pre-committed success and kill criteria, so nobody argues about the definition afterwards. Stage gates, so a weak candidate dies cheaply. A work-in-progress limit, so the program finishes things instead of starting them. That apparatus is why your portfolio no longer looks like the 95 percent of enterprise generative AI pilots MIT found producing no measurable profit-and-loss return.

It has one blind spot, and it is structural rather than accidental. Every one of those controls assumes the value of the work can be estimated before the work begins. Baselines assume you know which metric will move. Business cases assume you can bound the benefit. Gates assume you know what evidence to look for. Apply that apparatus to a candidate whose value genuinely cannot be estimated in advance and it does not fail loudly. It fails politely: it asks a reasonable question the proposer cannot answer, and either rejects a good idea or admits a fabricated estimate into a system built on honest ones.

The previous lesson's use-case engine is built for candidates that can answer those questions: a stated mechanism, a known readiness profile, a scoreable value estimate. That is most of the work, and it should be. But a serious portfolio also contains a category the engine cannot process: the genuinely uncertain idea, the capability class nobody has tried on this problem, the hunch a credible champion cannot justify in a spreadsheet because the thing she is uncertain about is exactly what the spreadsheet would have to assume.

The two bad answers organizations give

Faced with that category, organizations fail in one of two directions, and both are expensive.

Direction one: refusal. The program applies its discipline uniformly, demands a business case for everything, and quietly stops taking uncertain bets. This looks like maturity for about eighteen months. The portfolio delivers, the numbers hold, the governance is clean. What has actually happened is that the organization has become superb at optimizing processes it already understands and structurally blind to capability changes it has not met. When a capability class arrives that reorganizes a piece of your industry, the refusing organization does not learn about it late; it learns from a competitor's press release. Its discipline did not fail. It worked perfectly on the wrong scope.

Direction two: smuggling. The champion, who believes in the idea, reverse-engineers the business case: a plausible baseline, a vendor benchmark applied as if it were a forecast, a document that satisfies the gate. The damage lands in two places, and both are worse than the money. First, it corrupts the business-case discipline everywhere: once one case is known to have been fiction, every case is read as negotiable. Second, and this is the mechanism that produces the zombie, an experiment dressed as a pilot cannot be killed on its own terms. A pilot is killed by failing to hit its projected value, so when exploratory work returns an ambiguous result, the honest report ("we learned the model cannot do this reliably on our documents") reads as a missed target rather than a finding. Nobody wants to write that. So it goes under review, renews its licence, and joins the line items that make the next AI proposal harder to fund.

Gartner's forecast that over 40 percent of agentic AI projects will be canceled by the end of 2027 names the two causes precisely: escalating costs and unclear value. That is a description of exploratory work with no cap and no honest framing. Escalating cost is what happens with no ceiling; unclear value is what happens when work was never designed to produce an answer. Neither is prevented by more business cases.

Pilot and Experiment: Two Instruments, Not One

The answer is not to loosen the discipline. It is to build a second instrument with its own rules, and to be pedantic about which one you are holding. Most organizations use these two words interchangeably, and the interchange is the problem.

A pilot tests whether a designed solution delivers expected value under production conditions. You already believe the thing works; what you do not know is whether it works here, at your volumes, with your people, inside your controls. A pilot is justified by a business case, because the case states the value you expect and the pilot tests that expectation. It carries a charter, a baseline, an adoption plan, pre-committed criteria, and a named owner accountable for a number moving.

An experiment tests whether something is true. You do not know whether the capability works on this problem, whether humans would accept its output, or whether the workflow can absorb it at all. An experiment is justified not by projected value but by the value of the answer: what the organization gains by knowing, cheaply and soon, whether this path is open or closed. It carries no business case, because the case would forecast the quantity the experiment exists to discover.

DimensionPilotExperiment
Question it answersDoes this deliver the value we projected?Is this true?
Justified byA business caseThe value of the answer
ProducesValue (and evidence of value)An answer (and a documented finding)
Governing instrumentCharter, baseline, stage gateCost cap and time box
Design goalAs representative as possibleAs small and fast as possible
Needs an adoption planYes, from day oneNo
Success looks likeThe projected number movesA confident answer, including no
Failure looks likeValue did not materializeCap consumed, still no answer

Four practical consequences follow, and each will feel wrong the first time you enforce it.

An experiment needs no return-on-investment projection, and asking for one is a category error. Return on investment (ROI) is a ratio of expected benefit to expected cost, and in an experiment the expected benefit is knowledge whose worth depends entirely on the answer you have not received yet. Demanding a projection does not produce rigor; it produces fiction, reliably, because the only way to satisfy the demand is to invent the number. Governance that forces well-intentioned people to invent numbers has stopped being governance.

An experiment needs no adoption plan. Change management, training, super-users, comms: that is the machinery of making a proven capability stick. Attaching it to an experiment triples the cost and, worse, creates a constituency with a stake in the answer being yes. Once thirty people have been trained on something, "we learned it does not work" becomes a political statement rather than a finding.

An experiment should be as small and fast as possible, not as representative as possible. This inverts the pilot instinct completely. A pilot wants production-like conditions because it tests production performance. An experiment wants the cheapest configuration that can still change your mind. If a two-day test on forty redacted samples would move your belief, a three-month integration testing the same proposition is not more rigorous. It is more expensive for the same information.

An experiment's success is a confident answer, and no counts. This is the sentence that protects the entire system, and it has a timing requirement: the leader must say it out loud, with the sponsor and the finance partner in the room, before the first experiment runs. Said afterwards, on the day a probe returns a no, it sounds exactly like what it is not: an excuse. Said in advance, it becomes the standing definition against which every result is read. There is no way to retrofit it.

An experiment that returns a clear no, on budget and on time, is a success. Say that before the first one runs, not after the first one fails.

The Artifact: The Experiment Brief

Here is the lesson's named artifact, and its whole point is that it fits on one page and takes about forty minutes to write. If your version runs to four pages you have rebuilt the business case under a different name. Five fields, in order.

Field one: the question, stated so an answer is possible

Most exploratory ideas arrive as a topic wearing a question mark. "Can AI help with contract review?" is a topic. No result could ever settle it, which is precisely why it survives budget cycles indefinitely. Compare: "Can a grounded assistant identify our five highest-risk clause deviations, on our contracts, at a rate our lawyers consider useful?" That names the capability, the scope, the population, and the judge. Run it and you get an answer.

The sharpening technique is one repeated question: what would I see? Ask it of your draft until the draft names something observable. "It would find risky clauses." Which ones? "The five deviation types on our risk register." Find them where? "On a sample of our own executed contracts, not vendor demo documents." And how would I know the flagging is good enough? "Our lawyers would say they would act on it." Then say that, and name the lawyers.

Three rounds and a topic has become a testable proposition. The technique also works as a filter: some ideas cannot survive the sharpening at all. When a champion cannot say what she would see, the idea is usually a wish about a technology rather than a hypothesis about a problem, and you have saved the whole cap by spending forty minutes.

Field two: the cheapest test that could answer it

Once the question is sharp, the reflex of every capable team is to build the real thing: integrate the systems, ingest the corpus, stand up the pipeline, then evaluate. That reflex is the single biggest waste in enterprise experimentation, and it is expensive precisely because it feels responsible. Forty contracts and two lawyers' afternoons beat a three-month integration, because the integration answers a question nobody asked (can we build it?) at ten times the cost of the one that matters (would it help if we did?).

The brief asks for the minimum informative test: the smallest, fastest configuration that could still change what you believe. Three formats cost almost nothing and answer a surprising share of enterprise questions.

  • The manual walkthrough. An off-the-shelf tool, a redacted sample, a human doing the steps by hand and recording what happens. No integration, no build, no environment. Cost measured in afternoons. It answers capability questions ("does the model handle our language at all?") faster than any pipeline.
  • The wizard-of-oz test. A human silently simulates the capability so you test the workflow rather than the technology. An analyst produces the flags manually and they reach reviewers exactly as an automated system would deliver them. You learn whether the workflow can absorb the output, who acts on it, and where it stalls. That question is independent of whether the model can do it, and it is often the one that decides the idea.
  • The paper test. Show people the output format, on real examples, and ask what they would do with it. No system exists. You are testing the last mile: whether the artifact you propose to produce would actually change a decision.

Note the pattern in the last two. Very often the workflow question, not the model question, kills the idea, and the workflow question is answerable in a day. This matches MIT's autopsy of stalled pilots: failures clustered in missing workflow integration and adoption without transformation, not in model capability. If a workflow fact will eventually kill your idea, discovering it in a two-day paper test rather than in month seven of an integration is the entire return on this artifact.

Field three: the cap, in both dimensions

Write two numbers: money and elapsed weeks. Both are required, and the elapsed cap matters more. The direct cost of a drifting experiment is usually modest against an enterprise portfolio. The real cost is attention and credibility: the standing agenda item that never resolves, the leader saying "still looking into it" for the fifth month, the slow conversion of an interesting question into a running joke. Money you can absorb. A program that visibly cannot finish things loses the authority it needs.

Illustrative tiers, to adapt to your own scale:

TierMoney capElapsed capTypical shape
Level-one probeUnder 10,0003 weeksPaper test, manual walkthrough, wizard-of-oz on a small sample
Level-two experimentUnder 40,0008 weeksSmall sandboxed build, structured evaluation on a real corpus, several reviewers
Anything largerNot an experimentNot an experimentThis is a pilot: it needs a charter, a baseline, and a business case

That third row is not a formality. It is what stops the instrument being abused in the other direction, where "experiment" becomes the label people use to escape governance on a large piece of work. Above the ceiling, the old rules apply, without exception.

The cap is a promise made in advance and honored at expiry, and the hard case is not the one that fails early. The hard case is the one that is tantalizing at week eight: signal is emerging, the team is close, one more sprint would surely settle it. "We were so close" is the most expensive sentence in exploration. It turns a 40,000 experiment into a 300,000 open-ended effort one reasonable extension at a time, each individually defensible and collectively catastrophic.

The correct response is neither to extend nor to abandon. It is to write a new brief with a new cap. That is not bureaucratic theater; it is the difference between justified and unjustified spend. A re-briefed experiment has been re-justified: someone restated the question in light of what was learned, chose a test, set a cap, and put their name on it in front of whoever governs the fund. An extended experiment has merely continued, carried by sunk cost and the enthusiasm of the people running it.

Field four: what each outcome would trigger

Before the test runs, write down what each of the three possible results causes. This takes ten minutes and prevents the most common exploratory pathology.

  • Clear yes. The idea routes to the use-case engine's shaping stage with everything learned attached: the sample, the reviewer feedback, the observed rates, the workflow surprises. It does not become a project by acclamation. It becomes a candidate with unusually good evidence, entering the same intake as everything else, where it will do well precisely because the uncertainty has been retired.
  • Clear no. Documented and published. One paragraph: what we asked, what we did, what we found, what it cost. A deliverable, not an apology, and the publication step converts a dead end into an organizational asset that stops the same idea being re-proposed by three other functions over the next two years.
  • Ambiguous. Pre-commit to exactly one of two paths: one re-briefed experiment with a fresh question and a fresh cap, or a decision to stop. Not two more. Not "keep an eye on it."

Pre-committing the ambiguous case is what earns this field its place. Ambiguity is the natural output of exploratory work and the perfect fuel for an indefinite sequence of "just one more test," each small, each reasonable, the sequence unbounded. Deciding the rule while nobody is invested costs nothing. Deciding it afterwards, among people who have just spent six weeks on something interesting, is a negotiation you will lose.

Field five: the disclosure plan

The fifth field is the one people skip, and it is the one the lesson's title is about. Write down who hears about this experiment, when, and in what terms. The rule is short: experiments are visible while running and reported when finished, including the failures, in the program's regular channels. Not a separate innovation newsletter nobody reads: the same quarterly note, the same steering pack, the same page as delivered work.

The dynamic underneath is counterintuitive. An organization that hears about experiments only when they succeed learns a false model of how this work goes: that AI works, that the program picks winners, that every announced effort arrives. That model is comfortable and it is a trap, because it makes every eventual failure a scandal. When the first public failure lands in an organization with an unbroken record of announced successes, the reaction is not "one of these did not pan out." It is "what else have you been telling us?"

An organization that hears "we ran four probes this quarter, one advanced to shaping, three answered no, and here is what each taught us" develops an accurate model instead. It learns the real hit rate. It learns that a no is a normal Tuesday. It stops being surprised, because you have been calibrating its expectations continuously and for free. Publishing failures is not humility. It is expectation management with a compounding return: each published no makes the next one cheaper in credibility, until the organization treats exploration the way it treats a research and development line, which is what it is.

The Trust Economics: What Actually Burns Trust

Leaders protect exploration budgets badly because they misidentify the threat. Believing failure is what burns trust, they avoid failure, which means either not exploring or over-engineering experiments until they cannot fail visibly. Both responses make it worse.

Failure is not what burns trust. Surprise burns trust, and unbounded failure burns trust. Run the two cases side by side.

A capped experiment answers no in six weeks. Cost: its budget, say 18,000, and a small team's part-time attention. The sponsor knew it was running because the disclosure plan said so, and the result appears in the quarterly note alongside everything else. What the organization concludes: the program tests things, finishes them, and tells the truth. The 18,000 bought an answer and something more durable, a demonstration that this program's numbers can be believed when they are unflattering.

An uncapped exploration consumes two quarters and ends ambiguously. Cost: its budget, plus the team's time, plus the sponsor's patience, plus the organization's willingness to fund the next one. That last item is the expensive one and it appears in no ledger. It is why the next honest experiment proposal dies in a corridor.

That asymmetry is why the leader's line to the board works, and it is a line you can steal verbatim: "We will run roughly six experiments a year. Most will answer no. The total is capped at 300,000, which is under two percent of the program envelope. And you will hear about all of them, including the ones that fail." That sentence converts an unbounded, faintly alarming activity into a governed one with a known ceiling and reporting rhythm. It pre-announces the failure rate, so failures confirm the model rather than violating it. And it is easy to keep, because every promise in it is under your control. Boards do not object to bounded, disclosed, capped exploration. They object to discovering exploration.

Who runs them, and the capacity rule

Experiments need a small standing capacity, not ad-hoc heroics. The natural home is the transformer group between deliveries: people who know the process landscape, who have run structured evaluations, and who will not mistake a demo for evidence. Two to four people with a few days a month each will sustain six briefs a year at these caps.

One rule makes this work and its absence breaks it: experiment work does not compete with committed delivery. Exploration capacity is allocated explicitly, in the same plan that allocates delivery capacity. An experiment that quietly borrows two days a week from a delivery lead is not free. It is an undeclared delivery slip, and it is exactly how a work-in-progress (WIP) limit gets violated invisibly: nothing on the board changed while the capacity behind the board did. When the delivery slips a month later, the cause gets attributed to the delivery, and the program draws the wrong lesson twice. So name the people, name the days, put them in the capacity plan, and treat a raid on delivery capacity as the same class of error as skipping a gate.

Worked Example: Two Experiment Portfolios

Norvik Group is the 2,400-person industrial distributor this level has followed: a scrap year, then a governed rebuild, now in year two of the funded program. All figures below are illustrative and rounded, useful mainly as a shape.

The log that worked

Norvik's explore fund is capped at 300,000. In year two the program wrote six experiment briefs and spent 34,000 of it. The underspend was reported deliberately, with an explanation: the cheapest-informative-test discipline meant most questions were answered by level-one probes rather than level-two builds, and the fund is a ceiling, not a target. Six outcomes:

  • Three clear nos. The most instructive was the contract-clause probe, which died on a workflow finding rather than a model finding. A two-day paper test costing about 1,200 in analyst time showed lawyers a mocked-up flag list on real deviations. Every reviewer said the same thing: they would not act on a flag without the supporting clause text and the reasoning displayed alongside it, and building that display was a different piece of work from the one proposed. The unbuilt integration was estimated at around 60,000. Two days and 1,200 retired the question.
  • One ambiguous, re-briefed once, then stopped. A demand-forecasting probe returned encouraging results on two product families and noise on the other five. Per the pre-committed rule it got exactly one re-brief, with a sharper question and a 15,000 cap. The re-brief answered no. The file closed with a paragraph, and nobody fought about a third attempt, because the rule was written before anyone cared about the answer.
  • Two clear yeses, routed to shaping. Both entered intake with unusually strong evidence attached: real samples, real reviewer feedback, an observed rate rather than a vendor benchmark. One became the highest-value delivery of the following quarter, and it scored well at intake for a reason worth noticing: the experiment had already retired the uncertainty that would otherwise have made its estimate a guess.

The quarterly note reported all six, one line each, in the same document as the delivered work. No separate innovation section, no celebratory framing on the yeses, no defensive framing on the nos. By the third consecutive quarter of published nos, the board did something nobody had asked for: it raised the explore cap. The stated reason was not that the experiments had succeeded. It was that the discipline was visible, and a capped activity that reports its own failures on schedule is one of the few things in a transformation portfolio a board can actually verify.

The failure story: the skunkworks that became a scandal

A division of a comparable company took the other path, starting with a sentence that sounds like protection: "let it prove itself before we make noise." A divisional leader funded an exploratory AI effort quietly out of operating budget. No brief. No cap, in money or weeks. No disclosure plan, by design, because visibility was understood as the risk rather than the control.

It ran eleven months. Two engineers and a business analyst, part-time at first and then not, plus vendor spend on an evaluation environment and a data pipeline: roughly 400,000 in salary and external cost by the time anyone added it up. It produced a genuinely impressive demo and no deployable outcome, because the workflow question nobody asked in month one proved unanswerable in month eleven without a system integration that had never been scoped or budgeted.

It surfaced during a routine budget review, the way these things always surface: someone asked what a cost line was. The reaction was not proportionate to the money, which was modest against that company's portfolio. It was proportionate to the concealment and the open-endedness. The chief financial officer's questions were not "did it work?" but "how long has this been running, who approved it, and what would have stopped it?" With no cap and no brief, the honest answer to the last question was "nothing."

The new rule arrived within a month: all AI exploration in that division would go through the full business-case process. Note the shape of that outcome. The rule is not unreasonable on its face, and it ended experimentation there for two years, because business cases cannot be written for genuinely uncertain work. The uncertain work went back where it came from: into the shadows, unfunded and unrecorded, waiting to become the next budget-review surprise.

The two stories differ by about 366,000 in spend, but that is not the lesson. They differ by two properties, secrecy and open-endedness, and those two properties convert an honest failure into a governance event. A published, capped no is a normal quarter. A concealed, uncapped ambiguity is a scandal at a fraction of the money.

Why the caution is rational

Concede openly that your finance partner's caution is rational. Only about 5 percent of MIT's custom enterprise tools crossed from pilot into production. S&P Global found 42 percent of companies scrapping most of their AI initiatives in 2025, up from 17 percent a year earlier. McKinsey found 88 percent of organizations using AI regularly while only about 39 percent could attribute any earnings impact to it. But every one of those numbers describes unbounded, unmeasured, undisclosed work. None describes a 1,200 paper test that answered a question in two days and appeared in the quarterly note. That distinction is your argument, and the Experiment Brief makes it checkable rather than rhetorical.

What to Do Monday Morning

You almost certainly have a candidate already: the idea you have not been able to justify as a pilot, the one that keeps resurfacing and keeps failing the business-case test.

  1. Write one Experiment Brief for that idea. One page, five fields, forty minutes. If it runs longer than a page, you are writing a business case again.
  2. Sharpen the question until it names an observable. Ask "what would I see?" repeatedly until the answer names a population, a capability, and a judge. If it will not sharpen, close the file and consider the forty minutes well spent.
  3. Design the cheapest informative test, not the realistic one. Ask explicitly: could a paper test, a manual walkthrough, or a wizard-of-oz simulation change my mind? If yes, that is the test. Resist the build.
  4. Set both caps and say them out loud to your sponsor. Money and elapsed weeks. Add the sentence that protects the system: a clear no, on budget, is a success. Say it now, while nothing has failed yet.
  5. Pre-commit the ambiguous outcome in writing: one re-brief with a fresh cap, or stop. Decide it while you do not care about the answer.
  6. Publish last quarter's experiments, including the ones that answered no. One line each, in the channel where your delivered work already appears. If you have never published a failure, this is the highest-leverage paragraph you write this month.
  7. Put exploration capacity in the capacity plan by name, and declare the rule that experiment work never borrows from committed delivery.

Key Takeaways

  • Recognize the blind spot in your own discipline: baselines, business cases, and gates all assume value can be estimated before work begins, which makes them the wrong instrument for genuinely uncertain ideas.
  • Reject both bad answers, refusing uncertain work entirely (which produces an organization blind to capability change) and smuggling it through as a pilot with an invented business case (which corrupts the case discipline and produces zombies).
  • Separate the two instruments rigorously: a pilot tests whether a designed solution delivers projected value and is justified by a business case, while an experiment tests whether something is true and is justified by the value of the answer.
  • Apply the four consequences: an experiment needs no ROI projection (demanding one produces fiction), needs no adoption plan, should be as small and fast as possible rather than as representative as possible, and succeeds when it returns a confident answer including no.
  • Write the Experiment Brief on one page: the sharpened question, the cheapest informative test, the cap in money and elapsed weeks, the trigger for each outcome, and the disclosure plan.
  • Reach for the tests that cost almost nothing, the manual walkthrough, the wizard-of-oz simulation, and the paper test, because the workflow question rather than the model question usually kills the idea and it is answerable in a day.
  • Honor the cap at expiry and re-brief rather than extend, because "we were so close" converts a capped experiment into an open-ended effort one defensible extension at a time, and a re-briefed experiment has been re-justified while an extended one has not.
  • Publish the failures in the program's regular channels, because what burns trust is surprise and unbounded loss rather than failure itself, and a board told "roughly six a year, most will answer no, capped at this number, and you will hear about all of them" funds exploration instead of fearing it.

The engine turns understood opportunities into delivered value and the Experiment Brief turns uncertain ones into answers, which together produce a steady supply of verified wins inside one function. That creates the next problem, the hardest one in enterprise transformation: taking a win that works in finance or in service operations and making it an enterprise capability without breaking the thing that made it work.