←
AI Readiness & Process Transformation
Proficient · M13 · lesson 13 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Measuring Against Baseline Honestly

15 min

The pilot's six weeks ended on a Friday. On Monday morning there is already a slide in the transformer's inbox, drafted over the weekend by the sponsor's chief of staff, and it is beautiful: a single fat arrow pointing down, and the words "Invoice-exception cycle time improved 44%." The transformer checks the arithmetic. The 44 percent is real: the baseline pack's median of 3.2 days against the pilot's 1.8. Every digit is true. And the slide is still wrong three different ways at once, because the 3.2 is the wrong baseline, the 1.8 is carrying an unexamined mix shift, and cycle time alone is not what the pilot cost or saved. "Cycle time improved 44%" can be true, false, and meaningless at the same time, depending on windows, mix, and what got controlled. Between the pilot's logs and the baseline's before-photo stands the most abused calculation in enterprise AI: the delta. This lesson is about computing it the way a hostile analyst would, and reporting it with its caveats attached, because a delta that arrives naked gets dressed by whoever presents it next.

The Most Abused Calculation in Enterprise AI

Take stock of where the running example stands. The invoice-exception pilot ran its six weeks on the cadence you built in the previous lesson: the weekly one-pager, the conservation line, the drift watch that caught and corrected the week-five event. The event log is full. Every state change of every exception, timestamped, attributed, queryable. On the other side sits the Level 2 baseline pack: the sworn before-photo, with its measurement notes and its segment tables. Two honest datasets. And now someone has to subtract one from the other, and this is the step where more pilot credibility dies than in any other, because subtraction looks like arithmetic and is actually a chain of judgment calls, every one of which has a flattering option.

Which baseline window? Which pilot weeks count? Aggregate or by segment? Gross time saved or net of the verification the redesign added? Each question has a defensible answer and a convenient one, and the convenient answers compound. Choose the flattering option four times in a row, each individually arguable, and a real 21 percent improvement presents itself as 44, and the first analyst who reverses one choice unravels all four, in public, with your name on the slide.

So set the standard now, before touching a number: compute every delta the way a hostile analyst would compute it, and attach the caveats yourself, before anyone else gets custody of the number. This is not humility theater. Honest measurement is not modesty; it is precision about what the evidence does and does not show. If the pilot genuinely worked, honest analysis proves it in a way no one can take apart. If it did not, you need to be the first to know, not the last. Either way, precision is the only version of the story that survives contact with the one person every steering committee eventually contains: the skeptic with a calculator and an afternoon.

The stakes have been on the wall since Level 1. MIT's finding that 95 percent of enterprise generative AI pilots showed no measurable return convicts, above all, measurement failure: organizations that could not prove value because they never built the instruments to find it. You built the instruments. McKinsey's State of AI survey shows the other half of the trap: 88 percent of organizations use AI, yet only about 39 percent can attribute any earnings impact to it. Attribution is exactly what this lesson manufactures. The program's oldest slogan, "baseline or it didn't happen," completes its arc here: the baseline was the first half, identical live counters were the second, and the honest delta with caveats attached is the finished sentence.

A delta that arrives naked gets dressed by whoever presents it next. Attach the caveats yourself, or lose custody of your own number.

The artifact this lesson builds is the Delta Table: one page, one row per metric, seven columns. The baseline value from the correctly chosen window. The pilot value with its controls applied. The delta. The controls that were applied, named. The caveats that survive after the controls, stated in plain words. And a confidence word (clear, probable, or suggestive) assigned by rules agreed before anyone saw the results. The Delta Table is the evidence base the next lesson's decision memo will stand on, which means every shortcut taken here becomes a crack in the verdict there. We will build it through five disciplines, slowly, each one worked against the invoice-exception numbers. All figures ahead are hypothetical, an illustration carried forward from this program's running example.

Disciplines One and Two: Like Windows, Then Segments

Discipline one: compare like windows

The first judgment call is the baseline window, and the first temptation is the number everyone already knows: the baseline pack's headline median of 3.2 days, computed across a full year of exceptions. It feels authoritative precisely because it is familiar. It is also the wrong comparator, for a structural reason: the pilot's six weeks were not an average six weeks. They spanned one month-end close, the week when finance floods the queue, vendors chase payment, and cycle times spike. The annual average smooths that spike across fifty-two weeks, month-end chaos diluted by calm mid-month stretches. The pilot had no such dilution; it had to survive its month-end at full strength. Comparing a window that contains a storm against an average that has smoothed its storms away is not a comparison; it is two different weather reports.

The like-window rule: reconstruct, from the baseline pack's underlying data, a comparison window with the same structure as the pilot window. Six consecutive weeks, spanning exactly one month-end, from the baseline period. Run the same queries on it that the event log runs on the pilot. In the invoice-exception case the difference is not subtle: the month-end-inclusive baseline median comes out at 3.4 days, not the annual 3.2, and the like-window rework rate comes out at 12 percent against the annual 11, because month-end reopens more items. Those are the comparators the Delta Table will carry.

Notice what just happened, because it teaches the rule's real nature. In this pilot, the correction moved the baseline up, which makes the pilot's win look slightly larger. Do not celebrate that; do not even enjoy it. The rule was chosen for its structure, before anyone knew which direction it would cut, and it binds both ways. In most pilots it cuts the other way: pilots get scheduled into calm periods, compared against spike-inclusive annual averages, and the like-window correction shrinks the claimed win. When that happens to you, volunteer the shrinkage, loudly and first. Shrinkage you volunteer is credibility you bank: the analyst who watches you make your own number smaller stops auditing you and starts believing you. A comparator chosen after seeing which one flatters is not a comparator; it is a costume.

Discipline two: segment before you average

The second discipline attacks the aggregate itself. The invoice-exception process handles three exception types with very different difficulty profiles: missing purchase order (PO, the reference document that authorizes a spend), price mismatch, and quantity mismatch. The baseline pack measured each segment separately, and the instrumentation kept that segmentation live, precisely so that this moment could happen: the honest default comparison is within segment, each type against its own like-window baseline, and only then rolled up.

Why? Because an aggregate delta inherits every shift in the mix of work, and the mix never holds still. Run the pilot's actual numbers. Price mismatches, the easiest type, made up 41 percent of baseline volume in the like window but 45 percent of pilot volume: a 4 percentage-point drift toward the easy lane, caused by nothing more sinister than a vendor's seasonal ordering pattern. Easy items finish faster, so a queue with 4 percent more of them posts a better median even if nobody improved anything. How much better? Compute it openly: reweight the baseline's segment medians by the pilot's mix, and the baseline "improves" by about 0.2 days on its own. That 0.2 days is fake improvement, and it was sitting inside the pilot's raw median of 1.6 days, quietly inflating the win.

The mix adjustment removes it, and the plain-words version of the technique is worth memorizing, because you will explain it to executives who will never sit through the arithmetic: "we asked what the baseline process would have scored if it had been handed the pilot's mix of work, and we compared against that." Reweighted, the honest pilot figure is 1.8 days against the like-window 3.4. Yes, the adjustment made the pilot's number worse, from 1.6 to 1.8, and yes, you show the work: the segment table, the weights, the 0.2 removed, all visible in the Delta Table's controls column. An adjustment performed in the open is a control. The same adjustment discovered later by someone else is an accusation.

Within-segment, the story holds everywhere, which is what a real improvement looks like: missing-PO exceptions ran 2.6 days against their own baseline of 4.4, price mismatches 1.2 against 2.7, quantity mismatches 1.9 against 3.3. When every segment improves against its own baseline, the aggregate is trustworthy. When the aggregate improves and the segments do not, you are looking at a mix shift wearing a costume, and you have just learned to strip it.

Disciplines Three, Four, and Five: The Ramp, the Whole System, and What You Cannot Claim

Discipline three: separate the ramp from the verdict

The pilot's first two weeks were not the redesigned process at steady state; they were a team learning a new process in public. Week one carried the enthusiasm bump the instrumentation lesson warned about, plus the calibration change to the extraction prompts made in week two. The charter anticipated exactly this: it pre-declared weeks one and two as ramp, with the primary measurement window running from week three. The analysis now simply honors that declaration. Ramp weeks appear in the Delta Table's backup, labeled as ramp, visible to anyone who asks. They are excluded from the primary delta.

Watch what including them would have done, because it cuts in both directions at once, and that is the tell that flattery, not principle, drives most window choices. Including the ramp would flatter the error-rate trend: week one's verified error rate was 3.9 percent, so a chart starting there shows a dramatic plunge to 1.4 that is mostly just the team learning the gate. And it would damage the cycle-time result: ramp weeks ran slow, so including them drags the median from 1.8 toward 2.1. A team choosing windows for effect would include the ramp on the error slide and exclude it on the cycle-time slide, and somewhere a hostile analyst would notice the windows differ and correctly conclude that everything else is negotiable too. The pre-declaration decides, once, for every metric, whichever way it cuts. That is the entire value of having made it before launch: in week seven it is not a choice, it is a clause.

Discipline four: count the whole system

Now the discipline with the largest gap between the flattering number and the true one. The redesign did not just speed up handling; it added machinery, and the machinery costs money every week it runs. The human gate spends 3.8 reviewer minutes per reviewed item. The audit layer samples 100 documents a week and burns real hours verifying them. The validation layer bounces 3.1 percent of items into a rejected lane that gets full manual handling. None of that existed in the baseline process, and a delta that counts the savings but not the machinery is the vendor arithmetic you learned to catch in Level 1, now pointed at your own slide. The fullness rule does not retire when the numbers become yours.

So the honest headline metric is system cost per exception, baseline versus pilot, with everything inside it. Illustratively: the baseline's 31.20 dollars per exception (18.70 of handling time, 12.50 of rework and escalation cost) against the pilot's 24.70 (8.30 of clerk handling and gate minutes, 0.90 of audit sampling amortized per item, 0.65 for the rejected lane's manual handling, 4.10 of license and platform cost per item, 10.75 of residual rework and escalation). That is a 21 percent net improvement. Real money, at roughly 260 exceptions a week about 1,700 dollars weekly, and it is honest. It is also nowhere near the 44 percent the cycle-time-only slide claimed, and the gap between 44 and 21 is exactly the gap the CFO's analyst exists to find.

Two bookkeeping distinctions keep this number clean. First, one-time versus run-rate, the distinction every chief financial officer (CFO) applies by reflex: the 38,000 dollars of transition cost (integration work, schema build, training hours, the IT ticket) is reported separately as a one-time investment with a payback line (at 1,700 dollars a week, roughly 22 weeks), never smeared into the per-exception run-rate where it would poison the steady-state comparison. Second, resist what this program calls the two-slide temptation: the big gross number on the main slide for applause, the net buried in the appendix for deniability. Committees remember the applause number and forget the appendix, which means the appendix was decoration and the gross number was the claim. One table. Net first. Gross available underneath, labeled as gross, for anyone who wants the decomposition. The transformer who leads with the smaller, truer number is making a purchase: the room's permanent trust, at the price of one moment of lesser applause.

Discipline five: say what you cannot claim

The final discipline is a paragraph, and the rule is about authorship: the transformer writes the residual-confound paragraph before anyone else writes it for them. Every pilot has residual confounds that no control removes, and they will be named eventually, either by you in the report or by a skeptic in the meeting. Written by you, they are boundaries on the evidence. Written by the skeptic, they are holes in it. Same facts, opposite verdict.

For the invoice-exception pilot, the paragraph names three. Six weeks is one season: the window contains one month-end and no quarter-end, so no claim extends beyond the observed window's structure, and quarter-end behavior remains untested. One site: the pilot ran in one shared-services center with one team's tenure profile, so no cross-site claim exists yet. And the team knew it was measured: some observer effect survives even good design, so the paragraph bounds it rather than denying it. The bounding argument is the week-five drift event, and it is the strongest sentence in the report: "we cannot fully separate the redesign's effect from the attention the pilot brought; the week-five drift event argues most of the effect is structural, because performance recovered by mechanism, not morale." When extraction quality degraded in week five, the counters caught it, the grounding pack was corrected, and the numbers recovered within days, without a rally, a pep talk, or extra effort. Effects that recover by mechanism belong to the design. Effects that recover by morale belong to the novelty, and fade with it. Name the confound, bound it with evidence, move on.

Then give the committee a vocabulary for weighing all of it, because executives cannot act on error bars they do not have and should not pretend to statistics nobody ran. The Delta Table assigns each metric one of three confidence words, by rules agreed in the measurement plan before results existed. CLEAR: the delta survives every control (like-window, mix, ramp), and the mechanism producing it is understood and named. PROBABLE: the delta survives the controls, but the window limits it; one season or one cycle observed, mechanism plausible but not fully isolated. SUGGESTIVE: the direction is consistent, but the volume is too thin to hold weight; a real signal or a coincidence, and honesty cannot yet say which. Three words, defined by rule, assigned before the meeting, defended row by row. An executive can weight evidence with this vocabulary in a way no p-value theater would allow, and because the rules predate the results, nobody can accuse the words of being chosen to sell.

The Delta Table: Six Weeks of Evidence on One Honest Page

Now assemble the artifact in full. Every figure is illustrative; every row shows the shape your own table should take. Pilot values are weeks three to six, mix-adjusted where marked; baselines are the like-window reconstruction.

MetricBaseline (like window)Pilot (wks 3-6)DeltaControls appliedCaveat that survivesConfidence
Median cycle time3.4 days1.8 days-47%Like-window baseline; mix-adjusted (0.2d removed); ramp excludedOne month-end observed; no quarter-end in windowCLEAR
p90 cycle time9.1 days4.2 days-54%Like-window; ramp excludedTail sample is small by nature; one seasonCLEAR
Verified error rateNot measured1.4%Not computableDouble-check protocol, 100-doc weekly samplesBaseline never verified errors; nearest proxy is its 12% rework rate, a different definitionReported as a finding
Rework rate12%7.8%-4.2 ptsLike-window; identical frozen definition; ramp excludedOne month-end observed; reopens can lag the windowPROBABLE
Net cost per exception$31.20$24.70-21%Verification tax counted (gate, sampling, rejected lane); one-time costs separatedCost model assumptions documented; one site's loaded ratesPROBABLE
Clerk overtime hoursBaseline levelFlatNo changeGuardrail metric per charterNone; guardrail heldCLEAR
Gate queue latency4.0h budget2.1h avgWithin budgetFull window incl. month-end peakBudget assumes current staffing of the gateCLEAR

Walk the rows that carry teaching weight. The p90 row first, in operator terms: p90 is the 90th percentile, the boundary of the slowest tenth of items, the exceptions that age in the queue while a vendor calls twice a week. It fell further than the median did, 9.1 days to 4.2, and the mechanism is named, which is what earns the CLEAR: the redesign's queue restructure routes aged items to the front of the gate instead of letting them sink, so the tail compressed by design. A delta with a named mechanism is an explanation; a delta without one is a coincidence you are taking credit for. When you can say this caused that, and point to the design decision, the confidence word upgrades itself.

Now the row that shocks every committee that sees its first honest Delta Table: verified error rate, baseline not measured. The old process never verified its own outputs. Nobody sampled closed exceptions and checked them; the only error signal the baseline ever produced was its rework rate, the 12 percent of items that came back loudly enough to be reopened. Rework is not an error rate; it is the subset of errors noisy enough to return. The pilot's 1.4 percent verified error rate comes from actually checking a random 100-document sample every week, a measurement the baseline has no twin for. So the row refuses to fabricate a delta, states the definitional caveat in full, and reports the absence itself as a finding, because it is one, and a big one: the organization ran this process for years without knowing its error rate, and the pilot's first contribution was making the question answerable. Sometimes the baseline's silence is the loudest number on the page. Report it as one; never paper over it by quietly comparing 1.4 against 12 as if they measured the same thing.

And the net cost row carries the headline, 21 percent, wearing PROBABLE rather than CLEAR, and the discipline is in accepting that word for your best number: the delta survives every control, but the cost model's assumptions and the single observed season cap the confidence at window-limited. A transformer who marks their own headline PROBABLE has told the room something more valuable than the number itself: that the words mean what the rules say they mean, even when it costs the story some shine.

Reporting Language, and the Hostile Pre-Brief

A table does not present itself; a paragraph presents it, and the paragraph is where honest analysis is most often undone by enthusiastic prose. So write the paragraph with the same discipline as the rows, in the three-register structure that survives any room: what we know, what we believe, what we cannot yet say. Here it is for the invoice-exception pilot, verbatim, as copyable scaffolding for your own:

"What we know: against a baseline window matched for month-end structure, and adjusted for the pilot's easier work mix, median cycle time fell from 3.4 to 1.8 days and the slowest tenth of items improved even more, from 9.1 to 4.2 days; the mechanism is the queue restructure, and both results survive every control we applied. What we believe: the pilot process is cheaper to run, 24.70 dollars per exception against 31.20 with all verification costs counted, and rework is down from 12 to 7.8 percent; both figures are marked probable because we have observed one six-week season at one site. What we cannot yet say: how the process behaves at quarter-end, whether the results transfer to other sites, and precisely how much of the effect the pilot's own visibility contributed, though the week-five drift event argues the improvement is structural, because performance recovered by mechanism, not morale. The old process's true error rate was never measured; the pilot's was, weekly, and it ran at 1.4 percent verified."

Read it twice and notice what it never does: it never rounds up, never borrows the gross number for drama, never claims a season it did not observe, and it spends its strongest sentence on the mechanism, not the percentage. This is what evidence sounds like when it expects to be cross-examined.

Then arrange the cross-examination yourself, on your own schedule, in private. This is the hostile pre-brief: a week before the steering committee, invite the CFO's analyst, the sharpest skeptic available, to spend one hour attacking the draft Delta Table. Hand over the table, the segment data, the queries, the cost model, and one instruction: break it. This is the Level 2 pre-wire discipline applied to numbers, and its arithmetic is ruthless in your favor. Every catch the analyst makes in private is a catch that does not happen on stage; a flaw found in pre-brief costs you an edit, while the same flaw found in committee costs you the table's authority and a piece of your name. And the deeper prize is allegiance: an analyst who spent an hour trying to break your delta and helped tighten the mix adjustment is no longer the delta's natural predator. When someone in the meeting squints at the 21 percent, the most credible defender in the room is the person everyone knows was paid to doubt it. In the running example the analyst finds one real issue (two late reopens that belonged inside the rework window) and the corrected rate, 7.8 to 8.1 percent, goes to committee with the correction noted. The number got worse; the report got stronger. That trade is available every single time, and only before the meeting.

The failure story: the 38 percent that became 11

Now the mirror, assembled from the standard patterns with illustrative figures. A service-desk pilot at another company reports "resolution time down 38 percent" to a delighted committee. Applause; scaling approved on the spot. Weeks later, an analyst assigned to size the rollout budget starts pulling threads, because sizing requires the real number. The 38 percent, it turns out, compared pilot weeks against the single worst baseline quarter on record, chosen "because it was the most recent complete data," a phrase that sounded like rigor and functioned as selection. And nobody adjusted for a ticket-mix change that everyone in the room had personally discussed months earlier: an entire category of password resets had been automated away before the pilot began, draining the easy tickets out of the baseline and flattering every comparison since. Recomputed against a like window, mix-adjusted, the delta is 11 percent.

Sit with the cruelest part: 11 percent is a good result. Still positive, still fundable, still a rollout worth doing. But the number in everyone's memory is 38, so 11 does not land as a solid win; it lands as a 27-point retraction. The rollout proceeds under a cloud. The steering committee's trust does not reset to neutral; it resets to audit: the transformer's next table is checked line by line, every window questioned, every adjustment suspected, and the checking tax persists for quarters. The 38 percent bought one meeting of applause and paid for it with a year of doubt, on a pilot that honest arithmetic would have carried anyway. That is the general law, and it is worth engraving: an inflated delta is a loan against your next report, at loan-shark rates. The honest number, reported first, with its caveats attached, is the only version of the win you get to keep.

And with the Delta Table standing, pre-briefed, and survivable, the chapter reaches its final question. The table is evidence, not a verdict; it says what happened, not what to do. Turning evidence into a decision (scale this, iterate on it, or kill it, in writing, against the charter's pre-committed criteria) is the decision memo, and it is the next lesson.

What to Do Monday Morning

If you have a pilot ending, or a delta already circulating, this week's work is to make the number survivable before the meeting where it must survive.

  1. Rebuild your delta against a like-window baseline. Take whatever comparison your pilot currently claims and reconstruct the baseline window to match the pilot window's structure: same length, same month-end or seasonal content. If the number moves, in either direction, the move itself goes in the report, because volunteering it is what makes the rest believable.
  2. Run the mix adjustment and show your work. Compare each work segment to its own baseline, reweight the baseline to the pilot's mix, and state the plain-words version: "we asked what the baseline would have scored on the pilot's mix." Publish the segment table beside the aggregate, always.
  3. Compute the net, including the verification tax. Add up gate minutes, sampling hours, and the rejected lane's handling; separate one-time transition costs from run-rate with a payback line. Make system cost per unit the headline and refuse the two-slide temptation: one table, net first, gross underneath.
  4. Write the cannot-claim paragraph before your sponsor drafts the slide. One season, one site, observer effect: name each residual confound, bound it with your best evidence (a drift-recovery event is gold here), and assign every metric its confidence word by the pre-agreed rules: clear, probable, or suggestive.
  5. Book the hostile pre-brief. One hour, this week, with the most skeptical analyst you can recruit, ideally the CFO's. Hand over the draft Delta Table and the queries behind it, and ask them to break it. Every private catch is a public catch prevented, and every catch they help fix makes them the table's defender in the room.

Key Takeaways

  • Compute every delta the way a hostile analyst would compute it, and attach the caveats yourself, because a delta that arrives naked gets dressed by whoever presents it next.
  • Compare like windows: match the baseline window's structure (length, month-end, season) to the pilot's, never the smoothed annual average; in the running example the honest comparator was 3.4 days, not 3.2, and the rule binds whichever way it cuts.
  • Segment before you average: compare each work type to its own baseline, then reweight; the pilot's 4 percent drift toward easy items was worth 0.2 days of fake improvement, computed openly and removed with the plain-words mix adjustment.
  • Honor the pre-declared window: report ramp weeks labeled as ramp and exclude them from the primary delta for every metric alike, because including them flatters error trends while damaging cycle time, and flattery must never pick the window.
  • Count the whole system: the honest headline is net cost per unit with the verification tax inside it (31.20 to 24.70 dollars, 21 percent) and one-time costs separated, not the 44 percent the cycle-time-only slide would claim; one table, net first, gross available.
  • Write the residual-confound paragraph first, bounding one season, one site, and the observer effect, and assign each metric a rule-based confidence word (clear, probable, suggestive) so executives can weight evidence without pretending to statistics.
  • Report a missing baseline as a finding: the old process's error rate was never measured, and comparing the pilot's verified 1.4 percent against the baseline's 12 percent rework rate requires the definitional caveat spelled out in full.
  • Pre-brief the skeptic: an hour with the CFO's analyst before the committee converts every private catch into a public catch prevented, and remember the failure story's law: an inflated delta (the 38 percent that recomputed to 11) is a loan against your next report, at loan-shark rates.