Instrumenting the Process: Baselines and Counters
Month three of the pilot, and the steering committee wants numbers. The team lead has none, so she offers what she has: "the team says it feels much faster, and the vendor dashboard shows heavy usage." Across the table, the delegate the chief financial officer (CFO) sent instead of coming himself writes exactly four words in his notebook: "no numbers, month three." That meeting was not lost in month three. It was lost fourteen weeks earlier, in the quiet week before launch, when nobody built the measurement plan because everyone was busy with the exciting parts. Here is the unforgiving physics of pilot evidence: you cannot instrument the past. The log entry that was never written cannot be queried. The timestamp nobody captured cannot be reconstructed. Every number this pilot will ever report, to the committee, to the CFO, to the auditor who shows up in eleven months, is being decided right now, before launch, by three unglamorous choices: what gets logged, where, and by whom. This lesson is the last construction task of the redesign, and it is the one that decides whether everything else you built can ever be proven to have worked.
You Cannot Instrument the Past
Across this chapter you have built the operating system of the redesigned invoice-exception process: the standard operating procedure (SOP) rewritten around the AI step, the grounding pack that feeds the model your process truth instead of its imagination, the input and output schemas that force structure onto every handoff, and the verification architecture whose layers catch errors before they compound. One construction task remains, and it is the one teams skip most often because it produces no visible behavior at launch: the instrumentation. The counters. The measurement plan.
Why must it be built before launch rather than bolted on when the committee starts asking? Because measurement is the one component with no retroactive install. You can tighten a prompt in week four. You can add a validation rule in week six. But if nobody logged when items entered the gate queue in weeks one through five, that data does not exist and never will. A pilot that launches uninstrumented has not postponed its measurement decision; it has already made it. It has chosen to argue from anecdotes in month three, and anecdotes lose to the first person in the room who asks "compared to what?"
Set the standard where it actually needs to be. The standard is not "we collect data." Almost every failed pilot collected data; the vendor dashboard alone generates plenty. The standard is numbers honest enough to survive two specific stress tests: a skeptical CFO's analyst with an afternoon to spend, and an annual audit with the power to demand lineage on any figure. Surviving those tests means three things, and each one is a design requirement, not a hope. Every reported figure has a counter behind it: a defined, owned, queryable count, not a recollection. Every counter has a written definition: what is included, what is excluded, over what window. And every definition matches the baseline pack's definition exactly, because comparability is the whole game. In Level 2 you built the baseline pack, the sworn before-photo of this process: volume, median cycle time, error and rework rate, cost per unit, each with measurement notes. If the live instrumentation defines "rework" even slightly differently than the baseline did, the before-and-after comparison silently dies, and it dies in the worst possible way: it keeps producing numbers that look comparable and are not.
This is where the program's oldest slogan completes itself. "Baseline or it didn't happen" was always only half the sentence. The full sentence is: baseline, plus live counters defined identically to the baseline, equals evidence. Either half alone is decoration. And this is also the operational meaning of the MIT finding you have carried since Level 1: 95 percent of enterprise generative AI pilots showed no measurable return. Read the adjective, not the noun. Measurable is not a property the universe assigns to your pilot in month nine; it is a property you either build in before launch or permanently forfeit. McKinsey's State of AI survey shows the same coin from the other face: only about 39 percent of organizations using AI can attribute any earnings impact to it at all. Attribution is instrumentation, purchased in advance.
You cannot instrument the past. Every number the pilot will ever report is being decided before launch, by what gets logged, where, and by whom.
The artifact this lesson builds is the Measurement Plan: one page, four parts. Part one is the counter list, every number the pilot will report, each with a definition, a source, an owner, and its baseline-pack twin. Part two is the event log specification, the single data structure all counters are computed from. Part three is the weekly tally ritual, the standing thirty minutes that turns the log into the one-page weekly report. Part four is the anti-gaming notes, where you name in advance the pressures that will bend the numbers and the structural countermeasures already in place. One page. Roughly three days of work. It is the cheapest insurance in the entire transformation program, and we will build it slowly, part by part.
The Counter List: Three Sources, in Priority Order
A counter is a number the pilot commits to producing on demand: defined, sourced, owned. The temptation is to brainstorm counters, which produces forty of them, thirty of which nobody will ever read. Resist it. The counter list is not brainstormed; it is derived, from exactly three sources, in priority order.
Source one: the charter's success criteria and kill condition
Start with the pilot charter you signed at the end of Level 2, because the charter is a set of promises and every promise needs a measuring instrument. Run the mapping exercise explicitly, criterion by criterion. The invoice-exception charter promises verified extraction accuracy at or above 85 percent at week six: that maps to a counter (verified accuracy, measured by the double-check protocol on a 100-document random sample, not by the tool's self-reported score). It promises median cycle time under 2.0 days by week ten: that maps to a counter over the event log's timestamps. It caps the rework rate at the baseline's 11 percent: another counter. And the kill condition, verified accuracy below 85 percent at week six, maps to the same accuracy counter, which is precisely why that counter's independence matters more than any other's.
The rule of the exercise is blunt: every charter criterion must map to a counter, or the criterion is unmeasurable, and an unmeasurable criterion makes the charter fiction. If you find a criterion that maps to nothing ("improved team confidence in the process"), you have two honest options: define a real instrument for it, or amend the charter now, before launch, while amendment is still cheap. What you may not do is launch with a promise no counter can check, because that promise will be graded by whoever tells the best story in month four.
Source two: the verification architecture's layer outputs
The previous lesson built verification into the flow as layered controls, each with an error budget. Every layer of that architecture throws off events, and those events are counters waiting to be named: validation rejections (items the schema checks bounced before any human saw them), gate catches (items the human gate corrected or rejected, with the override reason code attached to each), audit findings (errors the sampling layer found in items that had already passed the gate), and downstream escapes (errors that got all the way to a vendor, a payment, or the general ledger before anyone noticed). You need these counters for a structural reason: the error-budget arithmetic does not work without them. A 2 percent verified error budget is meaningless unless you can actually count errors at each layer and show where in the stack they were caught. A verification architecture without its counters is a smoke detector with no test button: possibly working, never demonstrable.
Source three: the baseline pack's four metrics, re-measured identically
Finally, the baseline pack's four metrics (volume, median cycle time, error and rework rate, cost per unit) each get a live counter, and here the discipline is a single word: identically. Same definitions. Same measurement windows. Same segmentation. The baseline pack's measurement notes recorded exactly how each number was produced: rework defined as any exception reopened after closure or reversed downstream, cycle time as created-to-closed clock time including weekends, volume segmented by exception type. Do not paraphrase those notes into the measurement plan. Print them into it, verbatim, and mark each live counter with its baseline-pack twin. Think of it as the discipline of the identical camera: the before-photo and the after-photo are only comparable if they are taken by the same camera, same lens, same settings, same framing. A pilot that measures cycle time business-days-only against a baseline that measured calendar days has not improved anything; it has changed cameras and called the blur progress.
Three sources, and notice what is not among them: the vendor's dashboard. Vendor metrics may be interesting, but nothing the vendor self-reports satisfies a charter clause, feeds the error budget, or twins with the baseline. They are commentary, not evidence, and the plan should say so in writing.
The Event Log: Counters Are Queries, Not Tallies
Now the foundation underneath every counter. Where do the numbers physically come from? The answer, and this is the lesson's central engineering idea, is one data structure: the event log. One row per item per state change, with a timestamp, the item's identifier, the state it entered, and who or what moved it there. An invoice exception is received: row. It passes schema validation: row. It is routed to the AI lane: row. The AI proposes an extraction: row, carrying the confidence band. The gate reviews it: row, carrying the reviewer and any override reason code. It is released: row. It gets reopened two weeks later: row, and that row is the one that keeps the rework counter honest.
Once the log exists, every counter on the list becomes a query over the log rather than a number a person keeps. Median cycle time is the median of released-timestamp minus received-timestamp. The gate catch rate is gate rows with corrections divided by gate rows. The rework rate is items with a reopened row divided by items released in the window. This distinction, counters as queries versus counters as hand-kept tallies, deserves a paragraph of its own, because it separates auditable pilots from anecdotal ones. Hand tallies drift: they get updated from memory on Friday afternoon, they quietly exclude the awkward item, they cannot be recomputed when someone challenges them, and two people keeping the same tally produce two different numbers. Logs add up: a query run today and re-run in front of the auditor next year returns the same figure from the same rows, and anyone who doubts a number can be handed the query and the log and invited to check. When the CFO's analyst asks "how do you know?", the difference between "Marta keeps a spreadsheet" and "it is a query over the event log, here is the definition" is the difference between an opinion and evidence.
Here is the good news that makes this the cheapest part of the build: the event log usually already half-exists. Your ticketing or workflow system has been recording created, assigned, closed, and reopened timestamps for years; the enterprise resource planning (ERP) system knows when payments and reversals happened. The instrumentation task is mostly turning on and structuring what your systems record anyway: agreeing which system field means which state, making sure state changes are actually captured rather than skipped, and adding the two or three fields nobody currently captures. In the invoice-exception process, those missing fields are exactly two: the AI's confidence band on each proposal, and the override reason code the gate reviewer selects when correcting one. And notice where they come from: both travel into the log through the output schema you built earlier in this chapter. The schema forces the model to declare its confidence in a structured field; the gate spec forces the reviewer to code every override. The chapter's artifacts interlock one last time: the schema feeds the log, the log feeds the counters, the counters prove the charter. That is what a designed process looks like from the inside.
Specify the log in the measurement plan as a short table: the list of states, the fields carried on each row (timestamp, item ID, state, actor, confidence band where relevant, reason code where relevant), and the system of record for each. If a state change happens today by email with no system trace, that is your instrumentation gap, and closing it is a launch prerequisite, not a fast-follow.
Where the Numbers Lie: Five Traps and Their Countermeasures
If instrumentation were only plumbing, this would be a short lesson. It is not, because pilot numbers face pressures that plumbing does not: everyone in the room wants them to be good. Nobody has to lie for the numbers to bend; the numbers bend on their own, along five well-worn grooves. A measurement plan that does not name these traps and pre-install the countermeasures is a plan for producing impressive numbers that die under questioning. Learn all five; they are the lesson's centerpiece.
Trap one: the enthusiasm bump
The first weeks of any pilot measure novelty, not steady state. The team is excited, attention is maximal, and, more subtly, everyone routes their cleanest work into the new lane: the tidy exceptions, the cooperative vendors, the well-formed documents. Week-two numbers are the process on its best behavior, and reporting them as the pilot's performance is like timing a commute on a public holiday. Countermeasure: pre-declare the measurement window in the plan, starting after the ramp period the charter already defines, and report ramp weeks separately, clearly labeled as ramp. The declaration must be made before launch; a window chosen after you have seen the data is not a window, it is a selection.
Trap two: mix shift
The pilot period never receives the same work mix as the baseline period. Quarter-end floods the queue with a different exception profile; a major vendor changes invoice format; a seasonal product line spikes one segment. If the pilot months happen to draw an easier mix than the baseline months, aggregate cycle time improves with no help from the AI at all, and the improvement evaporates the moment the mix reverts. Countermeasure: segment the live numbers by the same segments the baseline pack used (for the invoice process: missing purchase order, price mismatch, quantity mismatch) and compare within segment, not just in aggregate. This is the volume-mix control from the Level 2 baseline lesson made operational: the baseline was segmented precisely so that this comparison would someday be possible. Report the aggregate too, but never without the segment table beside it.
Trap three: definition drift
Somewhere around week seven, someone helpful proposes a clarification: "surely items reopened just to attach a document shouldn't count as rework?" The definition tightens, the rework rate improves, and no process changed: the improvement is made of vocabulary. Definition drift is the quietest trap because each individual clarification sounds reasonable and the cumulative effect is a metric that no longer means what the baseline meant. Countermeasure: definitions are frozen in the measurement plan at launch, copied verbatim from the baseline pack. Changing one requires a logged amendment with a stated reason, and during any transition the plan requires dual reporting: the metric under the old definition and the new one, side by side, until the committee formally retires the old line. Drift that must be logged and dual-reported mostly stops being proposed.
Trap four: selection laundering
The most dangerous trap, because it can happen with everyone acting in good faith. Items that the new lane handles badly get quietly rerouted out of it: to a "special handling" queue, to a senior specialist, back to the old manual path. Each reroute is individually defensible ("this one was weird"). Collectively they launder the denominator: the measured lane looks superb because everything difficult has been selected out of it, while the process as a whole has not improved at all. Countermeasure: the conservation check, and it deserves to be taught as what it is, the accountant's oldest trick applied to process measurement. Every week, the report must balance one line: items in equals items out plus items pending, with every exit route named and counted (released, reverted to manual, escalated, cancelled). Nothing may leave the denominator unexplained. Accountants have used balance checks for centuries because totals catch what stories hide, and this single line catches most measurement gaming ever attempted: the moment eight items slide into an unmeasured side queue, the conservation line stops balancing and the report itself raises its hand.
Trap five: the unmeasured verification tax
The pilot reports the AI's time saved while the new human verification time goes uncounted. The gate you designed two lessons ago costs real reviewer minutes on every item it inspects; the audit layer costs sampling hours every week. Reporting "22 minutes of handling reduced to 6" while omitting the 4 minutes of gate review on every item is not a rounding error, it is the oldest trick in the vendor playbook, and you learned to catch it in the Level 1 return-on-investment lessons: the fullness rule, count all the costs, including the boring ones. Now the rule points at you. A transformer who counts the vendor's license and the model's speed but not their own gate's minutes is doing to the steering committee exactly what vendors spent years doing to them. Countermeasure: gate minutes and audit hours are counters on the list, first-class, reported every week, and the headline efficiency number is always the honest net: time saved minus verification time added. If the net is genuinely good, honesty costs you nothing. If the net is bad, you need to know before the committee's analyst discovers it for you.
The Weekly Tally Ritual and the Anti-Gaming Notes
A counter list and an event log produce nothing by themselves; something has to read them, on a rhythm, in public. Part three of the measurement plan is the weekly tally ritual: thirty minutes, same day each week, same person running the same saved queries, producing the same one-page weekly. The page carries five blocks: volumes in and out with the conservation line, verified error rate against the error budget, cycle time against the baseline (aggregate and by segment), escapes and audit findings, and the status of the escalation ladder from the verification architecture. Five blocks, one page, every week, no exceptions, including the bad weeks. Especially the bad weeks.
Why ritualize it? Because the ritual matters as much as the numbers. A metric nobody reads weekly is a metric that gets backfilled monthly, and backfilled metrics are fiction with a lag: reconstructed from memory, smoothed by hindsight, and assembled under exactly the deadline pressure that bends numbers toward what the assembler hopes is true. The weekly rhythm is also what makes the counters trustworthy at decision time: by week twelve, the one-pager has been produced eleven times by the same queries, and its credibility is a matter of record rather than assertion. Hold the format steady, resist the urge to redesign the page every month, and file every edition. The next chapter, which takes this pilot live, builds its whole operating cadence on this page: the weekly review, the steering updates, and ultimately the scale, iterate, or kill decision will consume the weekly one-pager as their raw material. You are not writing a report; you are laying the evidence trail the verdict will stand on.
Part four of the plan, the anti-gaming notes, is the section people feel awkward writing and are always glad they wrote. Its premise: honest numbers are a system property, not a virtue. You will be under pressure by month three, and the plan should assume it. So name the pressures in advance, in writing, while nobody is applying them yet: the sponsor wants green slides for their own chain of command; the team wants the pilot to survive because their credibility is attached to it; the vendor wants a case study with a big percentage in the headline. None of these people are villains, and every one of them will, entirely sincerely, prefer the flattering window, the convenient definition, the aggregate without the segment table. Then write down the structural answers already in place: measurement independence (the charter's role clause already says whoever measures has no compensation or reported performance tied to the outcome; the plan names the person), the conservation check (gaming the denominator now breaks a published line), and the audit layer sampling the reporters as well as the process (layer three's random sample re-verifies items the weekly report already counted, which means the report itself is subject to spot-checks). When month-three pressure arrives, you do not have to win an argument about integrity; you point at a section everyone approved before the pressure existed. That is the entire trick of this program, applied one more time: pre-commitment beats willpower.
The Worked Example: The Invoice-Exception Measurement Plan, and the Number That Died in the Room
Time to render the artifact in full. All numbers that follow are hypothetical, an illustration carried forward from this program's running example, not research findings.
The transformer builds the measurement plan in roughly three days of focused work, plus one information-technology (IT) request. The counter list comes out at eleven counters: three from the charter (verified accuracy, median cycle time, rework rate), five from the verification architecture (validation rejections, gate catches by reason code, gate minutes per item, audit findings, downstream escapes), and the baseline four re-measured identically (volume folds into the charter counters' denominators, plus cost per unit as the honest-net computation). Here are four of the eleven in full detail, the way each row of the plan should look:
| Counter | Definition (frozen) | Source | Owner | Baseline twin |
|---|---|---|---|---|
| Verified accuracy | Share of AI-proposed extractions confirmed correct by the double-check protocol on a 100-document random weekly sample; tool self-reports excluded. | Audit sample results logged against event-log item IDs | Assessor | None (new metric); governs charter Clause 3 and the kill condition |
| Median cycle time | Median of released minus received timestamps, clock time including weekends, all in-scope items in window; segmented by exception type. | Event log query | Assessor | Baseline pack: median 3.2 days, p90 9 days, same definition verbatim |
| Rework rate | Items with a reopened row after release, or reversed downstream (credit note, repayment), divided by items released in window. | Event log reopened rows cross-checked against ERP reversals | Process owner | Baseline pack: 11 percent, definition copied verbatim from its measurement notes |
| Gate minutes per item | Reviewer minutes logged at the gate divided by items gate-reviewed; feeds the honest-net efficiency line (handling time saved minus verification time added). | Gate queue timestamps in the event log | Team lead | Baseline pack: 22 minutes average total handling per exception |
The event log specification lists nine states, each carried as one row with timestamp, item ID, and actor: received, validation-passed, validation-rejected, routed (AI lane or manual lane), AI-proposed (carrying the confidence band), gate-reviewed (carrying the override reason code where applied), released, reopened, and escalated. Eight of the nine already exist as ticketing-system events. The IT request covers the two missing fields, confidence band and reason code, both delivered into the log by the output schema and the gate form. IT sizes the request at four days of queue time and under a day of work. Three days of transformer time and one small ticket: the cheapest insurance in the program, purchased before the only week in which it can be purchased.
Week five: the plan defends itself
Fast-forward into the pilot to watch the instrumentation earn its keep. The week-five one-pager reads, in condensed form: volume in 268; conservation line clean (268 in equals 224 released plus 31 pending plus 9 reverted to manual, and reverted items stay in the denominator of the process-level numbers, so the reversion is visible, not laundered). Verified error 1.6 percent against the 2.0 percent budget, from the 100-document sample. Median cycle time 1.7 days against the 3.2-day baseline, with a mix-shift note attached: missing-PO items ran 2.4 days against their baseline segment's 4.1, price mismatches 1.1 against 2.6, and the note flags that this week's mix leaned toward price mismatches, so the aggregate flatters slightly and the segment lines are the honest comparison. Gate minutes: 3.8 per reviewed item, honest net still strongly positive. And one line under "amendments": a well-meaning proposal to stop counting document-only reopens as rework was logged, discussed, and refused; definition unchanged, dual reporting not triggered. That last line is the plan visibly defending itself: the drift attempt is not a corridor conversation that quietly won, it is a logged event that publicly lost.
The mirror: month four, three questions
Now the failure story, the same movie with the instrumentation left out, assembled from the standard patterns with illustrative figures. Another company, another document-heavy pilot, no measurement plan. In month four the team presents a triumphant slide: cycle time down 40 percent. Applause. Then the CFO's analyst, who has an afternoon and a calculator, asks three questions. What window does that figure cover? (Answer, once reconstructed: from launch day, meaning the enthusiasm bump is inside the average.) What was the work mix versus baseline? (Answer: nobody segmented; it turns out the pilot months drew heavily on the easiest document type.) And where did the 31 reopened items go? (Answer, after an uncomfortable pause: they were rerouted to a "special handling" queue that sits outside the measured lane.) The number dies in the room, and with it the pilot's credibility. Note carefully what did not happen: nobody lied. Every choice had a plausible reason at the time. The instrumentation had simply been designed by optimism, window by hope, mix by accident, denominator by convenience. The bitterest part is the epitaph: the pilot was probably genuinely good. The tool likely worked; the team likely improved the process. And nobody will ever know, because the evidence that could have proven it was never collected and cannot be collected now. You met this story's smaller sibling in the Level 2 baseline lesson, where a pilot without a before-photo could not prove its gains. This is the same funeral one level up: a before-photo existed, and the after-photo was taken with a different camera, pointed at a curated subject.
And with that, the chapter closes. Look at what is now standing: a process redesigned around the AI step rather than beneath it, an SOP that tells humans and models exactly who does what, a grounding pack that keeps the AI inside your process truth, schemas that structure every handoff, a verification architecture with layers, budgets, and a ladder, and now the instrumentation that will testify to all of it. The process is built and wired. Chapter 3.3 pushes the button: pilot design, operating cadence, honest measurement in flight, and the decision every pilot must eventually face, scale, iterate, or kill, made on the evidence your counters are about to start writing.
What to Do Monday Morning
The measurement plan is three days of work, and Monday is day one. Here is the sequence.
- Run the charter mapping exercise. Take your pilot charter (or the nearest thing your initiative has to one) and map every success criterion and the kill condition to a named counter with a definition and a source. Any criterion that maps to nothing gets fixed this week: define an instrument or amend the charter.
- Spec the event log. List the states an item passes through in the redesigned process, check which state changes your ticketing or workflow system already records, and identify the missing fields. Expect roughly two: in the invoice-exception example they are the confidence band and the override reason code. File the IT request now; queue time is why this is a Monday task, not a launch-week one.
- Freeze the definitions, copied verbatim. Open the baseline pack, copy its measurement notes word for word into the measurement plan, and mark each live counter with its baseline twin. Any change from here forward requires a logged amendment and dual reporting during transition.
- Put the weekly tally ritual on the calendar. Thirty minutes, same day every week, named owner, saved queries, one-page output with the conservation line on it. Recurring, starting the pilot's first week, with ramp weeks labeled as ramp per the pre-declared measurement window.
- Write the anti-gaming notes while nobody is pressuring you yet. Name the three pressures (sponsor, team, vendor), then document the structural answers: who measures and their independence, the conservation check, and the audit layer's right to sample the reported numbers themselves. Get the page approved alongside the rest of the plan, before launch, while every clause is still hypothetical and therefore still cheap to sign.
Key Takeaways
- Build the measurement plan before launch, because you cannot instrument the past: every number the pilot will ever report is decided by what gets logged, where, and by whom, and an uninstrumented pilot has already chosen to argue from anecdotes.
- Set the standard at audit-grade, not "we collect data": every reported figure has a counter, every counter has a frozen definition, and every definition matches the baseline pack verbatim, because changing the camera kills the before-and-after comparison silently.
- Derive the counter list from exactly three sources in priority order: the charter's success criteria and kill condition (every criterion must map to a counter or the charter is fiction), the verification architecture's layer outputs, and the baseline pack's four metrics re-measured identically.
- Found every counter on the event log, one row per item per state change with timestamps and IDs, so counters are reproducible queries rather than hand-kept tallies: tallies drift, logs add up, and the log mostly already exists in your workflow system plus two missing fields the schema delivers.
- Name the five traps in the plan and pre-install their countermeasures: the enthusiasm bump (pre-declared window), mix shift (within-segment comparison), definition drift (frozen definitions, logged amendments, dual reporting), selection laundering (the weekly conservation check: items in equals items out plus pending), and the unmeasured verification tax (count the gate's minutes and report the honest net).
- Run the weekly tally ritual without exception: thirty minutes, same day, same queries, one page with volumes, error versus budget, cycle time versus baseline, escapes, and ladder status, because backfilled metrics are fiction with a lag.
- Treat honest numbers as a system property, not a virtue: write the anti-gaming notes in advance, naming the sponsor's, team's, and vendor's pressures and the structural answers of measurement independence, conservation, and audit sampling.
- Remember what the MIT 95 percent actually convicts: measurability is a choice made before launch, and the roughly three days this plan costs is the difference between a pilot that ends in evidence and one that ends in an epitaph.
Skill.re