Building the Process Baseline Pack
Month six of the pilot, and the steering committee has finally arrived at the question everyone knew was coming. The AI triage tool for invoice exceptions has been live since spring. The team likes it. The vendor's dashboard shows healthy usage. And the finance director, who has been quiet for twenty minutes, asks the only question that matters: "How much faster are we, exactly, and compared to what?" The pilot lead opens a slide that says "significant improvement" and realizes, in real time, in front of nine people, that there is no number on it. Not because the tool failed. Because six months ago, when the process was still running the old way, nobody wrote down what the old way cost. The before-photo was never taken, and now the before is gone forever. This lesson is about the two weeks of unglamorous work that would have saved that meeting, that pilot, and possibly that career: building the Process Baseline Pack.
The Argument You Lose in Month Zero
Here is the uncomfortable arithmetic of pilot politics. Every "is this even working?" argument that erupts in month six was actually decided in month zero. If the baseline exists, the argument takes ten minutes: here is what the process cost before, here is what it costs now, here is the delta, signed and sourced. If the baseline does not exist, the argument cannot be won by anyone, ever, because there is nothing to compare against. The most common outcome is not a draw. It is a loss for the pilot, because in the absence of evidence, the default position of any finance function is that the money was spent and nothing provable came back.
This program has been building toward this lesson since its first hour. The evidence rule you met in Level 1, "baseline or it didn't happen," has so far been a principle. Today it becomes a deliverable with a page count. And the stakes are exactly the ones the research keeps confirming. MIT's GenAI Divide study found that 95 percent of enterprise generative AI pilots delivered no measurable profit-and-loss return, and as we noted when we first read that number honestly, it is as much an indictment of measurement as of performance: in most of those organizations, nobody instrumented the work so that value could be found even if it existed. McKinsey's State of AI survey shows the same failure from the other end: 88 percent of organizations use AI regularly, yet only about 39 percent can attribute any earnings impact to it. You cannot attribute what you never baselined. The 61 percent are not all failing. Many of them simply cannot prove they are succeeding, and in a budget review those two conditions are indistinguishable.
Notice that this cuts in both directions, and this is the part pilot enthusiasts miss. A missing baseline does not just let bad pilots survive on vibes. It kills good pilots that deserved to live. Remember the program's rule that a documented kill is a win: even that rule needs a baseline, because "we killed it" only counts as a disciplined decision if you can show the numbers that triggered the kill. Without a before-photo, every possible ending of your pilot is an evidence failure. The zombie pilot that should die but cannot be proven dead, and the genuine success that cannot be proven alive, are mirror images of the same month-zero omission.
One more thing about who checks this work, because it shapes how you build it. Your baseline will not be reviewed by a friendly colleague who wants the pilot to succeed. It will be reviewed, eventually, by the least friendly person in the room, usually holding a CFO title or acting on behalf of one, whose professional obligation is to disbelieve you until the numbers force otherwise. That is not hostility; the chief financial officer is doing their job, and doing it well is what protects the company from the 95 percent. Build the pack for that reader. If it survives them, it survives anyone.
The Before-Photo: What Goes in the Process Baseline Pack
The named artifact of this lesson, and the highest-stakes document you have built in this program so far, is the Process Baseline Pack: the sworn record of what one process costs to run today, before any tool touches it. It is the before-photo that every later claim about the pilot will be judged against. It has three layers, and all three are load-bearing.
- The four core metrics. Cycle time, volume, error and rework rate, and cost per unit, measured for the process as it actually runs, not as the SOP (standard operating procedure, the official written instructions) says it runs. The next section treats each one in operator depth.
- Measurement notes for every number. Each figure carries its definition (what exactly was counted), its window (which weeks or months), its source (which system, which export, which sample), and its known blind spots (what this number cannot see). A number without its notes is an opinion wearing a suit.
- A data-lineage line for every number. One sentence tracing the figure from raw source to final value: "Median cycle time computed from the created-to-closed timestamps of all 3,412 exception tickets in the workflow system, Jan through Mar export, file EXC_Q1.csv, recomputed by hand on a 40-ticket sample on April 9." Lineage is what turns "trust me" into "check me," and inviting the check is precisely what makes the document credible.
Why so much apparatus around four numbers? Because of an asymmetry you must internalize before you collect a single figure: a baseline with one wrong number is worse than no baseline at all. No baseline is an honest gap; everyone can see it and discount accordingly. A wrong baseline is a manufactured falsehood that compounds. If your baseline cycle time is accidentally too high, the pilot will show fake improvement and you will scale a tool that does nothing, which is how six-figure zombies are born. If it is accidentally too low, the pilot will show no improvement even when there is one, and a working tool dies. Either way, when the CFO's analyst rechecks your figures against the raw exports, and on any pilot worth real money they will, one broken number poisons the other three. The pack is built to survive that recheck, which is why the notes and the lineage are not decoration. They are the load-bearing walls.
Baseline or it didn't happen: the before-photo you take in month zero is the only evidence that will exist in month six, and a baseline with one wrong number is worse than no baseline at all.
The Four Metrics, Measured Like an Operator
Anyone can list these four metrics. The craft is in the definitions, because every one of them has a lazy version that will be taken apart in review and a disciplined version that holds. Take them one at a time, slowly.
Cycle time: report the median and the tail
Cycle time is elapsed clock time from trigger to completion: from the moment the work item exists (the exception is flagged, the request lands, the ticket opens) to the moment it is genuinely done, not merely touched. Two disciplines separate the operator's version from the amateur's. First, clock time, not effort time. A task that takes 25 minutes of actual work but sits in queues for three days has a cycle time of three days, and the customer, the auditor, and the month-end close all experience the three days. Second, never report the average alone. Averages hide tails, and in process work the tail is usually where the money is. Report the median and the 90th percentile. In plain terms: the median is the value in the middle, half of all items finish faster and half slower; the 90th percentile (p90) is the value that 90 percent of items beat and the worst 10 percent exceed. If your median is 3 days and your p90 is 9, you have two different processes wearing one name: a routine path and a nightmare path, and the nightmare path is where escalations, complaints, and late-payment penalties live. You already know why this matters from the previous lesson: the verified map's walkthrough found a queue where exceptions wait up to two days for a specialist. An average would smear that wait invisibly across every ticket. The p90 puts it on the record.
Volume: count it, then segment it
Volume is units of work per period: exceptions per month, requests per week, cases per day. The count is the easy half. The operator's half is the mix, because raw volume treats every unit as equal and units are never equal. Segment by type, and look specifically for the concentration pattern that almost every operational process exhibits: a minority of item types consuming a majority of effort. If 20 percent of your invoice exceptions (say, the missing purchase order ones) consume 60 percent of the team's handling time, that fact must be visible in the baseline, for two reasons. It changes where a pilot should aim, and it protects your comparison later: if the mix shifts between baseline and pilot (a big supplier onboards, a policy changes), an unsegmented baseline lets a skeptic claim the whole improvement was mix, and you will have no way to answer. Segmented volume is your insurance against the "you measured a different workload" attack.
Error and rework rate: define the error before you count it
Here is the metric where more baselines die in review than any other, and they die for a definitional reason, not a data reason. What counts as an error? Is a returned-for-clarification item an error? Is an exception that was resolved correctly but reopened by the supplier a failure of the process or of the supplier? If you do not answer these questions in writing before you measure, you will answer them in month six, live, in front of the steering committee, while the two sides of the argument each pick the definition that flatters their position. Write the definition first, get the process owner to agree to it, and put it in the measurement notes. A workable pattern: an error is any item that had to be reworked, reopened, reversed, or corrected downstream after being marked complete. Then use your verified process map to aim the measurement: the rework loops you documented in the mapping lesson are literally arrows pointing at where errors accumulate, so pull your samples from the steps those loops return to. Rework rate is the honest cousin of error rate, and often easier to extract from systems: count how many items traveled a loop more than once.
Cost per unit: the honest version, with assumptions stated
Cost per unit is the number the CFO reads first, and it is built from the other three plus two inputs of your own: loaded labor time per transaction multiplied by a loaded labor rate, plus allocated system costs. "Loaded" means the fully burdened figures: the labor rate includes salary plus benefits, employer taxes, and overhead (finance will hand you the standard loaded rate; use theirs, not your estimate), and the labor time includes all the touches, not just the main one, so the three-minute status check and the five-minute follow-up email count. Now the integrity test. Where does labor time per transaction come from? The lazy version asks the veteran, who says "about twenty minutes," and the veteran is wrong in the same way experts are always wrong about their own work: they quote the clean path and forget the interruptions, the lookups, and the second touch. The honest version measures: a short time study or a sampling exercise, where a handful of staff tally their actual touches on a set of real items for one or two weeks. And whichever loading assumptions you use, state them in the notes: which rate, whose overhead allocation, what the system-cost line includes. An unstated assumption is a tripwire you set for your own future self; the reviewer who finds it will wonder what else you did not state.
Where the Numbers Hide: AI-Assisted Collection
Two weeks was the promise, and the reason it is achievable alongside your day job is that AI does the part that used to take a month: reading thousands of rows and finding the patterns. Your organization is already recording most of what the baseline needs; it is just recording it in places nobody reads. The collection craft is knowing where to look and what to ask.
Where the raw material lives. Ticketing and workflow systems carry created, assigned, and closed timestamps, which are cycle time in raw form, plus status-change histories, which are your queues and rework loops. The ERP (enterprise resource planning system, the core software where transactions like invoices and orders actually post) carries volumes, categories, values, and reversal or credit records, which feed both volume mix and error rate. Standard operational reports carry monthly counts you can cross-check totals against. And for the handoffs that no system tracks, the ones your verified map marked as email-and-spreadsheet territory, email timestamps are the forensic fallback: the gap between "sending this over to you" and the reply is a measured queue wait, recoverable from a sample of threads.
What to ask the AI to do with an export. Pull a system export (a few months of exception tickets, say, with identifiers removed or masked per your organization's data rules, a discipline the next chapter treats properly), and put the assistant to work with prompts in the four-part pattern you built in the prompt library lesson. Three prompt jobs earn their place in the pack:
- Timing analysis: "Using the created and closed timestamps in this export, compute the cycle time for each ticket in days. Report the median, the 90th percentile, and a distribution in one-day buckets. List the ten longest-running tickets with their IDs so I can inspect them." That last clause matters: you are always asking for the trail back to raw rows, never just the summary.
- Category and mix analysis: "Group these tickets by the exception-type field. For each type, report the count, share of total volume, and median cycle time. Flag any type whose median cycle time is more than double the overall median."
- Anomaly and rework hunting: "Identify tickets that were reopened after closure or that returned to an earlier status more than once. Report how many, what share of total volume, and which types they concentrate in. List every assumption you made about the data while answering."
The assistant will do this in minutes, and it will also, some percentage of the time, do it wrong: misread a date format, silently drop rows it could not parse, invent a plausible-looking percentile, or average when you asked for a median. You know this. It is the entire reason the next section exists.
The Certification Pass: Every Figure Traced, Every Gap Confessed
The Verification Habit lesson gave you a tiered sampling rule with one absolute clause: 100 percent of all figures get verified, regardless of tier. In the baseline pack, that clause is the whole law, because the pack is nothing but figures, and every downstream decision leans on them. The certification pass is non-negotiable and it has a simple standard: every AI-extracted number gets traced to its raw source and recomputed on a sample before it enters the pack.
In practice this is less painful than it sounds. If the AI says the median cycle time is 3.2 days, you sort the export yourself and check the middle of the distribution, or hand-compute cycle times for 30 or 40 random tickets and confirm they are consistent with the claimed median and p90. If it says exception type A is 34 percent of volume, you run one pivot table and confirm the count. If it flagged 138 reopened tickets, you open ten of them and confirm they were genuinely reopened rather than administratively touched. An hour or two of recomputation, total, for a pack that will be attacked for a year. Then the human act that gives this section its name: you certify each figure. Your name goes on the pack. Not the model's, not the vendor's, yours. AI accelerated the collection; the accountability never moved.
The measurement note is the armor. Every number in the pack travels with its four-part note: definition, window, source, blind spots. The blind-spot line is the one inexperienced builders omit and the one that most reliably wins over a hostile reviewer, because it demonstrates you know your own instrument's limits. "Cycle time excludes the initial email triage before a ticket is created, estimated at 2 to 4 hours, because no timestamp exists for it" is not a weakness in the pack. It is proof that the pack was built by someone who cannot be caught out, because they caught themselves first. Hostile review destroys documents that pretend to be complete. It respects documents that state their own edges.
And when the data simply does not exist? This will happen; some part of your process lives entirely in inboxes and habit. The answer is never the veteran's confident guess, and never silence. The answer is the two-week manual sampling sprint: pick a representative sample of items, arm the people who touch them with tally sheets and timestamps (a shared spreadsheet with four columns does the job: item ID, step, start, end), and measure the sample directly. Two weeks of tallies on 60 real items beats a decade of "about twenty minutes" because the tally can be audited and the memory cannot. If even a sprint is impossible in the time you have, a documented estimate is acceptable on one condition: it is labeled as an estimate, with its basis and its uncertainty range stated in the measurement note. A documented estimate labeled as an estimate is honest. An unlabeled one is a trap that will spring in the exact meeting where you can least afford it.
The Worked Example: The Invoice-Exception Process Gets a Denominator
Time to pay off the storyline this chapter has been building. Across the last four lessons you inventoried the invoice-exception process, mapped it with AI assistance, caught the invented step, verified the map with the people who live in it, and wrote the SOP. Now the same mid-size company builds the baseline pack. All numbers that follow are hypothetical, an illustration to make the method concrete, not research findings.
The analyst runs the build in roughly two weeks alongside her normal work: three system exports, the three mining prompts, a certification pass on every figure, and one manual sampling exercise for the untracked email handoff. Here is the summary table of the finished pack.
| Metric | Baseline value | Measurement notes (condensed) |
|---|---|---|
| Volume | 1,150 exceptions per month | Window: Jan to Mar workflow export, 3,449 tickets. Segmented: missing PO, price mismatch, and quantity mismatch together are 78 percent of volume; missing PO alone is 34 percent and carries the longest cycle times. Blind spot: exceptions resolved informally by phone before ticket creation are not counted. |
| Cycle time | Median 3.2 days, p90 9 days | Created-to-closed timestamps, all tickets in window; recomputed by hand on a 40-ticket sample. The specialist queue wait found in the map walkthrough accounts for roughly 60 percent of the p90 tail. Clock time, including weekends. |
| Error / rework rate | 12 percent | Definition, agreed with the process owner before measurement: any exception reopened after closure or reversed downstream (credit note, repayment). Source: workflow reopen flags cross-checked against ERP reversal records; 10 flagged tickets inspected by hand. |
| Cost per unit | ~ $31 loaded | Labor time 22 minutes average per exception from a two-week tally-sheet sample of 64 items across 3 staff (the veteran's estimate had been 15 minutes); loaded rate $68 per hour per finance's standard table; plus $6 allocated system cost per ticket. Assumptions stated in full in the pack. |
One derived line completes the pack: 1,150 exceptions per month at roughly $31 each is an annual run rate of about $428,000. Feel what that number does to every future conversation. Before the pack, a pilot proposal for this process was "AI could help with invoice exceptions," a sentence with no stakes and no denominator. After the pack, it is "we spend about $428,000 a year on a process with a 12 percent rework rate and a nine-day tail, concentrated in three exception types; here is exactly what the pilot must move, and here is the sworn record we will measure it against." One version gets a polite nod. The other gets a budget, pre-committed success criteria, and, if the pilot fails, a documented kill that enhances rather than damages the team's credibility. Same process, same company. The only difference is two weeks of measurement.
Now the mirror image, the failure story from the opening scene, played out to its end. A comparable team at another company launches an AI triage pilot for the same kind of process with no baseline; the tool "feels faster," and for five months that feeling is allowed to stand in for evidence. When finance asks, the team does the only thing left: reconstructs a baseline from memory and whatever old reports they can find, after the fact. The CFO's analyst does what analysts do, pulls the raw exports, and discovers the reconstruction used a quarter that included a system migration and a staff shortage: the slowest quarter in two years, cherry-picked, whether by intent or by convenience. From that moment the team's every number is presumed cooked. The pilot, which by the team's own honest belief was genuinely working, is killed for unprovable claims. Understand what died there: possibly a true result, executed well, lost not to technology but to evidence. That is the zombie pilot's mirror image, the success nobody can defend, and both are the same failure wearing different endings: nobody took the before-photo.
One bridge before you go. The pack you just built covers the process half of the story; Chapter 2.3 turns the same discipline on the data those processes run on, because Gartner's finding that 63 percent of organizations lack AI-ready data practices is about to become your problem to assess. And keep the pack close: together with the verified map, the SOP, and the inventory row, it forms the process half of the Level 2 capstone assessment pack. You are, quietly, already halfway to the capstone.
What to Do Monday Morning
The pack takes about two weeks alongside your regular work. The sequence starts in the first hour.
- Pick the process your Process Inventory Register scored most painful. One process. The baseline discipline is per-process, and the first one teaches you the craft.
- Write the four metric definitions before touching any data. What starts and stops the cycle-time clock, how volume will be segmented, what counts as an error (get the process owner's written agreement on this one), and which loaded rate and loading assumptions cost per unit will use.
- Pull one system export and run the three mining prompts: timing analysis with median and p90, category and mix analysis, and anomaly and rework hunting. Always ask for the row-level trail, never just the summary.
- Run the certification pass: hand-verify 100 percent of extracted figures. Trace each number to its raw source and recompute it on a sample of 30 to 40 items. Where data does not exist, launch the two-week tally-sheet sampling sprint rather than accepting a guess.
- Draft the Process Baseline Pack: the four metrics, a measurement note (definition, window, source, blind spots) and a one-line data lineage for every number, and the derived annual run rate. Label every estimate as an estimate.
- Add a "baseline exists: yes/no" column to your Process Inventory Register. From today, that column is a gate: no process advances toward a pilot conversation while its column says no.
Key Takeaways
- Treat the baseline as the argument you win or lose in month zero: MIT's 95 percent no-measurable-return finding and McKinsey's 39 percent EBIT-attribution figure both describe organizations that could not prove value because they never captured the before.
- Build the Process Baseline Pack as three layers: the four core metrics, a measurement note (definition, window, source, blind spots) for every number, and a one-line data lineage that invites the reviewer to check your work.
- Measure cycle time as elapsed clock time from trigger to completion and report the median and the 90th percentile, because averages hide the queue-wait tail your map walkthrough already found.
- Segment volume by type so concentration is visible, like the 20 percent of exception types consuming 60 percent of effort, and so a mix shift can never be used to dismiss your pilot comparison.
- Define what counts as an error in writing, with the process owner's agreement, before measuring anything, or the definition fight happens in month six in front of the steering committee.
- Compute cost per unit the honest way: sampled or time-studied labor time, never the veteran's guess, times the finance-standard loaded rate, plus system costs, with every loading assumption stated.
- Certify every AI-extracted figure by tracing it to raw source and recomputing it on a sample, because a baseline with one wrong number manufactures fake ROI or buries real ROI, and it will be rechecked by the least friendly person in the room.
- Run a two-week tally-sheet sampling sprint where data does not exist, and label every estimate as an estimate: the documented gap survives hostile review, the confident unlabeled guess does not.
Skill.re