←
AI Readiness & Process Transformation
Proficient · M16 · lesson 16 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Running the Pilot: Cadence, Logs, and Drift Watch

15 min

Week three of the invoice-exception pilot, Tuesday, 4:40 in the afternoon. The launch meeting was three weeks ago and it was great: the sponsor spoke, the demo landed, someone brought pastries. Now the sponsor is traveling, the issues list has fourteen open items nobody has ranked, the analyst who was supposed to pull the weekly numbers got borrowed for a quarter-close fire drill, and the sampling review that was "definitely happening Thursday" has quietly become "next week for sure." Nothing dramatic has occurred. No outage, no scandal, no angry email. And that is exactly the problem, because pilots do not fail in the launch meeting. They fail here, in week three, silently, when the adrenaline is gone and nobody owns Tuesday. Three months from now, someone will ask what the pilot proved, and the honest answer will be: we are not sure, we stopped writing things down around day nineteen. This lesson is about the dull, repeating, calendar-anchored rhythm that prevents that answer, and it may be the least glamorous and most valuable thing in this entire level.

The Quiet Death of Week Three

You have done the hard design work. The previous lesson gave you a pilot with a frozen scope, a captured baseline, and pre-committed success and kill criteria. Chapter 3.2 gave you a redesigned workflow with a human gate, a Verification Architecture with layered checks, an escalation ladder with pre-negotiated yellow, orange, and red rungs, and an override log at the gate. The instrumentation lesson gave you counters and a weekly one-pager. On paper, this pilot cannot fail quietly.

Except it can, because paper does not run pilots. Calendars do. Every artifact you built is a machine that only produces value when someone turns the crank on schedule, and the forces working against the crank are not villains. They are ordinary gravity: a busy week, a traveling sponsor, a reallocated analyst, a review that slips once and then slips forever. MIT's autopsy of the 95 percent of GenAI pilots with no measurable return found that the missing learning loop was a defining failure: tools and teams that never systematically captured what was going wrong and fed it back. The learning loop is not a philosophy. It is a meeting that happens every week whether or not anyone feels like it. When Gartner found that 30 percent of generative AI projects were abandoned after proof of concept, the abandonment mostly did not follow a catastrophic event. It followed a long silence during which nobody could say what the pilot had actually done.

So here is the thesis, stated plainly. What separates a pilot that produces evidence from a pilot that produces anecdotes is not talent, not the model, not the vendor. It is an operating rhythm: the same review, the same log, the same sampling, the same questions, every single week, at the same hour, with named owners. Discipline is the antidote to drama. A six-week pilot run on rhythm ends with six weeks of dated, comparable records. A six-week pilot run on enthusiasm ends with a folder of screenshots and a deck written from memory.

The artifact you take from this lesson is the Pilot Operating Rhythm: a one-page week template that names the four standing mechanisms, their owners, their agendas, and their time budgets. Here it is in summary form; the rest of the lesson teaches each row slowly.

MechanismCadenceTime budgetOwnerOutput
Weekly reviewSame day, same hour, every week45 minutesPilot lead (you)Decisions logged, the week's one change chosen
Issue logContinuous, triaged before the review10 minutes dailyPilot leadOne row per problem, closed only with evidence
Output samplingFixed weekly appointment60 minutesNamed sampling ownerStratified sample scored, results into the one-pager
Drift watchWeekly, produced before the review30 minutesInstrumentation ownerSegment table, canary results, novelty count, format check

On top of the four mechanisms sits the two-speed clock. In weeks one and two, the pilot runs on a daily beat: a 10-minute floor check at the gate every morning, because the first days surface things no log will ever show you. From week three onward, the pilot shifts to the weekly steady state. The transition is planned, not accidental: you stop going daily because the daily checks stopped finding new things, not because you got busy. Write both speeds into the calendar before launch, because after launch the calendar fills itself.

Cadence is not bureaucracy. Cadence is how a pilot remembers what happened, and a pilot that cannot remember cannot prove anything.

The Weekly Review: The Same 45 Minutes, Every Week

The weekly review is the heartbeat, and it is deliberately boring. Same day, same time, same room or call link, same fixed agenda, 45 minutes, hard stop. Attendees: the pilot lead, the team lead whose people work the gate, the instrumentation owner, and whoever owns the AI tool relationship. The sponsor attends when a decision on the ladder requires them; otherwise they get the one-pager. The agenda has five items, and the order matters.

1. Counters versus budget (10 minutes)

The weekly one-pager from the instrumentation lesson gets read together, out loud, line by line: volume in, volume routed to the gate, verified error rate against the error budget, cycle time against baseline, override rate. Reading it together is not ceremony. Numbers nobody reads together get read alone, and numbers read alone get spun: the optimist remembers the good week, the skeptic remembers the bad one, and by month two the team is running on two incompatible private narratives. Ten minutes of shared reading produces one shared record. When the number is uncomfortable, it gets read anyway, in the same flat voice as the comfortable ones. That habit, more than any single metric, is what makes the eventual scale-or-kill decision credible.

2. Override-cluster review (10 minutes)

The human gate's override log captures every case where a reviewer rejected or corrected the AI's output, with a reason code. Weekly, those codes get aggregated and the question gets asked: where is the model weak this week? Not in general, this week. Three overrides on freight-charge misreads from one vendor is a pattern; a pattern is a candidate for the week's one improvement. This is the MIT learning loop with a drumbeat attached: the 5 percent of pilots that produced measurable value were the ones where correction flowed somewhere and changed something. Your override log is only a learning loop if someone reads it on a schedule. Otherwise it is a diary.

3. Issue log walk (10 minutes)

Not the whole log. Only items that moved since last week or have aged past their severity's time limit. Everything else is noise, and a review that reads all fourteen open items every week trains people to stop listening. The walk exists to catch two things: fixes claiming to be done without evidence (more on that shortly) and old items quietly rotting under a mid-tier severity label.

4. Decisions needed (10 minutes)

If a metric is approaching a rung on the escalation ladder, the consult happens here, in daylight, against the pre-negotiated thresholds. The ladder was built precisely so this conversation is short: the trigger levels and responses were agreed before launch, so the review is checking a reading against a gauge, not renegotiating physics under stress. The only time the ladder gets pulled outside this meeting is when a trigger actually fires mid-week; then it fires per its own protocol, and the review does the post-mortem.

5. The week's one improvement (5 minutes)

Every week, the review picks exactly one change to the pilot: a prompt adjustment, a threshold move, a screen fix, a new rule. One. It gets logged with a date and an expected effect. Here is why the number is one, taught with the confounding example that makes it stick. Suppose in the same week you raise the confidence threshold for auto-release and rewrite the extraction prompt for freight charges. Next week the error rate drops from 1.8 to 1.3 percent. Which change did it? You cannot know. Worse: suppose the threshold change improved things by 0.7 points and the prompt change made them worse by 0.2. The combined number looks like success, you keep both changes, and you have just permanently installed a defect you paid to discover. One change per week is not slowness; it is experimental hygiene, the only way a six-week pilot can attribute effects to causes. The other failure direction matters equally: zero changes per week means the learning loop is dead and you are running a demo, not a pilot. More than one and you cannot attribute; zero and you are not learning. Exactly one, logged, every week.

The Issue Log and the Sampling Hour

The next two mechanisms are the pilot's memory and the pilot's eyesight, and they fail in opposite ways: the log fails by lying gently, the sampling fails by simply not happening. Take them in turn.

The issue log: the pilot's memory

The issue log's job description is one sentence: the issue log is the pilot's memory. Six weeks from now, the decision meeting will ask what went wrong, when, how bad it was, and whether it was truly fixed. The team's recollection will be warm mush by then. The log will be exact, if and only if it was kept exactly.

The format is deliberately minimal, one row per problem: date found, found-by (which verification layer caught it: the schema validator, the reconciliation check, the human gate, the sampling review, or, worst of all, a downstream customer), description, severity, owner, status, and closure evidence. The found-by column earns its place because it tells you whether your Verification Architecture is working as designed: issues should be getting caught by the cheap early layers, and any issue found downstream of the gate is automatically a second issue about the layer that should have caught it.

The rule that gives the log its teeth is close-with-evidence: an issue closes when its fix is verified by the same layer that found the problem, not when someone says it should be fixed now. That phrase, "should be fixed now," is the most expensive sentence in pilot operations. It is how a week-2 extraction bug gets marked resolved, drifts out of everyone's attention, and resurfaces in week 5 inside the evidence pack, at which point you no longer have a bug, you have a credibility problem. Close-with-evidence is mechanical: if the sampling review found the issue, the next sampling review must show it gone before the row turns green. If the reconciliation counter found it, the counter must run clean for a defined window. The closure evidence cell contains a pointer to that proof, not an adjective.

The log also carries a meta-signal worth reading monthly, and it is subtle enough to teach explicitly. Watch two rates together: the discovery rate (new issues opened per week) and the escape rate (issues found downstream of the gate, by customers or by the ledger). If discovery is falling while escapes stay flat and low, the pilot is maturing: you are finding fewer problems because there are fewer problems. But if discovery is falling and, in the same weeks, the sampling hour got skipped twice, you have not matured. You have stopped looking. Those two conditions produce identical charts and opposite realities, and the only way to tell quiet from blind is to check attendance before you celebrate silence. A falling issue count is good news only if the finding mechanisms ran at full fidelity while it fell.

Output sampling: the standing appointment

The Verification Architecture lesson designed your stratified sampling plan: which segments get pulled, at what rates, scored against what standard. That was Chapter 3.2's job, and it is done. This lesson's job is narrower and harder: execution fidelity. The plan only produces evidence if it runs every week, on schedule, at full size, and sampling is, empirically, the first thing a busy week skips. It has no angry stakeholder attached. Nothing visibly breaks when you skip it. The invoices still flow, the gate still runs, the counters still count. The only thing that dies when sampling skips is the pilot's evidence quality, and it dies invisibly, exactly where you can least afford it: in the verified error rate, the single number the scale-or-kill decision leans on hardest.

So the operating rule is blunt. The sampling hour goes on the calendar as a recurring appointment with a named owner, the same way payroll processing does, and a skipped week gets logged in the issue log as a defect with a severity, because that is literally what it is: missing evidence is a defect in the pilot's product, and the pilot's product is evidence. Writing "week 4 sample skipped, quarter-close conflict, severity 2, owner reassigning hour" feels pedantic the first time. It is also the difference between a decision meeting that says "verified error 1.4 percent across six full weekly samples" and one that says "around 1.4 percent, though two of the six weeks are, um, interpolated." One of those sentences survives a skeptical CFO. The other does not, and you already know which.

Two practical notes on keeping the appointment alive. First, protect the owner, not just the hour: name a backup sampler in week zero so a vacation does not become a gap. Second, feed the results forward the same day: the sample scores go into the one-pager and any new finding opens an issue row before the weekly review, so the review always works from this week's evidence, not last week's.

The Drift Watch: The Mechanism Unique to AI

The first three mechanisms would improve any process pilot. The fourth exists because this is an AI pilot, and it needs to be taught carefully, because nothing in a traditional operations career prepares you for it.

Mechanical systems degrade loudly. A conveyor bearing whines for weeks before it seizes; a worn die produces visibly rougher parts. Model performance decays differently: silently, with nobody touching anything. This is drift, the word you met in the Level 1 terminology lesson, now made operational. The model has not changed. The world underneath it has. A vendor updates their invoice portal and the layout your extraction was tuned on quietly changes. A new exception type appears because procurement signed a new supplier category. Volumes shift seasonally and the mix of easy and hard cases moves. The upstream ERP (Enterprise Resource Planning system, the system of record the pilot reads from) gets patched and a field format shifts by one character. Each of these degrades accuracy without a single error message, because from the model's point of view nothing failed; it is confidently applying old patterns to a world that no longer matches them. Drift is silent by nature, which means it will never announce itself. It must be watched, by design, on a schedule. That is the drift watch: a standing 30-minute weekly production of four instruments, delivered to the weekly review.

Instrument 1: the accuracy-by-segment table

Aggregate accuracy is a blanket that hides bodies. Overall accuracy can hold rock steady while one segment collapses, because the collapsing segment is small or because an improving segment offsets it. Your Failure Mode and Effects Analysis (FMEA, the pre-launch exercise that ranked failure modes by severity, likelihood, and detectability) flagged this as the lane-blindness risk; the drift watch turns that row into a weekly practice. The table is simple: verified accuracy by vendor group, by exception type, by amount band, week over week. You are not reading it for the level; you are reading it for movement. Any segment down materially two weeks running gets an issue row, even if the aggregate looks fine. Especially if the aggregate looks fine.

Instrument 2: canary inputs

Now define the term properly, because it is the cheapest drift detector in existence and most teams have never heard of it. Canary inputs are a fixed set of roughly 20 known-answer items, real historical cases whose correct outputs you have verified by hand, run through the live system every week. The power is in the word fixed: because the inputs never change, any change in the outputs is a change in the system, not a change in the case mix. Your live accuracy number confounds the two; a dip might mean the model got worse or might mean this week's invoices were harder. Canaries cannot be confounded. Same 20 invoices in, different answers out means something moved: the model version, a prompt, an upstream format, an integration. Twenty items take minutes to run and score, and they convert "accuracy feels lower lately" into "canary 14 changed its answer on Tuesday." One maintenance rule: refresh the canary set quarterly, retiring stale cases and adding current ones, so the set stays representative of the work the pilot actually sees.

Instrument 3: the novelty counter

Track the rate of never-seen-before combinations entering the pilot: new vendor plus format pairs, new exception subtypes, new currency or entity codes. Rising novelty predicts drift before accuracy shows it, and the distinction is worth one paragraph because it is the difference between leading and lagging indicators. Accuracy is a lagging indicator: by the time it falls, the unfamiliar inputs have already been processed, some wrongly. Novelty is a leading indicator: the unfamiliar inputs announce themselves on arrival, before their errors land. A novelty rate climbing from 2 percent of volume to 7 percent is the model telling you, in the only language it has, that the world is moving away from its training. You get to respond while the error is still in front of the gate instead of behind it.

Instrument 4: the input-format monitor

Go back to your FMEA and find the failure mode with the worst detectability score, the one the exercise concluded you would learn about last. For most document pilots it is an upstream format change: silent, systemic, and invisible until outputs rot. That row becomes a standing weekly check: field-population rates, layout fingerprints, character-length distributions on key fields, whatever cheap signature detects "the input changed shape." You are instrumenting the exact blind spot your own analysis identified, which is the FMEA doing what it was always for.

Finally, drift response: when the watch flags something, the response routes through the escalation ladder, starting at yellow. No heroics, no hotfix at midnight, no freelance threshold changes. Yellow exists precisely for this: widen the sample, tighten the gate on the affected segment, investigate with a deadline. The pre-committed rungs were negotiated in calm weather for exactly this moment of weather. Gartner's finding that organizations pairing AI deployment with explicit risk controls see materially better outcomes is this paragraph in institutional form: detection without a pre-agreed response path just produces better-informed panic.

Weeks One and Two: The Daily Regime

The steady-state rhythm above starts in week three. The first two weeks run hotter, on the second speed of the two-speed clock, and the extra investment is two mechanisms.

First, the daily 10-minute floor check: you (the transformer) and the team lead, physically or virtually at the gate, watching live reviews for ten minutes every morning. The goal is to see the first 50 gate reviews with human eyes, because the logs cannot tell you what the first weeks most need you to know. Logs record decisions; they do not record hesitation. They do not show the reviewer squinting at a cramped screen region, or scrolling past the evidence pointers because they load below the fold, or the workaround being invented in real time (the reviewer keeping a private spreadsheet because the tool's history view is two clicks too deep). This is the ethnography hour, and it is the cheapest redesign insurance you will ever buy: a friction you catch in week one is a screen fix; the same friction discovered in week five is a retraining program plus five weeks of subtly contaminated data.

Second, the week-2 calibration check, which the handoff lesson already scheduled as a named task: validate the model's confidence bands against live volume, because the vendor's stated calibration was measured on their data, not yours. If the tool's "95 percent confident" items are landing at 91 percent on your invoices, your routing thresholds are misrouting cases into the light-touch lane every day. The check is straightforward: bucket two weeks of live cases by stated confidence, compare stated to observed accuracy per bucket, and adjust the routing thresholds accordingly, once. That adjustment is logged as the week's one change, per the rule, so its effect is attributable in week three's numbers.

Six Weeks of the Invoice-Exception Pilot

Here is the rhythm doing its job, week by week, in the running example. Every number is illustrative, but the shape of the six weeks is the shape you should expect.

Week 1. The daily floor checks pay for themselves by Thursday. The logs show healthy gate throughput and a 4 percent override rate, all fine. What the logs do not show, and the floor check does: reviewers are approving from the AI's summary paragraph without opening the evidence pointers, because the pointers render below the fold. They are not being lazy; the screen is teaching them to skip verification. The fix, evidence made default-visible above the approve button, is logged as week 1's one change. Had this run unobserved for six weeks, the pilot's headline claim ("human-verified outputs") would have been quietly false the entire time.

Week 2. The calibration check runs on 900 live cases. The tool's high-confidence band is optimistic by about 3 points on this invoice mix, so the auto-release threshold moves up 3 points: week 2's one change, logged with its expected cost. Routing to the gate rises from 22 to 26 percent of volume, which costs cycle time, and the review accepts the trade explicitly: error-budget headroom now is worth more than speed now, and the decision is written down with its reasoning.

Week 3. The classic trough, arriving on schedule. Launch energy is gone, the issue log peaks at 17 open rows, and the verified error rate touches 1.9 percent against the 2.0 budget. The yellow rung is consulted at the weekly review: per the pre-negotiated threshold it has not fired, but sampling is widened on the two worst segments per the yellow protocol's cheap first step. The sponsor, traveling, gets the honest one-pager anyway, 1.9 and all, with the widened-sampling response noted beside it. That is a credibility deposit: in week 6, when the numbers are good, the sponsor will believe them, because they saw the bad week reported in the same flat format.

Week 4. Month-end arrives and volume spikes 40 percent with a harder mix, exactly as the pilot design predicted, because the baseline period included a month-end and the scope was frozen around it. The segment table holds within budget across the spike. One genuinely new exception subtype (a consignment-stock reconciliation case) appears, trips the novelty counter, and routes to the gate instead of auto-release, precisely as the FMEA-derived routing rule intended. The catch is undramatic, which is the point: the drama was spent at design time.

Week 5. The canaries earn their keep. Two of the 20 canary invoices, both from the same vendor, change their extraction outputs between Monday and the prior week with no model or prompt change on record. The input-format monitor confirms it: that vendor's portal silently updated its invoice layout over the weekend. Live extraction accuracy on that vendor's segment is down 11 points; aggregate accuracy barely moved, which is exactly why the segment table and canaries exist. Yellow rung, this time fired: the vendor's invoices route 100 percent to the gate, the vendor is notified, a temporary extraction rule ships as the week's one change, and the issue closes with evidence (canaries green again, segment accuracy recovered) in 4 days. Total drift-to-detection time: under one week. The industry norm, absent a watch, is discovering this in month three inside an angry reconciliation meeting.

Week 6. Steady state. Verified error rate 1.4 percent against the 2.0 budget, median cycle time 1.8 days against a 3.1-day baseline, override clusters shrinking, discovery rate falling while sampling attendance is perfect, so the quiet is real. The evidence pack for the decision meeting, already on the calendar, is not being written; it has been accumulating, one honest week at a time, since day one.

The pilot with no rhythm

Now the failure story, and it is uncomfortably common. A team at another company, same quarter, comparable technology, launches an invoice-automation pilot with a strong kickoff. Then: no standing review, no issue log, no sampling schedule, no drift watch. Just usage, and silence. Month three, the sponsor requests a status. The deck gets written from memory in one afternoon: anecdotes, screenshots, one before-and-after example chosen because it looked good. The sponsor, no fool, asks whether the model's accuracy has changed since launch, and whether anything upstream changed. Nobody knows. Nobody can know: nothing was logged, sampled, or watched, so the information does not exist anywhere in the organization. The pilot is not killed, because there is no evidence it failed. It is not scaled, because there is no evidence it worked. It is extended, which is the fate worse than death: more months of burn producing the same non-evidence, the zombie state you learned to fear in Level 1, the exact silhouette of MIT's 95 percent. And the team's real lesson, the one they will not put on a slide, is that they cannot tell whether their own pilot works. Both teams spent six weeks. One bought evidence. The other bought an extension.

What to Do Monday Morning

If your pilot launches soon, or launched already, build the rhythm this week.

  1. Calendar the weekly review before anything else: same day, same hour, recurring, 45 minutes, with the five-item agenda pasted into the invite (counters versus budget, override clusters, issue log walk, decisions needed, the week's one change).
  2. Open the issue log today with the seven columns (date, found-by, description, severity, owner, status, closure evidence) and write close-with-evidence into it as policy, in words, at the top of the sheet.
  3. Appoint the sampling owner and a named backup, put the weekly sampling hour on both calendars, and agree in advance that a skipped week is logged as a defect with a severity.
  4. Build your 20 canary inputs from known-answer historical cases spanning your main segments, verify their correct outputs by hand, and schedule the weekly canary run inside the drift-watch half hour, with a quarterly refresh reminder.
  5. Stand up the other three drift instruments: the accuracy-by-segment table, the novelty counter, and the input-format check on your FMEA's worst-detectability row, all feeding the weekly review.
  6. Book the week-1 and week-2 floor-check slots now, ten minutes every morning with the team lead, plus the week-2 calibration check as a named task, before launch week fills your calendar for you.

Key Takeaways

  • Run the pilot on the Pilot Operating Rhythm: four standing mechanisms (weekly review, issue log, output sampling, drift watch) with named owners, fixed time budgets, and a two-speed clock that goes daily in weeks 1 and 2, weekly after.
  • Hold the weekly review to its fixed five-item agenda and read the one-pager together, because numbers nobody reads together get read alone and spun into private narratives.
  • Enforce the one-change rule as experimental hygiene: exactly one logged change per week, since two simultaneous changes make effects unattributable and zero changes mean the learning loop is dead.
  • Close issues only with evidence from the verification layer that found them, because "should be fixed now" is the phrase that reopens in week 5 what was declared fixed in week 2.
  • Read the issue log's meta-signal correctly: falling discovery with steady sampling attendance means maturing, while falling discovery with skipped sampling means you stopped looking; check attendance before celebrating silence.
  • Treat a skipped sampling week as a logged defect with a severity, because the pilot's product is evidence and missing evidence is a defect in that product.
  • Watch for drift by design with four instruments (segment table, roughly 20 fixed canary inputs, novelty counter, input-format monitor), since drift is silent by nature and aggregate accuracy can hold while one segment collapses.
  • Route every drift response through the pre-committed escalation ladder starting at yellow, and remember the alternative fate: the rhythm-free pilot is neither killed nor scaled but extended, burning months to produce non-evidence.