Catching Hallucinations and Drift in Live Workflows
On the Monday of week seventeen, the person who had watched the invoice-exception workflow every single day of its pilot cleaned out her desk and moved to the next project. Nobody panicked, because nothing was wrong. The workflow was live at three sites, the vendor-inquiry agent was answering its routine lane politely, and the dashboards were green. Four months later, at an operations review, a controller asked a small question: "Who is actually looking at the quality numbers now?" The answer around the table was the most dangerous phrase in production AI: "the team." Not a name. Not a calendar slot. A vapor. This lesson is about the machinery that replaces the vapor: the difference between a pilot, which had a full-time guardian, and production, which gets a rota, a schedule, and a set of monitors boring enough to run for years. Because quality decay in a live AI workflow is not an if. It is a when, and a how fast, and the only number you truly control is how many days pass between the decay starting and someone knowing.
The Week the Guardian Left
During the pilot, quality had a bodyguard. The transformation lead read the issue log daily, ran the canary set weekly, eyeballed the segment table on Fridays, and knew the workflow's moods the way a parent knows a toddler's cough. That intensity was correct for a six-week experiment, and it is completely unsustainable for a workflow that will run for years across three sites while the same person moves on to the next process. Production does not get vigilance. Production gets whatever discipline you designed before the vigilance walked out the door.
Here is the uncomfortable arithmetic of what happens next, drawn from the pattern the postmortem literature keeps repeating. In one anonymized case that circulates in operations circles (illustrative, but painfully typical), a document-processing lane degraded for nine weeks before anyone noticed, because the aggregate accuracy number stayed flat while one lane quietly collapsed inside it. In another, a classification model drifted for roughly eighteen months, discovered only when an audit sampled old outputs and found the last year's decisions resting on categories the business had stopped using. Months. The well-instrumented shop measures the same gap in days. And the difference between months and days is not that the second shop hired more careful people. The difference is regime: a designed, budgeted, deliberately boring set of monitors that run whether anyone is inspired or not.
Call the number what it is: detection latency, the days between quality starting to slip and a named human knowing it slipped. It is this chapter's founding KPI (key performance indicator, the number a process is managed by). Everything else in this lesson exists to shrink it. MIT's autopsy of the 95 percent of stalled GenAI pilots found missing learning loops at the center of failure: systems nobody watched, corrected, or improved. A production workflow without a monitoring regime is the same disease in a later, more expensive stage, because now the workflow touches real money at three sites and nobody is even assigned to notice. Gartner's warning that over 40 percent of agentic AI projects will be canceled by the end of 2027 is usually read as a launch statistic; read it instead as an operations statistic. Projects do not only die on the launchpad. They die in year two, when an unwatched workflow finally embarrasses someone important.
The artifact this lesson leaves in your portfolio is the Quality Watch Regime: your production monitoring stack on one page. Four monitors, their cadences, their owners, their alert thresholds, and the monthly quality review that consumes them. You already built the ancestors of every piece during the pilot chapter: sampling, canaries, segment tables, the escalation ladder. This lesson does not re-teach them. It promotes them: from a six-week sprint watched by one devoted human to a multi-year regime carried by a rota of part-time reviewers, across the invoice-exception workflow at three sites and the vendor-inquiry agent beside it.
A monitor without a cadence, an owner, and a consumer is decoration. Detection latency, measured in days, is the number that tells you whether you built a regime or a dashboard.
Monitor One: The Sampling Program, Sized and Scheduled
In the pilot, sampling was generous because attention was generous: the guardian pulled cases whenever curiosity struck, and curiosity struck daily. Production sampling has no curiosity to spend. It has a budget, so it must be sized: the arithmetic you learned in the verification lesson, now run backward. There you asked "given this sample, what can we claim?" Here you ask "given the error budget and the detection latency we will accept, how big must the weekly sample be?" Nothing heavier than multiplication is required.
Walk the reasoning with illustrative numbers for the invoice-exception workflow. The error budget says the workflow may carry a 2 percent defect rate; the alarm condition you care about most is a doubling, to 4 percent, and you want to detect that doubling within two weeks. The logic in plain words: a sample only reveals a change in the error rate if it catches enough errors for the change to stand out from ordinary luck. At a 2 percent error rate, a two-week sample of 100 items would show about 2 errors; a doubling would show about 4. Two versus four is a coin-flip difference, invisible in noise. But a two-week sample of 480 items shows roughly 10 errors at baseline and roughly 19 after a doubling, and ten versus nineteen is a difference a part-time reviewer can see without a statistics degree. So the sizing lands at about 240 sampled items per week: with three sites and two monitored lanes per site, that is roughly 40 items per site per lane per week, stratified so no site or lane hides inside an average, exactly the failure the nine-week lane story describes.
| Stratum (illustrative) | Weekly sample | Expected errors at 2% baseline | Expected at 4% (doubled) |
|---|---|---|---|
| Site A, routine lane / exception lane | 40 + 40 | ~1.6 | ~3.2 |
| Site B, routine lane / exception lane | 40 + 40 | ~1.6 | ~3.2 |
| Site C, routine lane / exception lane | 40 + 40 | ~1.6 | ~3.2 |
| Total per week / per two-week window | 240 / 480 | ~5 / ~10 | ~10 / ~19 |
The alert threshold falls out of the same table: if the rolling two-week window shows 16 or more errors (comfortably above the baseline's noise, comfortably below the doubled rate's expectation), the regime opens an investigation. Wanting faster detection or a smaller alarm condition costs a bigger sample; that is a budget conversation, not a virtue conversation, and the Quality Watch Regime page records the choice explicitly so nobody later pretends the coverage was infinite.
Then come the three production disciplines that separate a sampling program from sampling vibes. First, samples are pulled by schedule: a random draw generated every Monday by the system, not chosen by a human (humans unconsciously pick interesting cases, and drift lives in the boring ones). Second, every sampled item is reviewed by the rota and logged with a reason code: pass, extraction error, routing error, policy mismatch, judgment call, so the log can be read as data later. Third, and this is the production verse of the attendance-before-celebrating-silence hymn you learned at the gate: the sample never silently shrinks. A busy week does not quietly review 12 items instead of 40. A skipped or shortened week is a logged exception with a named owner and a catch-up plan, because the most common way monitoring dies is not a decision to stop; it is a slide from 240 to 180 to 90 to "when things calm down," with nobody ever having chosen it. Before you celebrate a quiet quality log, check attendance: was anyone actually sampling?
Monitor Two: The Canary Deck Gets a Lifecycle
The pilot's canary inputs were 20 fixed cases with known correct answers, run weekly to catch silent behavior change. Production keeps the idea and gives it what the pilot version lacked: lifecycle management, because a canary deck that never changes slowly stops resembling the work, and a canary deck that never grows learns nothing from its scars.
Three upgrades. First, the deck splits in two. The fast deck is about 10 items per workflow, curated to cover the highest-stakes behaviors (the largest-dollar extraction patterns, the trickiest routing boundary, the vendor-inquiry answers that touch refund policy), and it runs on every model, prompt, or configuration change, exactly the regression gate your rewritten SOP (standard operating procedure) listed among its change triggers. No change ships without the fast deck passing; it is the ten-minute seatbelt. The full deck is 50 or more items and runs monthly, deep enough to notice subtle reshaping the fast deck is too small to feel.
Second, the deck is refreshed quarterly and expanded after every incident. Quarterly refresh retires cases whose formats the business no longer sends and adds cases representing the newest vendors, templates, and phrasings, so the deck keeps tracking reality. And every postmortem contributes its trigger case to the deck permanently: the invoice that fooled extraction in March becomes canary 47 forever. Over two years the full deck becomes something no vendor can sell you: institutional scar tissue, a test set made entirely of the specific ways this workflow has actually been wrong.
Third, canary results are plotted over time, not just checked pass/fail. A deck that scores 50 of 50 for six months and then drifts to 48, 47, 46 across a quarter is not "still mostly passing." It is drawing you the drift curve with a pencil, one point per month, and a slow canary degradation is the single clearest early portrait of decay you will ever get, because the inputs are frozen: if the answers change, the system changed.
Monitor Three: The Distribution Watch
Sampling inspects individual cases; the distribution watch inspects the population, and it catches the class of problem samples are almost blind to: nothing is exactly wrong with any single case, but the shape of the whole flow has shifted. This is the monitor that would have ended the eighteen-month drift story in its first month.
On the input side, chart monthly the leading indicators your pilot lesson already named: the novelty rate (share of items unlike anything in the reference set: a rising line means the world is changing faster than your configuration), the format mix (which templates, channels, and document types arrive, and in what proportion), and volume by segment (the lane that quietly doubled is a risk event even if its accuracy holds, because your error budget was sized on the old mix).
On the output side, chart three things. The confidence-band mix: if the model that used to place 70 percent of items in its high-confidence band now places 55 percent there, the model has not become humble; something changed upstream, and this line often moves weeks before any error metric does. The category distribution: the share of items landing in each classification, watching for the category that quietly doubles or the one that fades (real behavior change in the system or in the business, either way worth knowing). And the sharpest instrument of the three: the override rate by reason code. Your human gate logs a reason code every time a reviewer overrides the AI, and that log is a drift seismograph. A rising "amount corrected" code means extraction is slipping. A rising "wrong queue" code means routing is. The gate you defended as a control turns out to be a sensor, which is one more reason its budget line is never the one to cut.
Now the operator skill: read these lines the way a production operator reads a control chart, not the way a nervous sponsor reads a stock ticker. Every line wiggles; a monitoring regime that treats every wiggle as an emergency burns out its rota in a quarter and teaches everyone to ignore alerts, which is worse than having none. The rule of thumb the regime writes down: establish each line's historical band (the range it has lived in, say the last six months' high and low), and investigate when a line crosses its band twice in a row. One excursion is weather; two consecutive is a pattern. And log the investigation either way: "checked, benign, cause was quarter-end volume" is a real entry that builds the regime's memory and calibrates the band. An investigation that finds nothing is not wasted; it is the price of the ones that find something.
Monitor Four: The Hallucination Tripwires
The first three monitors catch decay in quality. The fourth targets the failure mode this lesson's title names: fabrication, the confident invention of things that do not exist. You learned the rule in Level 2 as personal discipline: verify every AI-touched figure, step, and claim. Tripwires are that rule industrialized: automatic cross-checks that run on every item, cost nearly nothing, and never sleep. Four kinds:
- Reconciliation totals. Amounts must sum. The line items on an extracted invoice must add to its extracted total; the vendor-inquiry agent's quoted balance must match the ledger it cites. A model can hallucinate a plausible number; it very rarely hallucinates a whole set of numbers that reconciles. Arithmetic is the cheapest lie detector in operations.
- Referential checks. Every vendor, PO (purchase order), invoice number, contract clause, or document the AI cites must exist in the system of record, verified by an automatic lookup. This is the single cheapest hallucination detector there is: a fabricated reference fails a lookup, every time, in milliseconds. If your workflow lets an AI-cited identifier reach a decision without a lookup, you have chosen not to know.
- Impossible-value rules. Delivery dates in the future of the payment date, negative quantities, unit prices ten times the historical range, categories outside the codebook. Each rule is one line of logic and one recovered afternoon.
- The periodic citation audit. Tripwires verify that cited things exist; they cannot verify that a cited document actually says what the AI claims it says. So monthly, the rota opens a sample of evidence pointers (the paragraph the vendor-inquiry agent cited for a policy answer, the field the extraction cites for an amount) and reads them. This is Level 2's trained skepticism running on a schedule instead of on mood.
Two rules make tripwires part of the regime rather than a pile of scripts. Tripwire hits route into the same escalation ladder as any other error, with the same reason codes, so a spike is triaged rather than absorbed. And the tripwire hit rate is itself a monitored line on the distribution watch: a rising rate means something upstream is degrading, and a rate that falls to zero for a month usually means a tripwire silently stopped running, which is its own alarm. Gartner's persistent finding that organizations underinvest in AI risk controls relative to AI enthusiasm has a small, concrete refutation available to any operations team: four categories of automatic checks, none of which requires a data scientist.
The Review, the Rota, and the Day the Regime Fired
Monitors produce lines; someone must consume them, or the whole apparatus is a diary nobody reads. The regime's consumer is the monthly quality review: 60 minutes, all AI workflows in the division on one agenda, standing structure. It is the weekly pilot review relaxed to a steady-state cadence, with the same discipline it taught you: (1) the four monitors' dashboards read together, because drift usually shows as a pattern across monitors rather than a single alarm; (2) any escapes (errors that reached downstream) traced to the verification layer that should have caught them; (3) the drift-suspicion list triaged: every "crossed its band twice" investigation reviewed and closed or escalated; (4) canary deck maintenance: refresh due dates, incident cases added; and (5) one improvement decision per workflow, and only one, so causality stays readable when next month's numbers move. Sixty minutes, monthly, forever. Boring is the point; boring is what survives.
The rota
Quality duty rotates through trained reviewers: nobody's full-time job, everybody's scheduled turn. Rotation solves two failure modes at once: fatigue (fifty weeks of solo sampling dulls anyone's eye) and single-point knowledge (the shop where only one person can judge quality is one resignation away from the vapor). Who qualifies for the rota comes straight from the Level 2 skills matrix: reviewers who have demonstrated verification skill on this workflow, not whoever is free. And the rota is a budget line, written into the workflow's run cost and defended exactly like the gate's staffing, because it will be attacked exactly like the gate's staffing: quietly, in a busy quarter, by someone pointing at months of green dashboards that are green because the rota exists. In the division's illustrative regime below, the rota is four trained reviewers at roughly 3 hours per week each: about 12 hours weekly buying coverage of two workflows across three sites. Against BCG's 10-20-70 arithmetic (10 percent of AI success is algorithms, 20 percent technology and data, 70 percent people and process), this is what a slice of the 70 looks like in steady state: not a program, a rota.
When the regime fires: a six-day story
Here is the regime working, in an illustrative composite. Week 1, Tuesday: the distribution watch's monthly refresh shows Site B's exception lane override rate crossing its historical band, second consecutive reading, reason code "amount corrected" driving the rise. The rule fires; an investigation opens with a named owner from the rota. Per the escalation ladder, Site B's sampling doubles for the week (80 items instead of 40: the ladder's first rung, pre-committed months ago, no meeting required). By Friday the doubled sample has isolated the pattern: quantity extraction errors clustered on receipts from one warehouse team. Monday, a reviewer reads twenty raw inputs and finds it: the local team had adopted a new shorthand in a free-text field ("recv dock 3 / 2 short, adj per JL"), and the extraction was reading the shorthand's numbers as quantities. The fix: an input-handling rule for the field, one retraining note to the warehouse team, and the shorthand case added to the full canary deck as permanent scar tissue. Detection latency: 6 days from first band crossing to root cause in hand. That number is reported at the next quality review with genuine pride, because latency is the regime's KPI: the thing you measure about your measuring. Note what did not happen: no customer saw it, no month-end close absorbed it, and nobody needed to be a hero.
The Quality Watch Regime, on its one page
The division's regime page, all numbers illustrative:
| Monitor | Cadence | Owner | Alert threshold |
|---|---|---|---|
| Sampling program (invoice-exception: 240/wk stratified; vendor-inquiry: 60 transcripts/wk) | Weekly, scheduled draw | Rota (4 reviewers, ~3 hrs/wk each) | 16+ errors in rolling two-week window; any skipped sample logged as exception |
| Canary decks (fast 10-12 / full 50-60 per workflow) | Fast: every change. Full: monthly. Refresh: quarterly | Workflow owner | Any fast-deck failure blocks the change; full-deck score below 94% opens investigation |
| Distribution watch (novelty, format mix, volume, confidence bands, category mix, overrides by reason code) | Monthly chart refresh | Rota lead of the month | Any line outside its historical band twice in a row |
| Hallucination tripwires (reconciliation, referential lookups, impossible values, citation audit) | Continuous; citation audit monthly | Automatic; hits to escalation ladder | Hit rate above 1% or at 0% for a month (both are alarms) |
| Monthly quality review (consumer of all four) | Monthly, 60 min, standing agenda | Division process owner | One improvement decision per workflow per month |
And the page's bottom strip, last quarter's results: 3 investigations opened (2 closed benign and logged, 1 real drift caught at 6 days latency); tripwire hit rate steady at 0.4 percent, with referential checks catching supplier typos far more often than model fabrications, and here the regime page carries an honest note: most tripwire catches are data quality, not hallucination, and that is still value: a wrong vendor ID caught in milliseconds is a misdirected payment that never happened, whoever authored the error. Sampling: 100 percent of plan across all 13 weeks, two exceptions logged (a holiday week rescheduled, a reviewer absence covered). That last line is the quiet triumph. Anyone can sample in week one. The regime is whoever is still sampling in week thirteen.
Monitors Without a Regime Are Decoration
Now the mirror, assembled from real failure patterns into one illustrative firm. A distribution company of similar size launched a similar invoice workflow with, on paper, everything this lesson teaches. Monitoring appeared on the launch checklist and was marked complete. Ownership was assigned to "the team," which is to say to nobody. There was no rota, no named consumer, no exception log for skipped samples.
The slide was gradual and completely undramatic. Weekly sampling held for two months on launch enthusiasm, then slid to "when workload allows" by month 4; nobody decided this, so nobody could be asked about it. In month 7, the model vendor shipped a silent update: the contract had no change-notification clause, because the due-diligence question you learned to ask in Level 1 ("How will we know when you change the model?") had never been asked, and in month 7 the invoice for skipping it arrived. Extraction behavior subtly reshaped: a category boundary moved, and a few percent of exceptions began landing in a lane where they were auto-approved under a threshold meant for a different mix. The distribution watch that would have caught the category shift within one monthly refresh existed, technically: a dashboard, built at launch, that no calendar ever compelled anyone to open. Discovery finally arrived in month 11, from outside: a customer dispute over a pattern of incorrect charges, which is the most expensive monitoring channel there is. The cleanup ran an illustrative $220,000 in credits, reprocessing, and audit hours, plus a quarter of the team's attention, plus the intangible that costs most: the next AI proposal at that firm now gets negotiated against this memory.
The postmortem's finding deserves quoting because it is the lesson's closing distinction: every individual monitor existed on paper, and no regime existed. There was a sampling plan with no schedule, canaries with no runner, a dashboard with no reader, tripwires half-wired and unwatched. Detection latency: roughly four months from the vendor update to discovery, and the discovering instrument was a customer. Monitors without cadence, owner, and consumer are decoration. The regime is the cadence, the owner, and the consumer; the monitors are merely what they operate.
One bridge before you go. Quality decay is one production risk, loud in its consequences even when slow in its arrival. There is a second risk that is quieter in every way: a workflow can hold its error budget perfectly and still distribute its errors, its scrutiny, or its service unevenly across the people it touches. Systematic unfairness does not trip a reconciliation total. The next lesson gives it its own monitors.
What to Do Monday Morning
- Size your production sample from the error budget backward. Take one live workflow's error budget, decide the change you must detect (a doubling is the honest default) and the latency you will accept, and multiply out the weekly stratified sample as this lesson did: enough expected errors at baseline that the doubled rate stands clear of noise. Write the number, the reasoning, and the review-hours cost on one page.
- Split your canaries into a fast deck and a full deck, and date the refresh. Ten items that run on every change; fifty or more that run monthly; a quarterly refresh reminder on a real calendar; a standing rule that every incident donates its trigger case to the deck.
- Chart last quarter's override rate by reason code and read it like a control chart. Draw each code's line, mark its historical band, and apply the rule: twice outside the band in a row means investigate, and log the investigation even when it closes benign.
- Wire referential checks on every AI-cited identifier. Vendor IDs, PO numbers, document references, policy clauses: every one the AI cites gets an automatic existence lookup before it reaches a decision. It is the cheapest hallucination detector in operations; there is no defensible reason to run without it.
- Put the monthly quality review on the calendar with its rota named. Sixty minutes, standing agenda, all workflows, one improvement decision each. Name the four rota reviewers from the skills matrix, book their hours as a budget line, and report detection latency as the regime's headline KPI at the first session.
Key Takeaways
- Treat quality decay in live AI workflows as a when, not an if, and manage the one number you control: detection latency, the days between quality slipping and a named human knowing.
- Replace the pilot's full-time guardian with a regime: four monitors with cadences, owners, alert thresholds, and a monthly review that consumes them, captured on the one-page Quality Watch Regime.
- Size the production sampling program backward from the error budget and target latency, pull samples by schedule with reason-coded review, and log every skipped week as an owned exception so the sample never silently shrinks.
- Manage the canary deck as a living asset: a fast deck gating every change, a full deck run monthly, quarterly refreshes, and every incident's trigger case added permanently as institutional scar tissue, with scores plotted over time.
- Watch distributions, not just cases: novelty rate, format mix, volume by segment, confidence-band mix, category distribution, and override rate by reason code, investigating when a line crosses its historical band twice in a row and logging the outcome either way.
- Automate fabrication defense with tripwires: reconciliation totals, referential existence checks on every AI-cited identifier, impossible-value rules, and a monthly citation audit, with the hit rate itself charted as a monitored line.
- Defend the rota as a budget line: trained reviewers from the skills matrix rotating through quality duty a few hours a week, preventing both fatigue and single-point knowledge.
- Remember the closing distinction: monitors without cadence, owner, and consumer are decoration, and a firm can own every dashboard on paper while its real monitoring instrument turns out to be a customer complaint in month eleven.
Skill.re