←
AI Readiness & Process Transformation
Capable · M23 · lesson 23 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

The Verification Habit: Trust Nothing You Didn't Check

15 min

The readout deck is on slide fourteen when the operations director stops the meeting. She is looking at a quote attributed to her, in italics, with her name under it: "We stopped trusting the escalation queue in 2023." She never said it. She said something adjacent, something about response times, in an interview six weeks ago, and an AI synthesis tool smoothed her actual words into a punchier sentence that happens to be politically radioactive, because the executive who owns the escalation queue is sitting three chairs away. The assessor who built the deck used AI to synthesize twenty interview transcripts, which was smart, and shipped the output without checking the quotes, which was fatal. The project survives the meeting. The assessor's credibility does not. Nobody in that room will ever again read one of his deliverables without wondering what else was never said. The previous lesson taught you to see bad AI output when it is in front of you. This lesson installs the thing that separates professionals from casualties: a verification workflow with named steps, sampling rates, and a paper trail, so that "we used AI for this" becomes a defensible sentence instead of a confession.

The Job Is Not the Draft Anymore

Here is the quiet shift that happened the day AI entered assessment work, and most people have still not said it out loud. Your job used to be producing the draft: reading the transcripts, writing the synthesis, building the current-state summary, assembling the findings deck. That work took the hours it took, and the hours themselves were a kind of quality control, because you touched every sentence on its way in. When AI drafts in minutes what took you days, the drafting hours disappear, and so does the incidental checking that lived inside them. The work did not get smaller. It changed shape. The job is no longer "produce the draft." The job is "verify the draft," and the professionals who internalize that shift are the ones whose AI-assisted work survives contact with a steering committee, an external auditor, or a hostile Chief Financial Officer looking for a reason to distrust the whole program.

Notice what that sentence does not say. It does not say "re-do the draft." If verification means reading every transcript again and rebuilding the synthesis by hand to check the AI's version, you have spent the time savings entirely and added a coordination step; you would have been better off working manually. That is the trap on one side. The trap on the other side is the assessor in the opening scene: ship it raw, because it reads well and the deadline is Tuesday. Between those two failure modes sits the actual skill of this lesson: verification as a designed control, sized to risk, with a defined cost, the same way a factory does not inspect every bolt and does not inspect zero bolts. It inspects a designed sample, at a designed rate, tied to what a defect would cost downstream.

This is not a compliance nicety bolted on for the nervous. It is the load-bearing wall of the entire case for using AI in readiness work at all. Recall the number this program opened with: MIT's 2025 GenAI Divide research found that 95 percent of enterprise GenAI pilots delivered no measurable profit-and-loss return, and one of its three root causes was the missing learning loop, systems and workflows that never got better because nobody captured what went wrong. A verification workflow is a learning loop you build for yourself. Every error you catch and log teaches you where this tool fails, on your documents, in your domain. Skip verification and you are not just risking one bad deck; you are running an AI-assisted practice that never learns, which is the exact autopsy finding of the 95 percent.

One more framing before the mechanics, because it determines whether you will actually do this under deadline pressure. Verification is not a tax on the time you saved. It is the product. A synthesis that took five hours instead of sixteen, and that carries evidence of having been checked, is worth more than the sixteen-hour manual version, because it arrives faster and it arrives auditable. The manual version's quality claim is "trust me, I read everything." The verified AI version's quality claim is "here is what was machine-drafted, here is exactly how it was checked, here is what we found and fixed, and here is who signed off." In front of a steering committee, the second sentence wins every time.

Verification Is a Control, Not a Redo

Operations professionals already know the discipline this lesson is importing; it just lives in a different building. When a factory receives a shipment of ten thousand components, nobody inspects all ten thousand and nobody waves the truck through unopened. The receiving team pulls a sample whose size is set by an Acceptable Quality Limit, or AQL: a pre-agreed threshold that says, in effect, "we will tolerate at most this defect rate, and here is the sample size that lets us detect a worse rate with confidence." If the sample passes, the lot is accepted. If the sample fails, the entire lot is rejected and goes back to the supplier, not just the specific defective pieces found, because a bad sample impeaches everything you did not look at. Incoming inspection is cheap precisely because it is designed: sized to risk, statistically honest, and binary in its consequence.

Treat every AI output as an incoming shipment from a supplier with a known quality problem. The supplier is fast, tireless, cheap, and periodically inserts a fabricated component that looks identical to a real one. You would never accept that supplier's lots uninspected, and you would never bankrupt yourself inspecting every unit. You would design a control. That control, for AI-assisted assessment work, has three passes, and each pass exists because it catches a class of error the other two structurally cannot.

Trust nothing you didn't check, check nothing twice, and write down what you checked: that is the whole habit.

Before we walk through the passes slowly, hold the shape of the whole system in one view. The Source Pass catches fabrication: claims, quotes, and citations that point at nothing. The Spot-Audit Sample catches drift at volume: the error rate hiding inside four hundred processed tickets you cannot re-read. The Expert Walkthrough catches what only a resident of the process can see: the plausible-sounding step that does not exist, the critical exception the AI never mentioned because no document ever mentioned it either. Fabrication, drift, and omission. Three failure modes, three passes, and a one-line log entry at the end that turns all of it into evidence.

The Three Passes, Taught Slowly

Pass One: The Source Pass (Every Citation Checked)

In the previous lesson you learned to demand cite-your-source outputs: prompts that force the AI to label every claim with where it came from, "Transcript 7, paragraph 12" or "SOP-114, step 4," where an SOP is a Standard Operating Procedure, the written instruction set for how a task is done. The Source Pass is what you do with those labels. You walk each substantive claim in the output back to its cited source and confirm the source actually says it. Not approximately says it. Says it. The operations director in the opening story was smoothed, not invented from nothing; the AI took a real statement about response times and sharpened it into a fake statement about trust. A Source Pass catches exactly that, because the moment you put the quote next to the transcript, the mismatch is visible in five seconds.

The rules of the pass are strict and simple. Every direct quote gets checked against its transcript, word for word. Every figure, date, and named person gets checked against its document. And any claim that points nowhere, no label, or a label that does not contain the claim, gets one of exactly two treatments: you verify it manually from primary material and add the real citation, or you delete it. There is no third option where an unsourced claim stays in the deliverable because it "sounds right." Sounding right is the AI's core competency; it is precisely the property you cannot use as evidence.

Cost, so you can plan honestly: on a two-page assessment summary drawn from labeled sources, a disciplined Source Pass typically runs fifteen to twenty minutes. That is the price of never having a fabricated quote read aloud by the person it was pinned to. It is the cheapest insurance in this entire program.

Pass Two: The Spot-Audit Sample (Statistical Honesty at Volume)

The Source Pass works when the output is short and the claims are countable. It collapses when the AI has processed volume: classified 400 support tickets, extracted pain points from 20 interview transcripts, tagged 1,200 rows of process-exception data. You cannot re-read everything; that is the redo trap. And you cannot check nothing; that is the confession trap. The honest middle is sampling, and the discipline is the same AQL logic from the receiving dock: pull a sample sized to the risk, check it completely, and let the sample's error rate speak for the whole lot.

Here is a risk-tiered rule of thumb you can adopt today and tune with experience. These tiers are illustrative starting points, not an ISO standard; the point is to have pre-committed numbers before the deadline pressure arrives, because a sampling rate chosen at 11 p.m. the night before the readout will always be zero.

Deliverable tierSample rateWhy this rate
Internal working documents (your own notes, interim analyses)10 percent, minimum 10 itemsErrors here are recoverable; you will touch this material again before it matters.
Anything leadership-facing (readout decks, findings reports, steering-committee material)20 to 30 percentErrors here are read aloud in rooms with your name on the title slide; the cost of a defect is political, not clerical.
All figures, names, dates, and thresholds, in any deliverable, at any tier100 percent, alwaysNumbers and names are where fabrication is cheapest for the model and costliest for you; a wrong percentage or a misattributed statement is the error people remember.

Run the sample like an inspection, not a skim. For each sampled item, do the task the AI did, independently, and compare: read ticket 214 yourself and see whether "billing dispute" was the right classification; open Transcript 9 and see whether the extracted pain point is actually in it. Count your findings. And then apply the AQL consequence with a straight face: if the sample's error rate exceeds your tolerance, the whole lot goes back for rework, re-prompted, re-run, or re-done, not just the specific errors you happened to find. This is the step people flinch at, so anchor it in the factory logic: if 3 of your 30 sampled tickets were misclassified, the honest inference is that roughly 10 percent of all 400 are, which means about 40 wrong classifications are sitting in your deliverable and you found 3 of them. Fixing the 3 and shipping is not verification. It is laundering.

What tolerance should you set? For internal working material, many practitioners tolerate a sampled error rate up to around 5 percent, on the theory that downstream passes will catch survivors. For leadership-facing content, tolerance approaches zero for substantive errors: one fabricated quote in a sample of 25 items is not a 4 percent error rate, it is a stop-the-line event, because fabrication is a different species of defect than a mislabeled category. Write your tolerances down before you need them. That is what pre-committed means.

Pass Three: The Expert Walkthrough (Ground Truth Has an Address)

The first two passes are desk checks, and desk checks share a blind spot: they can only compare the output to documents. But the previous lesson taught you that the most dangerous AI failure in process work is not the wrong fact, it is the missing one, and the plausible one. Template gravity, the model's pull toward describing the generic version of a process instead of your organization's actual version, produces outputs that match every document and still describe a process that does not exist, because the real process lives in exceptions, workarounds, and tribal knowledge that no document ever captured. No sample rate catches an omission. You cannot cite-check a sentence that was never written.

The only instrument that detects this class of error is a human who lives inside the process. The Expert Walkthrough is thirty minutes, booked in advance, in which the process owner, the person who actually runs the escalation queue or closes the month or dispatches the trucks, reads the deliverable with one instruction: "Tell me what is wrong, what is missing, and what would make you wince if your boss read it." Not "please review," which produces a polite thumbs-up. A specific, adversarial ask. Watch where they slow down. The paragraph a process owner rereads twice is the paragraph the AI got subtly wrong.

Thirty minutes of a process owner's time has a real cost, which is why this pass is reserved for deliverables that describe how work actually happens: current-state summaries, process narratives, findings that will drive redesign decisions. It is also, quietly, the best relationship-building move in an assessment. Process owners who are asked to correct the record before publication become allies. Process owners who first see their process described wrongly in a steering-committee deck become the opposite, permanently.

The Verification Log: One Line That Makes You Auditable

Now for the artifact, the thing you will still be using years after you forget this lesson's prose. The Verification Log is a single running table, one line per AI-assisted deliverable, kept wherever your team already keeps working documents. Its columns:

  • Date and deliverable. What shipped, when, to whom.
  • What was AI-drafted. The specific portions machine-generated, stated plainly: "synthesis of 20 transcripts," "classification of 400 tickets," "first draft of current-state narrative."
  • Passes run. Which of the three passes executed: Source Pass, Spot-Audit (with sample size and rate), Expert Walkthrough (with the reviewer's name).
  • Errors found and fixed. Count and type: smoothing, fabrication, misclassification, omission. This column is the learning loop.
  • Sign-off. The named human who accepted the deliverable as verified. Accountability stays human; the log records whose.

Five columns, one line per deliverable, perhaps ninety seconds of writing per entry. Now look at what those ninety seconds buy you, in ascending order of importance.

First, defensibility. The day an auditor, a regulator, a legal hold, or a hostile CFO asks "was AI used to produce this, and how do you know it is right?", there are only two kinds of professionals: the ones who go pale and start reconstructing from memory, and the ones who open a table. The log converts "we used AI for this" from a confession into a documented, controlled practice with named accountability. This is also where the regulatory weather is headed; the European Union's AI Act phases in transparency obligations for AI-generated content through December 2026, and organizations that can already show what was machine-drafted and how it was verified will treat those obligations as paperwork rather than crisis.

Second, and this is the deeper payoff, the log is your personal learning loop, the thing MIT found missing in the 95 percent of pilots that returned nothing. After ten entries, your Errors Found column stops being a record and becomes a pattern. You will discover, concretely and for your own work, things like: the model fabricates most when synthesizing interviews about contentious topics; ticket classification is reliable above 95 percent except for one ambiguous category; quotes are the highest-risk output type you touch. Those patterns are instructions. They tell you which tasks need the 30 percent sample instead of the 10, which prompts need tightening, and which deliverable types should never ship without a walkthrough. Your verification gets cheaper and sharper every month, because it is learning. That is the difference between using AI and running an AI-assisted practice.

Third, the log feeds one small, high-leverage habit: the verification note. Every AI-assisted deliverable you ship carries one line, usually in the footer or the appendix: "AI-assisted draft; all sources checked; 25 percent spot audit, 3 findings, corrected; reviewed by [process owner]." Watch what that sentence does in a room. It preempts the gotcha question before anyone asks it. It signals a controlled process to exactly the audience (auditors, CFOs, steering committees) that is professionally trained to distrust uncontrolled ones. And it quietly raises the bar on everyone else's work, because the deliverables without a verification note now have a question mark that yours does not.

A Worked Example: Five Hours Instead of Sixteen, With Receipts

Here is the whole system running once, end to end, with illustrative numbers you can adapt. All figures are hypothetical; the arithmetic is the point.

The deliverable: an interview-synthesis report for a warehouse-operations readiness assessment. Twenty transcripts, averaging forty-five minutes each. The manual baseline, from your own history: reading, coding, and synthesizing this volume takes roughly sixteen hours of focused work, spread over four days you do not have.

The AI-assisted path, with the three passes designed in from the start:

  • AI draft: 1 hour. Cite-your-source prompt, transcripts fed in batches, output labeled claim by claim: theme, supporting quotes, transcript and paragraph references.
  • Source Pass: 1.5 hours. Every quote walked back to its transcript, every named attribution confirmed. Findings: two smoothing errors (real statements sharpened into words the interviewees did not use) and one fabricated quote, attributed to a shift supervisor, that appears in no transcript at all. All three corrected or removed. Sit with that middle finding for a second: that is the opening story, caught at a cost of ninety minutes instead of a career.
  • Spot-Audit: 2 hours. This is leadership-facing, so the 25 percent tier applies: five of the twenty transcripts re-read in full against the synthesis, checking that extracted themes are actually present and that nothing material was omitted. Findings: one theme overweighted (mentioned by two interviewees, presented as widespread), adjusted. Sampled error rate within tolerance; lot accepted.
  • Expert Walkthrough: 0.5 hours. The warehouse operations manager reads the draft with the adversarial instruction. She flags that the synthesis never mentions the seasonal-surge staffing workaround that dominates October through December, an omission no desk check could have caught, because the interviews barely touched it and the documents never did. One paragraph added.

Total: roughly five hours against a sixteen-hour manual baseline, an 11-hour saving on one deliverable even after paying the full verification cost. The report ships with a one-line verification note: "AI-assisted synthesis; all quotes source-checked; 25 percent transcript spot audit, 4 findings, corrected; reviewed by warehouse operations manager." One line in the Verification Log records the same. And here is the part that surprises people the first time: the verified AI version lands harder than the pure-manual version ever did, because the manual version never carried evidence of its own quality, and this one does. The steering committee is not being asked to trust a person's diligence. It is being shown a control.

Now the same deliverable on the other path, because this program's rule is that every success story travels with its failure twin. Same twenty transcripts, same deadline, but the assessor skips verification: the draft reads beautifully, the week is brutal, and checking feels like distrusting a tool that has been fine so far. The fabricated supervisor quote ships. Except in this version the smoothed quote is attributed to a named operations director, and it says something she never said about a function she does not own, and she is in the readout. The correction takes thirty seconds in the meeting. The damage does not. Every finding in the deck is now suspect; the walkthroughs that would have taken thirty minutes now happen anyway, after publication, adversarially, as a hunt for other errors. The assessor spends more hours defending the deliverable than verification would have cost by a factor of ten, and the political ledger is worse: for the remainder of the engagement, and plausibly for the remainder of that assessor's tenure, "check his numbers" is attached to his name. The deadline pressure that justified skipping verification saved four hours. The skip cost the thing assessments run on, which is the assumption that the assessor's documents mean what they say.

This, in miniature, is the workflow-redesign finding that anchors this whole program: McKinsey's research keeps showing that the organizations getting real value from AI are the ones that fundamentally redesign the workflow around the tool, roughly three times more likely among high performers than the rest, rather than bolting the tool onto old habits. Verification designed into the flow, with its hours budgeted up front, is workflow redesign at the scale of one professional. Verification bolted on afterward, when someone gets nervous, is the bad pilot pattern reproduced on your own desk. Level 3 of this program scales that idea to whole processes. You are practicing it now, on yours.

What to Do Monday Morning

This habit installs in one week or it does not install. Here is the sequence.

  1. Create your Verification Log before you need it. One spreadsheet, five columns: Date and deliverable; What was AI-drafted; Passes run (with sample size); Errors found and fixed (count and type); Sign-off. Put it where your team's working documents already live. An empty log you can open in one click will get used; a perfect log you plan to build later will not.
  2. Write down your personal sampling tiers and tolerances. Start with the illustrative defaults: 10 percent or minimum 10 items for internal working material, 20 to 30 percent for anything leadership-facing, 100 percent of all figures, names, dates, and thresholds regardless of tier. Add your rework rule: what sampled error rate sends the whole lot back. Pre-committed numbers survive deadline pressure; improvised ones do not.
  3. Run all three passes on your very next AI-assisted deliverable, even a small one, and time each pass. You are calibrating your own cost model: most people discover the full control costs 20 to 30 percent of the time the AI saved, which reframes verification permanently from "tax" to "obviously worth it."
  4. Book the Expert Walkthrough at the same moment you start the draft, not after it is finished. Thirty minutes on the process owner's calendar, with the adversarial instruction in the invite: "what is wrong, what is missing, what would make you wince." Booking it up front makes it part of the workflow instead of an optional extra the deadline will delete.
  5. Add the verification note to everything you ship this week. One line: what was AI-assisted, which checks ran, what was found and fixed, who reviewed. Then watch how the room treats the first deliverable that carries it.
  6. After five log entries, read your own Errors Found column. Look for the pattern: which task types produce fabrication, which produce omission, which are clean. Adjust one sampling tier based on what you find. That adjustment is your learning loop turning over for the first time.

Key Takeaways

  • Accept the job change: AI moved your role from "produce the draft" to "verify the draft," and the professionals whose work survives steering committees, auditors, and hostile CFOs are the ones who treat verification as the job, not an afterthought.
  • Design verification as a control sized to risk, like AQL incoming inspection on a factory line, never as a full redo (which erases the time savings) and never as a skipped step (which converts a deadline into a career event).
  • Run the Source Pass on every cite-your-source output: walk each quote, figure, and claim back to its labeled source, and give unsourced claims exactly two options, manual verification or deletion; budget 15 to 20 minutes on a two-page summary.
  • Sample at volume with pre-committed tiers: 10 percent or minimum 10 items for internal material, 20 to 30 percent for leadership-facing work, 100 percent of figures, names, and thresholds always, and send the whole lot back for rework when the sample exceeds tolerance.
  • Book the 30-minute Expert Walkthrough with the process owner for any deliverable describing how work actually happens, because omission and template gravity are invisible to desk checks and visible in seconds to a resident of the process.
  • Keep the Verification Log, one line per AI-assisted deliverable (what was AI-drafted, passes run, sample size, errors found and fixed, sign-off), so "we used AI for this" is a documented control with a named accountable human, not a confession.
  • Ship every AI-assisted deliverable with a one-line verification note; evidence of checking makes the AI-assisted version land harder than the pure-manual version, not softer.
  • Mine your own log as a learning loop, the mechanism MIT found missing in the 95 percent of failed pilots: your error patterns tell you which tasks need tighter sampling, which prompts need repair, and where your verification hours buy the most safety.