Recognizing Bad AI Output in Process Work
The summary is beautiful. You asked the model to condense your invoice-process interviews and the accounts payable standard operating procedure into two pages, and it delivered exactly that: clean headings, confident prose, a crisp bulleted flow from invoice receipt to payment. It reads like the work of a sharp consultant on a good day. You are ninety seconds from pasting it into the readiness report when your eye snags on one line: "invoices above the $5,000 threshold route to a second approver under the standard 48-hour SLA." A service level agreement, an SLA, with a number attached. Reasonable. Professional. And you cannot remember anyone, in any interview or any document, ever mentioning it. You search the SOP. Nothing. You search the interview notes. Nothing. The model invented it, dressed it in the same confident prose as everything else, and buried it in paragraph four where it looks exactly like a fact. That is the whole problem of this lesson in one line: bad AI output in process work almost never looks bad. It looks finished.
The Deliverable That Looks Done
In Level 1 you learned what an AI hallucination is: a fluent, confident, fabricated claim produced by a system that predicts plausible text rather than retrieving verified facts. That was the concept. This lesson is the craft: what those failures actually look like when they are sitting inside a process summary, a process map, or a stakeholder analysis, on your screen, an hour before a deadline, wearing a suit.
The mental model most people carry into AI work is wrong in a specific, dangerous way. They imagine bad output as visibly bad: garbled sentences, obvious nonsense, the kind of error that announces itself. That failure mode exists, and it is harmless, because nobody ships garbled nonsense to a chief financial officer. The failure mode that ends careers is the opposite one: the deliverable that is roughly 92 percent right, formatted immaculately, and wrong in the 8 percent that only someone who knows the process could catch. The polish is not incidental to the danger. The polish is the danger. A human analyst who was unsure about the approval threshold would hedge, flag it, or leave a bracketed question. The model does not hedge by default. It renders its guesses and its facts in the same confident register, which means the errors arrive pre-camouflaged.
Think about what a wrong number looks like in a spreadsheet: it sits in a cell, visibly a number, checkable against a source. Now think about what a wrong number looks like inside two pages of fluent prose: it looks like every other sentence. There is no red underline for "this SLA does not exist." Your eyes are the underline, and untrained eyes slide right past.
The stakes here are the program's spine. MIT's autopsy of the 95 percent of enterprise generative AI pilots that produced no measurable return found the failures were organizational: no workflow integration, no learning loop, adoption without transformation. A missing learning loop has a very concrete street-level meaning for you as an assessor: the tool that fabricated the 48-hour SLA today will fabricate something else tomorrow, and correcting it does not teach it anything. The verification burden never goes away. It is not a bug you wait out; it is a permanent feature of the work, which is why this program's rule is absolute: verify every AI-touched figure, step, and claim. But you cannot verify what you cannot spot. The next lesson installs the verification habit as a workflow. This lesson trains the eye that makes the workflow worth running. First you learn to see the five ways a polished draft lies; then you learn to check.
The dangerous AI output is not the one that looks wrong. It is the one that looks done.
Five Ways a Polished Draft Lies
Across process summaries, process maps, interview syntheses, and readiness analyses, AI failures cluster into five recognizable patterns. Each has a signature, a tell you can train yourself to notice, the way an experienced auditor notices a too-tidy expense report. Go slowly here. These five patterns are the vocabulary of everything that follows in this level.
Pattern One: Fabrication, the Invented Specific
Fabrication is the model producing a step, threshold, figure, system name, or policy that exists nowhere in your sources. It is the classic hallucination from Level 1, now wearing operational clothes: a "48-hour SLA," a "$5,000 threshold," a "monthly reconciliation meeting," a step where "the ERP system auto-flags duplicates" in an organization whose enterprise resource planning (ERP) system does no such thing.
Why does it happen? Because the model has read thousands of descriptions of invoice processes, and invoice processes in its training data very often have SLAs and thresholds. When your Source Pack, the curated bundle of SOPs, interview notes, and ticket samples you learned to build in the previous lesson, is silent on a detail, the model does not leave a gap. It fills the gap with the statistically typical answer. It is not lying in the human sense; it is completing a pattern. The result is indistinguishable from lying.
The tell: specifics that are suspiciously round, unattributed, or absent from the Source Pack. Real operational numbers are lumpy and traceable: the threshold is $7,350 because of an old policy memo, the turnaround is "usually two or three days, worse at quarter close" because a human said so in an interview. Fabricated numbers are smooth: 48 hours, $5,000, 99 percent, five business days. When you see a crisp, round, convenient specific, your first question is not "is this right?" but "where exactly in my sources does this come from?" If you cannot point to the line, treat it as invented until proven otherwise.
Pattern Two: Omission, the Missing Mess
Omission is the inverse crime: not what the model added, but what it silently dropped. You fed it forty ticket descriptions and it returned four tidy themes that cover perhaps 80 percent of them; the remaining 20 percent, the weird ones, the exceptions, the angry one from the warehouse, simply vanish. You fed it an SOP with a messy exception path for invoices that arrive without a purchase order, and the summary describes only the happy path, because the happy path is cleaner to write.
Omission is uniquely dangerous in readiness work because the exceptions are the work. Any process looks automatable if you only look at its happy path. The 20 percent of cases that are exceptions routinely consume 80 percent of the effort, and they are exactly what an automation pilot will choke on. A summary that drops them is not a shorter version of the truth; it is a different, sunnier process that does not exist, and any readiness score built on it is fiction.
The tell: the output is shorter and cleaner than the input mess warrants. You handed the model chaos and received serenity; be suspicious of the exchange rate. The single most effective probe costs one sentence: ask the model directly, "what did you leave out of this summary?" and watch it produce a list. That list is often startling, and the fact that the model can produce it on demand tells you the omissions were not accidents of ignorance but casualties of compression. Nothing in the default behavior surfaces them. You have to ask.
Pattern Three: False Confidence, Uncertainty Laundered
False confidence is what happens to hedges in the wash. Your interview notes say "Priya thinks the approval usually takes a day or two, but she wasn't sure about month end." Your ticket data shows wild variance. The summary says: "Approvals complete within one to two business days." The uncertainty went in; a declarative sentence came out. The model laundered your doubt into a fact.
This pattern matters because readiness assessment runs on calibrated uncertainty. A claim you are 95 percent sure of and a claim you are 55 percent sure of should be handled completely differently downstream: different verification effort, different caveats in the report, different weight in the scoring. Prose that erases the difference between them corrupts every decision built on it. And executives read confidence as evidence. A declarative sentence in a clean document, in front of a steering committee, becomes an organizational fact within a week, regardless of its pedigree.
The tell: no hedges anywhere, despite conflicting sources that you registered yourself. You lived the interviews. You know the SOP contradicted the tickets, that two stakeholders disagreed, that half the answers came with a shrug attached. If the draft reads as if none of that friction ever happened, if every sentence is declarative and nothing is flagged as uncertain, contested, or unverified, the confidence is cosmetic. Real synthesis of messy sources should visibly carry some of the mess.
Pattern Four: Smoothing, the Averaged Contradiction
Smoothing is false confidence's more insidious sibling. Where false confidence erases uncertainty about one claim, smoothing erases a conflict between two sources. The SOP says the accounts payable manager approves all invoices. The tickets show team leads approving most of them in practice. These two facts contradict each other, and the contradiction is the finding: it tells you the documented process and the real process have diverged, which is one of the most important things a readiness assessment can discover. The model, asked to summarize both sources, quietly picks one version, or blends them into a plausible middle ("invoices are approved by AP management"), and moves on. The contradiction does not appear in the output at all. It has been averaged out of existence.
Understand why this happens and you will never stop watching for it: the model is optimized to produce coherent text, and contradictions are incoherent. Its training pressure pushes it toward the smooth, single-narrative version of any story. But in process work, the friction is the signal. Divergence between SOP and practice, disagreement between stakeholders, variance between sites: these are precisely the findings your assessment exists to surface, and they are precisely what the model's coherence instinct deletes.
The tell: conflicts you know exist do not appear in the output. This tell is unique among the five because it lives in your memory, not on the page. Nothing in the document looks wrong; something you know is missing from it. Before reading any AI synthesis, write down the two or three conflicts you personally registered in the sources. Then check whether the draft names them. If the draft resolved them silently, it has smoothed, and you must ask what else it smoothed that you did not happen to know about.
Pattern Five: Template Gravity, the Generic Drift
Template gravity is the subtlest of the five. The model has ingested the textbook version of every standard business process: the industry-average invoice flow, the canonical onboarding journey, the best-practice procurement cycle. When it describes your process, that generic template exerts a constant gravitational pull, and details drift toward it. Steps you never described appear because the textbook process has them. Your draft mentions a "three-way match" between purchase order, receipt, and invoice; your organization does a two-way match. It mentions the "vendor portal"; there is no vendor portal, there is a shared inbox called AP-INVOICES and a woman named Marta who knows everything. The output describes a respectable, median company. It just does not describe yours.
Template gravity is dangerous for a reason specific to your role: a readiness assessment of the industry-average process is worthless, because pilots are not deployed against the industry average. They are deployed against Marta's inbox, the missing purchase orders, and the month-end pileup. Every imported best-practice step in your process map is a landmine: it makes the process look more standardized, more documented, and more ready than it is, which inflates the readiness score exactly where inflation does the most damage.
The tell: the output reads like a textbook and could describe any company. It mentions systems, roles, or steps your organization does not have. The probe is a simple substitution test: read each paragraph and ask, "could this sentence appear unchanged in a description of our nearest competitor?" A faithful process summary is full of your organization's fingerprints: the odd system names, the workaround, the person who is a single point of failure. A summary with no fingerprints has drifted to the template, and the specifics it does contain deserve double suspicion.
The Artifact: The Output Skeptic's Checklist
Here is this lesson's named artifact, the tool you will carry out of it: the Output Skeptic's Checklist. It is the five patterns turned into five questions, run against any AI deliverable before it leaves your desk. Not after the meeting. Not when someone asks. Before it leaves your desk, every time, the way a pilot runs a pre-flight check regardless of how good the plane looks.
| Pattern | The tell | The question you run |
|---|---|---|
| 1. Fabrication | Specifics that are round, convenient, and unattributed | Can I point to the exact line in the Source Pack for every number, threshold, system, and step named here? |
| 2. Omission | Output shorter and cleaner than the input mess warrants | Have I asked the model "what did you leave out?" and reviewed the list, especially exception paths? |
| 3. False confidence | No hedges anywhere despite conflicts I registered myself | Does every uncertain or contested claim from the sources still read as uncertain here? |
| 4. Smoothing | Conflicts I know exist are absent from the output | Did I list the known contradictions before reading, and does the draft name each one instead of resolving it silently? |
| 5. Template gravity | Reads like a textbook; mentions things we do not have | Does every paragraph carry our fingerprints, or could it describe any company in our industry? |
Three usage notes. First, the checklist is a reading protocol, not a form to file: it changes how you read, from "is this well written?" to "where would each pattern hide in this?" Well written is the default output of the technology; it is no longer evidence of anything. Second, run the questions in order, because they compound: fabrications found in question one sharpen your suspicion for question five, since an invented specific is often the template talking. Third, notice what the checklist does not do: it does not verify anything. It finds the claims that need verification. Spotting and checking are two different skills, and the next lesson, on the verification habit, is where the checking workflow lives. The checklist is the eye. The habit is the hand.
The Invoice Summary Autopsy: A Worked Example
Pattern recognition is built on reps, so here is a full rep. All numbers in this example are hypothetical, chosen to be realistic. An assessor at a mid-sized distributor asks a model to summarize the invoice-approval process from a Source Pack containing the accounts payable SOP, three interview transcripts, and a sample of 40 helpdesk tickets. The model returns a two-page summary. It is excellent prose. It also contains one planted instance of each pattern. Run the checklist with me.
The draft says: "Invoices are received via the vendor portal and logged automatically. Standard invoices are approved by AP management within one to two business days. Invoices above $5,000 route to a second approver under the standard 48-hour SLA. Approved invoices undergo a three-way match before payment. The process is consistent and well documented."
Question one, fabrication. Point to the line in the Source Pack for "$5,000" and "48-hour SLA." You cannot. The SOP mentions a second approval above $7,350 (an old policy memo amount, lumpy and real) and no SLA anywhere. Two round, convenient, unattributed specifics: both invented. Caught.
Question two, omission. The input mess included a fat exception path: invoices arriving without a purchase order, which the tickets suggest is roughly one in five. The summary's flow has no exceptions at all. You ask the model what it left out, and it promptly lists the missing-PO path, the disputed-quantity path, and the quarter-end backlog. Three exception paths, gone; the summary described a process about 20 percent of invoices never experience. Caught.
Question three, false confidence. "Within one to two business days" is declarative. Your interview notes say Priya was unsure and month end is worse; the tickets show variance from four hours to nine days. The uncertainty was in the sources and is gone from the prose. Caught.
Question four, smoothing. Before reading, you wrote down the conflict you knew about: the SOP says the AP manager approves; the tickets show team leads approving in practice. The draft says "approved by AP management," a phrase engineered to be compatible with both and to surface neither. The divergence between documented and actual process, your single most important finding, has been averaged into a plausible middle. Caught.
Question five, template gravity. "Vendor portal" and "three-way match": the company has neither. Invoices arrive in a shared inbox, and matching is two-way. Both phrases came from the industry-average process in the model's training data, not from your sources. The final sentence, "consistent and well documented," could describe any company on earth and describes this one least of all. Caught.
Five patterns, five catches, perhaps ten minutes of desk review. Now price the alternative, illustratively. Suppose the invented $5,000 threshold sails through and becomes an input to the readiness report's automation sizing: the volume of invoices needing second approval is estimated off a threshold that does not exist. The report reaches the chief financial officer, who knows the real approval rules cold, spots the phantom threshold on page six, and asks the question you never want asked in a steering committee: "if this number is wrong, which of the other numbers are right?" The answer costs you three weeks of line-by-line re-verification of the whole report, call it 60 hours of work, and something more expensive than hours: the report's credibility, and yours. Every future finding now gets the folded-arms treatment. Ten minutes at your desk, or three weeks and a reputation. That is the exchange rate this checklist trades at.
The Deck That Erased the Dissent: A Failure Story
The worked example shows the mechanics; this story, a composite of the pattern this program exists to prevent, shows the blast radius. An internal transformation lead, call him Marcus, ran fourteen stakeholder interviews ahead of a warehouse-automation pilot and used a model to synthesize them into a themes deck for the sponsor. Eleven interviews were broadly positive. Three, all from the operations team that would actually live with the pilot, were sharply negative: they flagged that the pick-list data was unreliable, that the proposed workflow ignored the night shift, and that a previous tool had been quietly abandoned for the same reasons. Real, specific, load-bearing dissent.
The model did what coherence-optimized systems do: it smoothed. The deck's theme slides reported "broad enthusiasm with some change-management considerations," a phrase that is technically compatible with the dissent and communicates none of it. Marcus, reading a polished deck that matched the meeting's optimistic mood, shipped it. No one in the steering committee ever saw the operations team's objections, so no one resourced them. The pilot launched, and hit, with almost comic precision, exactly the three problems the dissent had named: bad pick-list data, a night shift that refused a workflow designed without them, and a workforce that had pattern-matched this tool to the last abandoned one. Six months and a mid-six-figure budget later, the pilot joined MIT's 95 percent, and the post-mortem's ugliest finding was that the organization had paid for its own early warning, three times, in interview transcripts nobody re-read.
Connect this to the arithmetic you already know. BCG's 10-20-70 rule holds that AI success is 10 percent algorithms, 20 percent technology and data, and 70 percent people and process. Stakeholder dissent is the 70 percent talking. A synthesis step that smooths dissent out of the record is not a formatting glitch; it is a silent amputation of the largest success factor you have. Smoothing does not just corrupt documents. It corrupts the organization's picture of itself, one plausible middle at a time, and the bill arrives during the pilot, when it is most expensive to pay.
What to Do Monday Morning
Pattern recognition is a trained skill, and the training is cheap. Here is the sequence.
- Print the Output Skeptic's Checklist and put it where you work: the five patterns, their tells, and the five questions. On paper, deliberately. For the first few weeks this must be a physical ritual, not a memory exercise, because the whole failure mode is that polished output switches your skepticism off.
- Re-review one AI output you already accepted. Pick something recent that you used: a summary, a set of meeting notes, a synthesis. Run all five questions against it with the original sources open. Expect to find at least one pattern; finding it in something you shipped is the fastest possible cure for trusting polish.
- Run the plant-and-find drill with a colleague. Each of you takes a short process summary and seeds it with one fabrication: a plausible invented threshold, SLA, or step. Swap. The other must find the plant using the checklist. This drill does for output skepticism what a fire drill does for evacuation: it makes the rare event practiced.
- Ask "what did you leave out?" after every summarization, starting today, as a standing follow-up prompt. Read the list it returns and decide consciously, item by item, whether each omission is acceptable. One sentence of prompt against the most invisible failure pattern is the best trade in this lesson.
- Start a catch log. One line per catch: date, deliverable, pattern number, what the tell was. Within a month you will know which patterns your tools and your prompts are most prone to, and that log becomes both your personal calibration record and, later in this program, evidence for the human-gate designs you will build in governance work.
Key Takeaways
- Treat polish as camouflage, not evidence: the dangerous AI output in process work is the deliverable that is 92 percent right and immaculately formatted, because the wrong 8 percent is invisible to anyone who does not know the process.
- Learn the five failure patterns by name: fabrication (the invented specific), omission (the dropped mess), false confidence (laundered uncertainty), smoothing (averaged contradictions), and template gravity (drift toward the generic industry process).
- Memorize each pattern's tell: round unattributed specifics; output cleaner than the input warrants; no hedges despite known conflicts; missing contradictions you personally registered; textbook prose that mentions things your organization does not have.
- Run the Output Skeptic's Checklist on every AI deliverable before it leaves your desk, in order, as a reading protocol that replaces "is this well written?" with "where would each pattern hide?"
- Ask the model "what did you leave out?" after every summarization, because omissions are compression casualties the model can list on demand but will never volunteer.
- Write down known source conflicts before reading any AI synthesis, then check that the draft names them, since smoothing leaves no trace on the page and can only be caught from your own memory of the friction.
- Price the catch honestly: finding the invented threshold at desk review costs about ten minutes; finding it after the report reaches the CFO costs weeks of re-verification and the credibility of every other number you shipped.
- Remember the spine: MIT's 95 percent failed on missing learning loops and missing verification, the tool that fabricated today will fabricate tomorrow, and this lesson's trained eye is what the next lesson's verification habit is built on.
Skill.re