←
AI Readiness & Process Transformation
Proficient · M5 · lesson 5 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Failure-Mode Mapping Before Launch

15 min

The room smells like burnt coffee and laminated floor plans. It is 1996, a tier-one automotive plant outside Toledo, and a new door-panel line is six weeks from its first shift. Nothing is running yet. Instead, eleven people, a process engineer, two operators, a maintenance lead, a quality manager, a buyer, are on hour five of a meeting whose only agenda item is imagining disaster. Station by station, they ask the same three questions: how can this step fail, what happens downstream when it does, and would we know? The welder can cold-weld and the joint looks perfect until it snaps in the customer's hands: severity nine, write it down. The torque gun can drift out of calibration slowly: who would notice, and when? By the end of the second day they have a spreadsheet of 140 ways the line can fail, each one scored, ranked, and, for the worst of them, answered with a designed catch before a single panel ships. Manufacturing has run this meeting for seventy years. It has a name, it has a worksheet, and it is the single most transferable discipline that world owes yours, because you are now six weeks from launching a line of your own: the redesigned invoice-exception workflow, with its marked map, its Gate Specs, and its Handoff Contracts, and nobody has yet sat down to imagine how it dies.

The Meeting Before the Line Runs

The discipline is called FMEA: Failure Mode and Effects Analysis. If you came up through Lean or Six Sigma, you already know it, and this lesson is your home turf meeting your new problem. If you did not, here is the operator's definition: FMEA is a structured meeting, held before a new process runs, in which the team enumerates every way every step can fail (the failure modes), traces what each failure does downstream (the effects), scores each one for how bad, how often, and how visible, and then designs catches for the worst ones while failures are still hypothetical and therefore free. It was born in the US military in the late 1940s, matured in aerospace and automotive, and became a pillar of quality engineering for one reason: it is dramatically cheaper to catch a failure on paper than on the line, and dramatically cheaper on the line than in the customer's driveway.

Notice what FMEA is not. It is not pessimism, and it is not a compliance ritual. The Toledo meeting was not a group of people who doubted their line; it was a group of people who intended to run it for ten years and wanted the surprises up front, at conference-room prices. That posture, professional imagination of failure as an act of confidence rather than doubt, is exactly what most AI pilots skip, and the skipping shows up in the record this program keeps returning to. MIT found 95 percent of enterprise GenAI pilots deliver no measurable return, and Gartner projects that over 40 percent of agentic AI projects will be canceled by the end of 2027, with inadequate risk controls named among the leading causes. Inadequate risk controls is analyst language for a precise operational fact: nobody held the meeting. Nobody enumerated the failure modes, so nobody designed the catches, so the first real failure arrived unannounced, in production, wearing the pilot's budget as a name tag.

This lesson closes the chapter you have been building all along. You have a marked map of the invoice-exception process showing which steps the AI takes, which stay human, and which are shared. You have Gate Specs: the designed human review points, each with entry criteria and a review standard, because in this program a human gate is a designed control, never a gesture. You have Handoff Contracts specifying exactly what the AI passes to the human and what the human passes back. Every one of those artifacts is a claim about how the workflow behaves when things go right. FMEA is where you stress-test all of them against things going wrong, on paper, where a failure costs a marker and a sticky note instead of a clawback letter and a CFO briefing. The output is this lesson's named artifact, the AI-FMEA Worksheet, and it will be reused more times than anything else you build this level.

One caveat before you carry the classic tool across the bridge, and it is the caveat this whole lesson turns on. FMEA transfers to AI-augmented workflows almost perfectly, with one twist: AI steps fail differently from machines. A bearing screams before it seizes. A torque gun throws an error code. A jammed conveyor stops, visibly, and someone walks over. An AI step does none of this. Its characteristic failure is confident, fluent, well-formatted output that is wrong, delivered at exactly the speed and polish of its correct output. No smoke, no code, no scream. The classic FMEA scoring survives the trip for severity and occurrence, but the third dimension, detectability, has to be rebuilt from the ground up, and rebuilding it is where this lesson earns its place at the end of the chapter.

How AI Steps Fail: The Enumeration Engine

The hardest part of any FMEA is the blank page: how do you enumerate failures for a technology your team has run for eight months instead of eighty years? Manufacturing engineers inherit failure catalogs (fatigue, wear, contamination, drift) refined across generations. You need the equivalent for AI steps, and you already met its raw material in Level 2, where the Output Skeptic lessons taught you the individual failure patterns of AI output: the fabricated figure, the invented step, the confident wrong summary. What changes now is placement. Those were patterns in a document; these are failure modes in a flow, each attached to a specific step on your marked map, each with downstream effects you can trace along the swimlanes. Walk your map step by step and enumerate from the catalog that matches each step's type.

Extraction steps: reading the document

Extraction steps pull structured values out of unstructured inputs: the invoice number, the total, the purchase-order reference, the line items. Their failure catalog has three entries that look similar and behave very differently downstream. First, the confidently wrong value: transposed digits in an amount, the shipping total read as the invoice total, the right field from the wrong line item. Nothing about the output signals trouble; the number sits in its field looking exactly like a correct number. Second, the right value from the wrong document: a multi-invoice PDF, a stapled credit memo, an emailed thread with three attachments, and the extractor reads page two's total into page one's record. Third, the missed field, and here you must split one failure mode into two, because their downstream effects diverge completely. A missed field returned as blank is a kindness: blanks are easy to detect, easy to route, easy to fill. A missed field returned as hallucinated filler, a plausible date, a rounded amount, a vendor name pattern-matched from history, is poison, because it flows downstream indistinguishable from real data. When your team enumerates extraction failures, write blank-versus-filler as separate rows. One is a nuisance with detectability on your side; the other is among the most dangerous rows on the sheet.

Categorization steps: naming the exception

Categorization steps assign each invoice exception a type: price variance, quantity mismatch, missing purchase order, duplicate, contract-rate mismatch. Their signature failure is systematic bias toward frequent classes. The model has seen ten thousand price variances and forty contract-rate mismatches, so when a contract-rate mismatch arrives wearing ambiguous clothes, the model confidently calls it a price variance. The cruelty of this failure mode is that aggregate accuracy hides it perfectly: a model that is 96 percent accurate overall can be 40 percent accurate on the rare class that happens to carry your compliance exposure, and the weekly summary will read beautifully. If that sounds familiar, it should: the handoff lesson's weekly-summary failure story, where a clean-looking report concealed a lane quietly filling with mislabeled exceptions, was this failure mode wearing a narrative. On the worksheet it becomes a row with a score, which is the whole move of this lesson: stories become rows, rows get catches. Add two more entries to the catalog: novel-input misassignment, where a genuinely new thing (a new vendor's format, an exception type your taxonomy has never seen) gets confidently jammed into the nearest old category because the model has no way to say "none of the above"; and drift, where a categorizer that scored well at launch degrades over months as suppliers, formats, and exception mix shift underneath it. Machines wear out; categorizers go stale. Both belong on the sheet.

Generation steps: writing the words

Generation steps draft text: the resolution note, the vendor email, the exception summary for the approver. Their catalog is the L2 fabrication family relocated into a flow: invented specifics (a delivery date no one promised, a policy clause that does not exist, a discount the contract never granted), omitted caveats (the draft states the resolution as settled when the SOP, the standard operating procedure that governs the workflow, requires it to be framed as proposed pending approval), and tone failures in outbound text, which sound cosmetic until you enumerate the effect: a curt dispute email to your largest vendor's accounts team is a relationship incident with a dollar tail, and it deserves severity scoring like anything else.

Routing and confidence steps: deciding who sees it

Finally, the steps that decide the path: confidence scoring and routing. Their first failure mode is miscalibration, when the confidence bands your triage design relies on stop meaning what they claim, and items scored 0.9 are right only 70 percent of the time. You built a calibration check when you designed the triage thresholds; on the FMEA it gets reclassified as what it always was, a detection control for this exact failure mode. The second failure mode has no analog in machinery at all, and your sheet must state it in plain language: threshold gaming under operational pressure. Recall the gate lesson's insurance story, where a backlogged team facing a quarter-end push nudged the auto-approve threshold down "temporarily" and the temporary became policy. On the worksheet, that story becomes an explicit cause in the occurrence column: operational pressure is a failure cause, listed by name, exactly as a manufacturing FMEA lists "operator bypasses guard to save time." A good FMEA includes human and organizational failure causes on equal terms with technical ones, and an honest AI-FMEA doubles down on this, because most of what goes wrong around an AI step is done by people responding rationally to pressure.

Scoring the Sheet: Severity, Occurrence, and the Detectability Twist

Classic FMEA scores every failure mode on three ten-point scales, then multiplies them into a single rank. Two of the scales cross to AI workflows intact. The third breaks, and understanding why it breaks is the most important idea in this lesson.

Severity (S, 1 to 10) scores how bad the effect is if the failure reaches its destination, and it works exactly as your Six Sigma colleagues remember: money at stake, compliance exposure, customer impact, safety. A mislabeled internal note might score 2; an overpayment released to a vendor scores 7 or 8; anything touching regulatory reporting climbs from there. Severity is a business judgment, and it stays human. An AI can help you enumerate effects, but the question "how much do we care?" is answered by the people who own the consequences, which is this program's accountability rule wearing a scoring pencil: every AI-touched decision has a named owner, and severity is where the owner's judgment enters the arithmetic.

Occurrence (O, 1 to 10) scores how often the failure mode happens, and here AI workflows actually start with an advantage most new production lines never had: you already have data. The dirty-sample simulations you ran in Level 2's data-quality triage, where you pushed real, messy documents through the pilot task and counted failures, are occurrence estimates. If the extraction step misread 4 of 150 sampled invoices, you have a defensible starting rate for the confidently-wrong-value row. Where you have no data, estimate honestly, mark the estimate as such, and let the pilot's instrumentation replace it: occurrence scores improve every month you run, which is one reason the worksheet is a living document and not a launch ritual.

Detectability (D, 1 to 10) is where the classic scale breaks, so slow down here. In manufacturing, detectability asks: if this failure occurs, how likely is it to be noticed before it escapes? Machines cooperate with that question. Failures announce themselves: vibration, noise, error codes, scrap that looks like scrap. An AI failure does the opposite. It passes review. It arrives fluent, formatted, and plausible, and it will sail through every checkpoint that relies on someone sensing something is off, because nothing is off except the content. So the AI-FMEA recalibrates the scale with one rule: score detectability by the control that exists, not by intuition. Do not ask "would we probably notice?" Ask "which specific control catches this, and what is its actual catch probability?", and where the control is a sampling plan, compute it.

Here is what computing it looks like on the running example. Your redesigned workflow routes about 900 routine invoice exceptions per month past the gate on the strength of the AI's classification, and your Gate Spec commits to reviewing a 10 percent sample: roughly 90 items a month, four or five per working day. Now test that control against failure mode one from the catalog: a systematic miscategorization concentrated in one lane, say 5 percent of the routine flow, about 45 items a month, all of them contract-rate mismatches wrongly labeled as price variances. If your 90 samples are drawn at random from the whole routine lane, you will touch four or five of the bad items a month, scattered across days and reviewers, each one looking like an isolated miss rather than a pattern. Weeks pass before anyone connects them, if anyone does. But stratify the sample, and the arithmetic flips. Stratified sampling means you divide the flow into its lanes (here, the AI's predicted exception categories) and draw a fixed quota from every lane every week, including the thin ones, instead of letting the sample fall where volume is. Under a stratified plan, the price-variance lane gets, say, ten reviewed items a week no matter what, and if the model is dumping mislabeled contract-rate mismatches into that lane, two or three of those ten are wrong in the same way in the same week. That is not an anomaly; that is a cluster, and clusters get escalated. Same failure, same 10 percent budget: random sampling detects it in months, stratified sampling in days. Write D as 3 with stratification and 7 without, and you have just experienced the lesson's central mechanism: a bad detectability score is not a grade, it is a work order. You go redesign the catch, on paper, before launch, which is the entire point of holding the meeting now.

An AI failure does not announce itself; it passes review. So score detectability by the control that exists, not the vigilance you hope for, and treat every bad score as a catch waiting to be designed.

Multiply the three scores and you get the RPN, the Risk Priority Number: S times O times D, ranging from 1 to 1,000. Sort the sheet by RPN, descending, and the top rows are your design agenda. For each one, design a catch from three families, in order of preference. Prevention stops the failure from occurring: input validation that rejects malformed documents before extraction, format allowlists that only admit vendor layouts the system was tested on. Detection catches the failure after it occurs but before it lands: stratified sampling as above; reconciliation totals, where an arithmetic identity (line items must sum to the invoice total) checks the AI's output automatically; and canary inputs, which are known-answer test items seeded into the live flow so that if the system starts mishandling them, you learn immediately (Chapter 3.5 builds these into live monitoring). Containment limits the blast radius when a failure gets through anyway: the AI's output cannot release a payment, because payment release sits behind the human-only step on your marked map. Notice what just happened to that placement decision: back in the triage lesson you made it on principle; the FMEA now re-justifies it with arithmetic, because containment is what caps severity when prevention and detection both miss. When someone asks why the map is drawn the way it is, the worksheet is the answer.

The Artifact: The AI-FMEA Worksheet

The named artifact of this lesson, and the capstone artifact of the chapter, is the AI-FMEA Worksheet. One row per failure mode, grouped by step from your marked map, with nine columns: the step; the failure mode in plain language; the downstream effect traced along the flow; severity (1 to 10); occurrence (1 to 10, with its evidence source noted: simulation, pilot data, or estimate); detectability (1 to 10, scored against the named control, with the computation shown where a sampling plan is involved); the RPN; the designed catch, tagged prevention, detection, or containment; and finally the owner and the test, one named person accountable for the catch existing and working, and one concrete way to verify it does. A row without an owner is a wish. A row without a test is a hope. The worksheet tolerates neither.

Building it is a partnership, and the division of labor matters. AI is your enumeration partner. Give it the real inputs: "Here is the SOP for this step and the Handoff Contract that feeds it. List every failure mode for this step. Be exhaustive. Include human and organizational failure causes, not just technical ones." A strong model, prompted with your actual artifacts, will produce failure modes your team missed and will do it without the social cost of being the person who keeps imagining disasters in the meeting. AI is your devil's advocate on scores. After your team scores a row, ask: "Argue that detectability is worse than scored. What has to be true for this control to miss?" You will keep maybe one challenge in three, and the ones you keep are worth the exercise. AI is your completeness check. When the sheet feels done, ask the closing question: "What failure mode on this sheet has no catch? What failure mode is missing entirely?" What stays human is exactly what the accountability rule says must: every severity score is a business judgment signed by the owner of the consequence, and every catch has a budget that a human approved, because the AI will cheerfully design you a beautiful control regime that costs more than the workflow saves.

The Top Five Rows: Invoice Exceptions, Worked

Here is the centerpiece: the top five rows of the invoice-exception AI-FMEA, scored and caught. The numbers are illustrative, built for a hypothetical mid-size firm running about 1,200 exceptions a month with roughly 900 routed as routine, but the reasoning is the part you are meant to steal.

#Failure modeSODRPNDesigned catch
1Systematic miscategorization of one exception type into a common class74384Stratified sampling by predicted lane + weekly lane report
2Extraction transposes digits in the invoice amount83248Automatic reconciliation: line items must sum to invoice total
3Novel exception type confidently routed to the routine lane654120Novelty flag: unseen vendor/format combinations route to the gate regardless of confidence
4Reviewer fatigue degrades the gate's catch rate by month 3665180Batch caps, reviewer rotation, double-review sampling to measure drift
5Upstream format change (vendor portal update) silently degrades extraction746168Input-format monitor + the charter's re-score trigger

Row 1 is the weekly-summary story converted into arithmetic. Severity 7: a mislabeled contract-rate mismatch skips the scrutiny its compliance exposure demands. Occurrence 4: the dirty-sample simulation showed the rare class confused at meaningful rates. Detectability 3, and hold on to the qualifier: 3 after stratification. Before the sampling plan was stratified, this row scored D7 and an RPN of 196, top of the sheet; the redesign of the catch, one paragraph of arithmetic and a change to the Gate Spec's sampling clause, cut the RPN by more than half without touching the model. That is FMEA doing its job: the score fell because the design improved. Owner: the exception-process owner. Test: seed a batch of known contract-rate mismatches and confirm the weekly lane report surfaces the cluster.

Row 2 has the lowest RPN of the five and is on the sheet anyway, because severity 8 rows earn a catch regardless of rank: a transposed amount on a large invoice is real money out the door. The catch is the best kind that exists: an arithmetic identity. Line items must sum to the total; a transposition breaks the identity; the check runs on every invoice, costs effectively nothing, and never gets tired. Detectability 2, the best score on the sheet, and note the design principle it demonstrates: the best catches are arithmetic, not vigilance. Every time you can replace "the reviewer will notice" with "the equation will fail," take the trade. Owner: the finance-systems analyst. Test: submit test invoices with transposed amounts and confirm every one flags.

Row 3 is novel-input misassignment with the highest occurrence on the technical rows, because new vendors and new formats arrive constantly and the model will always have an opinion about them. The catch is an entry-criteria amendment to the Gate Spec: any invoice from an unseen vendor/format combination routes to the gate no matter how confident the classifier is, because confidence on novel input is exactly the thing you cannot trust. Watch what this row did: the FMEA fed back into an earlier artifact and changed it. The Gate Spec you finished two lessons ago just got its first amendment, and that is the sheet working as designed, stress-testing the chapter's artifacts and improving them while changes still cost nothing. Owner: the triage designer. Test: submit an invoice from a fabricated vendor in a never-seen layout and confirm it hits the gate at any confidence score.

Row 4 is the highest RPN on the sheet, and it contains no AI at all. Reviewer fatigue is a human failure mode, sitting in the FMEA on equal terms with the technical rows, exactly as the discipline demands. Occurrence 6, because it is close to certain: every reviewing team's catch rate degrades as the work becomes routine, and month 3 is when the gate quietly becomes a rubber stamp. Detectability 5, because degraded attention looks identical to efficient attention from the outside. The catches are the gate lesson's fatigue designs, now justified by numbers instead of intuition: batch caps so no reviewer sees more than a fixed run of items, rotation so no one lane becomes anyone's autopilot, and double-review sampling, a small slice of items independently reviewed twice, so that reviewer disagreement rates give you a measurable signal of drift. Owner: the gate's review-standard owner. Test: track the double-review disagreement rate monthly; a falling catch rate shows up there before it shows up in escaped errors.

Row 5 carries the worst detectability score on the sheet, D6, and deserves its reputation. When a vendor updates their portal and the invoice layout shifts, extraction does not stop working; it starts working slightly worse, silently, on one vendor's documents. No error, no blank fields, just a rising miss rate localized where your aggregate metrics will not see it for weeks. The catch is change management dressed as a control: an input-format monitor that fingerprints incoming document layouts and alerts on new ones, wired to the re-score trigger you pre-committed to in the Level 2 charter, the clause that says a material change in inputs forces a re-validation before the workflow keeps running at full autonomy. Owner: the pilot's technical lead. Test: replay a deliberately altered document format and confirm the monitor fires and the trigger opens.

One more thing about these five rows, and it is the thing that sells the effort to whoever asks why you spent two afternoons on a spreadsheet. Each row ends with an owner and a test, and those tests do not retire after launch. In Chapter 3.4, the testing lesson turns the worksheet's test column into the pilot's test plan. In Chapter 3.5, the incident playbook takes its scenarios straight from these rows, because an incident is just an FMEA row that happened. The AI-FMEA is the pilot's threat model, written once and reused three times: as a design review before launch, as a test plan during the pilot, and as an incident catalog in operations. Few artifacts in this program pay rent in three rooms at once. This one does.

The Tuition of Skipping the Meeting

Now the failure story, because there is always a team that skips the meeting, and their reasoning always sounds pragmatic. A regional distribution firm, hypothetical but assembled from patterns you will recognize, launches its own invoice-exception redesign on a compressed timeline. When the process lead proposes an FMEA session, the sponsor waves it off with the sentence that should be printed on warning labels: "We'll learn from the pilot." Understand what that sentence actually proposes. The pilot is the FMEA, run in production, at customer expense, with real invoices as the test cases and real money as the scoring system. The team was never choosing between doing the analysis and skipping it; they were choosing where to run it and who pays tuition.

Week 7, the invoice arrives. A large one: $47,150 from a major freight vendor. The extraction step reads it as $74,150, a clean transposition, confidently formatted, syntactically perfect. It is exactly row 2 from the worksheet the team never wrote. There is no reconciliation check, because no one enumerated the failure mode that would have demanded one, so the transposed figure flows through categorization, passes the gate inside a batch of routine approvals (a fluent wrong number passes review; that is what it does), and payment releases. The vendor, to their credit, flags the overpayment, eleven days later. The clawback is awkward: five weeks of emails, a credit note negotiation, and a hole in the month's cash forecast. The CFO briefing is worse than the money. "How did a $27,000 overpayment pass every checkpoint?" has no good answer when the honest one is "we decided to learn from the pilot," and the AI program's credibility takes the kind of dent that the S&P Global statistic is made of: 42 percent of companies scrapped most of their AI initiatives in 2025, and dents like this are how the scrapping starts.

Here is the detail that makes the story sting instead of merely cautionary. That afternoon, the same afternoon as the CFO briefing, the team designs the reconciliation check. It takes one meeting. Line items must sum to the total; flag the mismatch; done. The catch that would have cost one conference room hour in week 0 was always one meeting away; the pilot did not teach them anything the worksheet would not have, it just charged $27,000, five weeks, and a reputation for the same lesson. "We'll learn from production" is FMEA with the highest possible tuition, and the charter you signed in Level 2 already knew it: the kill condition you pre-committed to is watching precisely these rows. An FMEA row that fires repeatedly with no working catch is what a kill condition is for; the worksheet tells the kill condition what to look at.

And with that, step back and look at what you are holding, because this closes the chapter. A marked map that says which steps the AI takes and which stay human. Gate Specs that make every human checkpoint a designed control with entry criteria and a review standard. Handoff Contracts that specify what crosses each boundary and in what shape. And now an AI-FMEA Worksheet that has stress-tested all of it, ranked the ways it fails, and designed the catches, with owners and tests attached. The redesign is complete on paper, which is the cheapest place anything is ever complete. Chapter 3.2 builds it for real: rewriting the SOP for the AI-augmented team, grounding the AI in your actual documents, structuring its outputs so systems can check them, weaving verification into the flow, and instrumenting the whole thing so the occurrence column stops being an estimate. Everything in that chapter will move faster because of the sheet you just built, because building is easy when you already know where it breaks.

What to Do Monday Morning

  1. Book the meeting. Ninety minutes, the people who own the redesigned workflow's steps, the marked map on the wall. Call it what it is: the failure-mode session, held now because paper is cheaper than production.
  2. Run the enumeration prompt for each AI step. Feed the AI the step's SOP section and its Handoff Contract, and ask: "List every failure mode for this step. Be exhaustive. Include human and organizational failure causes." Merge its list with the team's, using the taxonomy: extraction, categorization, generation, routing.
  3. Score your top ten rows, and compute detectability instead of guessing it. Severity from the consequence owner, occurrence from your Level 2 dirty-sample rates where you have them, and for every row whose control is a sampling plan, do the arithmetic: sample size, error concentration, time to a visible cluster. If the numbers demand stratification, stratify.
  4. Design catches for the top five RPNs. Prefer prevention, then detection, then containment, and prefer arithmetic to vigilance every time the choice exists. Every catch gets an owner and a test, written in the row.
  5. Amend one earlier artifact based on what the sheet exposed. A Gate Spec entry criterion, a Handoff Contract field, a sampling clause. If the FMEA changed nothing upstream, it was a ritual, not a review; the feedback loop is the proof it worked.
  6. File the worksheet where its next two users will find it. The Chapter 3.4 test plan and the Chapter 3.5 incident playbook will both be built from these rows. Put the AI-FMEA in the pilot's shared folder, dated, versioned, next to the charter whose kill condition it feeds.

Key Takeaways

  • Hold the failure meeting before launch: FMEA (Failure Mode and Effects Analysis) enumerates every way every step fails, scores each one, and designs the catch while failures are still hypothetical and free, and it transfers from manufacturing to AI workflows almost perfectly.
  • Respect the one twist: AI steps fail without error codes or smoke, producing confident, fluent, wrong output that passes review, so classic detectability scoring breaks and must be rebuilt around controls.
  • Enumerate from the taxonomy, step by step: extraction (wrong value, wrong document, blank versus hallucinated filler), categorization (frequent-class bias, novel-input misassignment, drift), generation (invented specifics, omitted caveats, tone), and routing (miscalibration, threshold gaming under pressure).
  • Score detectability by the control that exists, not intuition, and compute it where you can: 90 random samples find a lane-concentrated 5 percent error in months, while the same 90 samples stratified by lane find it in days.
  • Rank rows by RPN (severity times occurrence times detectability) and design catches in preference order: prevention, then detection, then containment, choosing arithmetic checks over human vigilance wherever the option exists.
  • Keep human failure modes on the sheet on equal terms: reviewer fatigue at month 3 and threshold gaming under quarter-end pressure are FMEA rows with scores, owners, and catches, not footnotes.
  • Use AI as enumeration partner, devil's advocate on scores, and completeness checker, while humans own every severity judgment and every catch's budget under the accountability rule.
  • Reuse the AI-FMEA Worksheet three times, as pre-launch design review, as the Chapter 3.4 test plan, and as the Chapter 3.5 incident playbook's scenario list: it is the pilot's threat model, and "we'll learn from production" is the same analysis at the highest possible tuition.