Prediction, Generation, Extraction, Agents: Four Different Readiness Problems
Two proposals land on the COO's desk in the same week, and both of them say "AI" on the cover page. The first wants $120,000 to score incoming insurance claims by denial risk. The second wants $90,000 for an agent that will "autonomously resolve" those same claims end to end. The COO, a careful person with a full calendar, reads them as two prices for the same category of thing, picks the cheaper one, and funds a project whose actual requirements nobody in the room has listed. Eleven months later the agent pilot is quietly canceled, the steering committee concludes "we tried AI for claims," and the scoring project, the one that would have worked, is now unfundable because it sounds like the thing that just failed. This is the specific damage done by one word covering four technologies. The previous two lessons in this chapter gave you the taxonomy and the mechanics; this one gives you the consequence: each of the four AI species demands a different readiness bar, and conflating them is how the wrong tool gets bought and the wrong pilot gets funded.
One Word, Four Different Purchases
Recall the species from this chapter's first lesson. Prediction systems score a case against historical outcomes: which claim will deny, which customer will churn, which invoice will pay late. Generation systems produce new text or content from a prompt and source material: the draft letter, the summary, the first-pass SOP (standard operating procedure). Extraction systems pull structured fields out of unstructured documents: the invoice number off the PDF, the denial code off the remittance. Agents chain all of the above into multi-step action: read the case, decide, and act in your systems without a human touching each step.
What that first lesson did not do, and this one does, is answer the buyer's question: what does my organization have to already possess for each species to be safe to deploy? Because the honest answer is four different lists. A prediction tool is starved without years of labeled outcomes; a generation tool needs almost no dataset at all but consumes verification hours forever. An extraction tool lives or dies on document variance you have probably never sampled; an agent needs everything the other three need, plus permissions, rollback, and logs that most organizations cannot produce on request.
Treating these as one purchase category produces two recurring accidents. The first is buying the wrong tool: funding an agent when the process needed a score, or buying a generation tool for what is actually an extraction problem, and then blaming "AI" when it underdelivers. The second is funding the wrong pilot: green-lighting the species whose demo was most impressive rather than the species whose readiness bar your organization actually clears today. The market data says this second accident is epidemic: Gartner projects that over 40 percent of agentic AI projects will be canceled by the end of 2027, and separately found that 63 percent of organizations lack or are unsure of AI-ready data practices. Read those two figures together and you see thousands of organizations funding the species with the highest bar while unable to clear the lowest one.
To compare the bars cleanly, we will walk each species through the same four dimensions: data needs (what must exist before day one), oversight needs (what human control structure the species requires), integration depth (how far into your systems the tool must reach to produce value), and failure economics (what going wrong characteristically costs, and when you find out). Go slowly through each one. The matrix at the end only works if you can feel why every cell is there.
The Prediction Bar: History, Thresholds, and the Cost of a Wrong Score
Data needs. A prediction model is a pattern-matcher trained on your past, so the non-negotiable input is labeled historical outcome data in volume: not just the cases, but what actually happened to each one. To score claims by denial risk you need years of claims each tagged with its real outcome, denied or paid, and the reason. Thousands of examples at minimum, and they must come from a process that has stayed recognizably the same, because a model trained on how your process worked before the 2024 reorganization is a model of a company that no longer exists. Stable process definitions are as much a data requirement as the rows themselves.
There is a third, less obvious data requirement: a tolerance analysis of wrong scores, done before purchase. Every prediction system produces false positives (flagged cases that were actually fine) and false negatives (clean scores on cases that go bad). Those two errors almost never cost the same. A mid-sized lender scoring loan applications for fraud might find a false positive costs 20 minutes of analyst review, about $18, while a false negative costs an average $9,400 write-off. That asymmetry, written down as two dollar figures, is what lets you set a sane threshold. An organization that cannot state what each error type costs has no basis for accepting any threshold a vendor proposes.
Oversight needs. Prediction oversight is threshold and override design. Someone accountable must decide the score above which a case gets flagged, who may override a score, and how overrides are recorded, because override patterns are your earliest evidence that the model and reality are drifting apart. This is lightweight compared to what agents demand, but it must be designed, not improvised by whoever sees the first weird score.
Integration depth. Moderate, and unforgiving on one point: the score must land inside the system where the decision is made. A denial-risk score that lives in a separate dashboard nobody opens during claim submission is a decoration. If the biller works in the practice-management queue, the score appears in that queue, sorted by it, or the project produces adoption charts and nothing else.
Failure economics. Prediction's characteristic failure is silent drift. The model keeps emitting confident scores while the world it learned from changes underneath it: a payer rewrites its rules, a product mix shifts, a pandemic rewires customer behavior. Nothing errors out; accuracy just decays, and you discover it months later in a KPI you were not watching. The pre-check is cheap: does anyone own a monthly comparison of predicted versus actual outcomes? If nobody can name that owner, drift will be found by the CFO instead.
The Generation Bar: Exemplars, the Verification Tax, and Fluent Fabrication
Data needs. Here is the inversion that surprises buyers: generation needs almost none of what prediction needs. No labeled history, no thousands of outcomes. What it needs instead is source material and style exemplars: the policy documents the drafts must be grounded in, ten examples of what a good output looks like, the template the team already trusts. A procurement team wanting first-draft supplier letters needs its contract playbook and a folder of its best past letters, assets most teams already have. This is why generation is usually the lowest-barrier species on the data dimension, and why it is so often the right first pilot.
Oversight needs. What generation saves you in data it charges you in review. Every output needs a verification step sized to its consequence: the verification tax you met in Chapter 1. A low-stakes internal summary might need a skim; a customer-facing letter needs a qualified reader checking every claim against source; a regulatory filing might need review so heavy the draft saves nothing. The readiness question is arithmetic, not enthusiasm: minutes to verify and correct versus minutes to write from scratch, multiplied by monthly volume. If verification costs 80 percent of drafting, the project's entire margin is an illusion that survives only as long as nobody measures.
Integration depth. Modest, with one trap. The output must land where the work continues: in the case file, the CRM (customer relationship management system), the document repository, already attached to the record it belongs to. The failure shape is the copy-paste bridge, a human ferrying text from a chat window into the system of record all day. Bridges like that feel free in week one and collapse under volume, and they also destroy the audit trail of what the AI drafted versus what the human changed.
Failure economics. Generation's characteristic failure is fluent fabrication caught late: the invented clause, the confident wrong figure, the plausible policy that does not exist, discovered after the document shipped. The cost is not the error itself but its lateness; a fabrication caught in review costs a correction, one caught by a customer or a regulator costs trust and remediation at many times the price. The cheapest pre-check: take ten real outputs, have your best reviewer verify them against source, and time it. That one afternoon prices the tax before you sign anything.
The Extraction Bar: The Worst Fifty Documents, Not the Best Twenty
Data needs. Extraction does not need training history the way prediction does, but it has a requirement buyers routinely fake without meaning to: a document sample that spans the real variance of what arrives. Every extraction demo runs on clean documents, and every production feed contains the other kind: the faxed third-generation scan, the handwritten margin note, the supplier who redesigned their invoice layout last month, the two-hundred-line consolidated statement. The readiness test is blunt: assemble the worst 50 documents your process received this quarter, not the best 20, and run the tool on those. Chapter 1's demo-versus-production lesson gave you the general law; extraction is where it collects most reliably, because the gap between demo documents and real documents is the entire gap between the pitch and the P&L.
Oversight needs. Extraction oversight is exception queue design plus confidence thresholds. The tool will read most documents correctly and must route the rest, the low-confidence reads, to a human queue. Readiness means that queue is designed before go-live: who staffs it, what its SLA (service level agreement) is, what confidence score triggers routing, and what happens when the queue backs up. An extraction deployment without a designed exception queue is a deployment where exceptions become silent errors flowing into downstream systems as fact.
Integration depth. Deeper than it looks. Extracted fields are only valuable landing in the destination system, in the destination schema: the ERP (enterprise resource planning system) wants a vendor ID from the master file, not a vendor name string; the billing system wants the denial code from its own code table. Schema mismatch is the quiet killer here; a tool that extracts "Acme Corp." perfectly while your ERP requires vendor 40017 has produced a lookup problem, not an automation.
Failure economics. Extraction's characteristic failure is the per-document exception cost multiplied by volume. The percentages sound harmless and the arithmetic is not. At 96 percent field accuracy on 20,000 documents a month, 800 documents need human handling; at 10 minutes each, that is 133 hours a month, most of a full-time role, appearing in nobody's business case because the vendor slide said "96 percent accurate." The cheapest pre-check is exactly the worst-50 test above, scored not as accuracy but as exceptions per hundred documents, priced at your loaded hourly cost.
The Agent Bar: Everything Above, Plus Permission to Act
Data needs. An agent chains prediction, generation, and extraction into sequences of real actions, so its data bar starts at the union of everything above and keeps going. On top of the histories, exemplars, and variance-spanning samples, an agent needs a bounded action space written down: the explicit list of actions it may take, systems it may touch, and amounts it may commit. It needs its own credentials and permissions (an agent acting under a human's login is an audit finding waiting to be written), a rollback path for every consequential action, and logging complete enough that you can reconstruct any decision three months later. If your organization cannot produce an access-rights map for its human staff, it is not ready to define one for software.
Oversight needs. Approval gates on consequential actions. The readiness work is drawing the line between what the agent does autonomously (assemble, check status, draft) and what waits for a named human's approval (submit, pay, cancel, communicate externally), with entry criteria for each gate. An agent proposal that cannot show you this line drawn does not have an oversight design; it has a hope.
Integration depth. The deepest of the four, by construction. An agent must read from and write to every system in its workflow: tools, APIs (application programming interfaces), credentials, error handling for each connection. Each integration is a project, and the agent needs all of them working simultaneously before it produces any value at all. This is why agent timelines are quoted in quarters and delivered in years.
Failure economics. Errors compound across steps. A 4 percent error rate per step sounds tolerable until you chain eight steps: the probability that at least one step goes wrong reaches roughly 28 percent per case, and unlike a wrong draft, a wrong action has already happened in a live system when you find it. The wrongly canceled order, the duplicate payment, the customer email that should never have been sent: these are incidents, not edits. The market has already graded this species' skipped homework: Gartner's projection that over 40 percent of agentic AI projects will be canceled by end of 2027 is not a verdict on the technology so much as on organizations that funded the highest bar first. The same research house found "agent washing" rampant, with only about 130 of thousands of vendors claiming agentic products judged the real thing, which means the label on the box does not even reliably tell you which species you are buying.
"Are we ready for AI?" is an unanswerable question. "Which species, and is that column's bar met?" is a Tuesday-afternoon checklist.
The Artifact: The Readiness-Bar Matrix
Here is the entire lesson compressed into the instrument you will keep. First the master comparison, then how to run it as a pre-purchase check.
| Prediction | Generation | Extraction | Agents | |
|---|---|---|---|---|
| Data needs | Labeled historical outcomes in volume; stable process definitions; costed tolerance for false positives and negatives | Source material and style exemplars; no large dataset | Document samples spanning real variance: the worst 50, not the best 20 | All of the left, plus a written bounded action space, permissions, rollback paths, and full action logs |
| Oversight needs | Threshold and override design with recorded overrides | Verification step sized to consequence (the verification tax) | Exception queue design plus confidence thresholds | Approval gates on every consequential action, with entry criteria |
| Integration depth | Moderate: score must land in the system where the decision happens | Modest: output must land where work continues, never a copy-paste bridge | Deep on schema: extracted fields must match the destination system's format | Deepest: reads and writes across every system in the workflow, with credentials and error handling |
| Characteristic failure | Silent drift as the world changes under the model | Fluent fabrication caught late | Per-document exception cost compounding at volume | Errors compounding across steps into live-system incidents |
| Cheapest pre-check | Name the owner of monthly predicted-versus-actual review; state both error costs in dollars | Verify ten real outputs against source and time the review | Run the worst 50 documents; score exceptions per hundred and price the queue | Ask for the action list, the approval gate line, and the rollback path in writing |
To use the matrix as a pre-purchase instrument, run three moves in order. Move one: identify the species. Ignore the product name and the word "agent" on the brochure; ask what the tool actually does to a case. Scores it? Prediction. Drafts about it? Generation. Reads fields off it? Extraction. Acts on it across systems? Agent. Many products are two species stapled together; evaluate each column separately. Move two: check the five cells in that column, demanding evidence rather than intention for each: the labeled data exists or it does not, the exception queue is designed or it is not, the rollback path is written or it is not. Move three: price every unmet cell as remediation before buying. An unmet cell is not automatically a no; it is a cost the vendor's proposal omitted. "We lack the worst-50 sample" becomes two days of document pulling; "no override design" becomes a workshop and a policy; "no rollback path" becomes, frequently, the honest discovery that the agent project is a year premature. The matrix's output is either a true total cost of ownership or a documented deferral, and both are wins. What it prevents is the third outcome: a purchase priced off the demo.
One Use Case, Four Verdicts: A Worked Example
Watch the matrix earn its keep on a single, realistic case. Meridian Claims Group is a fictional 70-person regional healthcare billing office, a composite for this lesson. It processes about 11,000 claims a month for physician practices; roughly 12 percent deny on first submission, and a nine-person follow-up team works those 1,300 monthly denials at an average 40 minutes each, about 880 hours a month. The managing partner announces the goal in exactly the words that cause accidents: "we should automate claims follow-up with AI." The operations lead, who has this lesson, refuses to treat that as one project and walks it through all four columns.
As prediction: score claims likely to deny before submission. Data check: the clearinghouse holds four years of claims, roughly 480,000, each labeled with its adjudication outcome and denial reason code. Labeled history in volume: met. Process stability: the office's submission workflow has been materially unchanged for three years: met. Tolerance analysis: a false positive sends a clean claim to a 6-minute pre-submission review, about $4 of biller time; a false negative just means a denial that would have happened anyway, so the downside is bounded at today's status quo. Oversight: flag the top-scoring 15 percent for review, revisit the threshold monthly. Integration: the score must appear in the submission queue inside the practice-management system, which the vendor's API supports. Column verdict: bar met. Projected impact, stated conservatively for the business case: catching a quarter of denials pre-submission saves roughly 220 follow-up hours a month.
As generation: draft appeal letters for denied claims. Data check: payer appeal policies are on file, and the team's shared drive holds 200 past appeal letters, including a folder of wins to use as exemplars: met. Oversight arithmetic: a senior biller currently drafts an appeal in about 45 minutes; verifying and correcting an AI draft against the claim record takes about 8. At 300 appeals a month, that is roughly 185 hours saved against a verification tax the team can actually pay: met. Integration: drafts must be created inside the appeals module attached to the claim, not in a separate chat window: the vendor supports it. Column verdict: bar met.
As extraction: pull denial codes and amounts from EOB documents. An EOB (explanation of benefits, the payer's statement of what it paid and why) arrives from more than 40 payers in wildly different formats, including scanned faxes from two regional payers. Data check: the demo ran on 25 clean PDFs from the three largest payers. Nobody has assembled the worst 50. Column verdict: bar not met, and priceable. The operations lead runs the pre-check: on a genuinely representative sample the tool posts a 13 percent exception rate. At 8,000 EOB documents a month, that is 1,040 exceptions; at 11 minutes each, 190 hours a month of exception handling, about 1.2 full-time roles the vendor's ROI slide did not contain. The demo was not lying about the clean payers; it was silent about the other 37. The project is not killed; it is re-priced with a designed exception queue and a two-payer phased rollout, and approved at the honest number.
As agents: auto-correct and resubmit denied claims end to end. The pitch is seductive: the agent reads the denial, fixes the coding error, resubmits, and follows up. The matrix column collapses it in twenty minutes. Bounded action space: not written. Permissions: the agent would act under a shared team login in the clearinghouse, which compliance will never sign. Rollback: a wrongful resubmission can trigger duplicate-claim denials and payer audit flags, and nobody can describe how to unwind one. Logs: the clearinghouse records submissions but not reasoning. Approval gates: undefined. Column verdict: bar failed; deferred, with a written list of what would have to exist (credentialed system access, a resubmission approval gate, an action log) before the question is asked again in twelve months.
Count what just happened. One vague mandate became four verdicts: two funded on evidence, one re-priced to its true cost, one deferred with a documented reason. Meridian gets a sequenced roadmap (generation first for fast wins, prediction next as integration completes, extraction phased behind a priced queue, agents parked) instead of one doomed mega-pilot named "automate claims follow-up." The alternative timeline is easy to write, because Gartner already wrote its ending: the agent pilot gets funded first because its demo was the most cinematic, joins the 40-plus percent of canceled agentic projects, and takes the credibility of the other three columns down with it.
What to Do Monday Morning
- Take the AI proposal nearest to you and name its species out loud. Prediction, generation, extraction, or agent; if the answer is "several," split it into columns and evaluate each separately. Refuse to proceed on any proposal whose species nobody can name.
- Draw the Readiness-Bar Matrix for that proposal: the five cells of its column on one page, each marked met, unmet, or unknown, with one line of evidence per cell. "The vendor says" is not evidence; it is the absence of evidence wearing a lanyard.
- Run the column's cheapest pre-check this week. Time ten verified outputs, pull the worst 50 documents, write down both error costs in dollars, or request the action list and rollback path in writing. Each of these costs an afternoon and re-prices a six-figure decision.
- Price every unmet cell as a remediation line item and add it to the proposal's cost before it reaches the approval meeting. If the total changes the decision, the matrix just did its job.
- If the proposal is agentic, apply the sequence test: which prediction, generation, and extraction bars does this agent implicitly depend on, and are those met? An agent proposal in an organization that has not cleared a single simpler column is a request to run before the walking test.
- File the completed matrix in your readiness portfolio. It sits beside Chapter 1's failure-mode checklist, and Chapter 3 will deepen every row of it into the full readiness lens.
Key Takeaways
- Treat "AI" as four purchases, not one: prediction, generation, extraction, and agents carry different data, oversight, integration, and failure-economics bars, and conflating them is how the wrong tool gets bought and the wrong pilot gets funded.
- Demand labeled historical outcomes, stable process definitions, and a costed false-positive versus false-negative analysis before any prediction purchase, and put the score inside the system where the decision happens.
- Fund generation on exemplars and verification arithmetic, not datasets: size the verification step to consequence, and reject any deployment whose output travels by copy-paste bridge.
- Test extraction on the worst 50 documents your process actually receives, design the exception queue and confidence thresholds before go-live, and price exceptions per hundred documents at volume.
- Hold agents to the union of all three bars plus a written action space, real permissions, rollback paths, logs, and approval gates; Gartner's 40-plus percent agentic cancellation projection is the market grading organizations that skipped this.
- Run the Readiness-Bar Matrix as a pre-purchase instrument: identify the species, check the five cells in its column with evidence, and price every unmet cell as remediation before signing.
- Sequence instead of mega-piloting: one use case examined through all four lenses yields funded quick wins, honestly re-priced projects, and documented deferrals rather than one cinematic failure.
- Carry this bridge forward: the species taxonomy from this chapter meets Chapter 3's readiness lens (people, process, data, governance) exactly here, one bar per species, one column per proposal.
Skill.re