←
AI Readiness & Process Transformation
Aware · M21 · lesson 21 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

What AI Is and Isn't for Process Owners

15 min

The vendor's slide says "Autonomous Order-to-Cash" and the room leans forward. The demo shows an invoice appearing, a payment matching itself, a polite collections email writing itself, and the narrator says the sentence that will cost this company a year if nobody interrupts: "The AI runs the entire process end to end." Sitting third from the left is the order-to-cash process owner, the person whose name goes on the steering-committee deck if this gets bought, and she is about to ask the only question that matters. Not "how accurate is it?" Not "what does it cost?" Her question is: "Which AI?" Because she has learned the thing this lesson exists to teach you: for process work, "AI" is not one thing. It is at least four different things, four working species that consume different inputs, produce different outputs, and fail in four different ways. Buy the wrong species for a step, and you have purchased a very expensive way to join the 95 percent.

One Word Doing Four Different Jobs

Here is the trap built into the vocabulary. When your CFO says "AI," she may mean the model that forecasts demand. When your sales director says "AI," he means the chatbot that drafts his proposals. When your shared-services lead says "AI," she means the tool that pulls invoice numbers out of PDF attachments. When the vendor says "AI," he means whatever the prospect seems to want. Four people, one word, four entirely different machines, and a procurement process that treats them as interchangeable.

This matters to you specifically because you own processes, and a process is a sequence of steps, and each of these machines does something different to a step. A demand forecaster cannot draft a collections email. A drafting tool cannot reliably pull structured data out of ten thousand supplier PDFs. An extraction tool cannot decide anything. And the "autonomous agent" that claims to do all of the above in sequence is, in 2026, the single most oversold object in enterprise software: Gartner expects over 40 percent of agentic AI projects to be canceled by the end of 2027, and when it examined the thousands of vendors claiming to sell agents, it judged only around 130 of them to be selling the real thing. The rest is what Gartner calls agent washing: old automation wearing a new word.

In the previous chapter you watched the failure record assemble itself: MIT's finding that 95 percent of enterprise GenAI pilots produced no measurable return, and the autopsy behind it. One of the quiet routes into that 95 percent is simply buying the wrong species for the step: a generation tool deployed against an extraction problem, an agent deployed against a judgment problem. The tool works exactly as designed. It is just the wrong design for the work. This lesson gives you the taxonomy that prevents that purchase, and an artifact, the Four-Species Placement Map, that turns the taxonomy into something you can run against any process map you own by Friday.

The Four Species, in Operator Terms

Forget the research taxonomy (machine learning, deep learning, transformers, and the rest); that is the biologist's tree of life, and you are running a farm. You need to know what each animal eats, what it produces, and how it dies. For process work there are four species that matter, and every credible AI proposal you will ever review is one of them, or a combination you should insist on seeing decomposed.

Species one: prediction (scoring, classification, forecasting)

What it consumes: structured historical data, lots of it. Rows and columns: past orders, past defaults, past machine failures, past churn. What it produces: a number about the future or about an unseen case. A score (this customer is 82 percent likely to pay late), a class (this ticket is a warranty claim, not a billing dispute), a forecast (demand for SKU 4471 next month is 1,900 units, plus or minus 300).

Prediction is the oldest species, decades older than the chatbots, and it is the one your organization may already run without calling it AI: the credit score, the demand forecast, the fraud flag. In a process, prediction attaches to steps where a human currently makes a judgment call from pattern experience: which orders to hold, which claims to audit, which machines to service first. Characteristic failure mode: the wrong-but-confident score. A prediction model never says "I don't know"; it says "23 percent" with the same straight face whether it was trained on ten thousand relevant cases or two hundred irrelevant ones. It also decays silently as the world drifts away from its training data: the model that scored credit beautifully in 2024 quietly misprices risk after your customer mix shifts, and no error message ever fires. One process example: scoring incoming orders for credit risk so the credit analyst reviews the 30 risky ones instead of all 220.

Species two: generation (drafting text and artifacts)

What it consumes: an instruction and context, in plain language: "draft a past-due reminder for this account, firm but polite, referencing these three invoices." What it produces: a draft. Text, mostly: emails, summaries, SOP paragraphs, job descriptions, meeting notes, translations, code. This is the species the world met in late 2022, the one that made "AI" a boardroom word, and the one most people now mean by default.

In a process, generation attaches to steps where a human currently composes something: writes the email, summarizes the call, drafts the report section. It does not attach to steps where a human decides something, and confusing those two step types is the most common placement error you will see. Characteristic failure mode: fluent fabrication. Generation produces prose whose confidence is completely uncorrelated with its accuracy: the invented delivery date, the plausible-but-wrong policy citation, the summary that smooths over the one sentence that mattered. The draft reads beautifully, which is exactly what makes the error survivable all the way to the customer. Every generation step therefore ships with a permanent verification step attached, owned by a named human; you met this principle in the failure record, and you will meet it in every level of this program. One process example: drafting the first version of dunning emails for overdue accounts, which a collector reviews, edits, and sends.

Species three: extraction (structured data out of unstructured documents)

What it consumes: documents. PDFs, scans, emails, contracts, forms: the unstructured sediment that every real process runs on. What it produces: fields. PO number, line items, quantities, renewal date, liability cap, remittance reference: structured data your systems can actually use, extracted from documents only humans could previously read.

Extraction is the unglamorous species, and for back-office process work it is routinely the highest-ROI one, which should not surprise you: the MIT autopsy found the clearest returns exactly where document-heavy, high-volume, measurable work lives. In a process, extraction attaches to steps where a human currently reads a document and types what it says into a system. If your process map contains the words "keys into," you have found an extraction candidate. Characteristic failure mode: silent mis-extraction. The tool does not refuse the blurry scan or the oddly formatted PO; it extracts something, populates the field, and moves on, and the wrong quantity flows downstream wearing the uniform of clean data. The defense is boring and essential: confidence thresholds, sampling audits, and a reconciliation step that catches the 2 percent before it becomes a shipped order for 1,100 units instead of 100. One process example: pulling header and line-item data from emailed PDF purchase orders straight into the ERP entry screen for a clerk to confirm.

Species four: agents (multi-step, tool-using execution)

What it consumes: a goal, plus permissions to use tools: query this system, send that email, update this record. What it produces: a sequence of executed actions. Where the first three species each perform one kind of step, an agent chains steps: look up the order, check the inventory system, draft the response, update the ticket, escalate if X.

Agents are real, and they are also the species drowning in the deepest marketing. The honest 2026 description is this: agents execute bounded multi-step sequences under explicit permissions, with defined escalation rules, on processes whose steps and exceptions are already well mapped. They are a power tool for well-understood workflows, not a replacement for process ownership. Characteristic failure mode: runaway error compounding. Each step an agent takes has some error rate, and errors multiply across steps: a step reliability of 95 percent looks excellent in isolation and produces roughly a 60 percent success rate across a ten-step chain, before you account for the exception cases nobody mapped. Worse, an agent's errors are actions, not drafts: an email sent, a record changed, a refund issued. A generation error waits politely for review; an agent error has already happened. This is the arithmetic underneath Gartner's cancellation forecast, and it is why every agent deployment question in this program will come back to two words: bounded and permissioned. One process example: handling routine order-status inquiries end to end (look up, compose, send, log) with automatic escalation of anything involving a change, a complaint, or a credit.

SpeciesConsumesProducesCharacteristic failurePlacement test
PredictionStructured historical dataA score, class, or forecastWrong but confident; silent driftIs there a pattern to predict?
GenerationInstruction plus contextA draft artifactFluent fabricationIs there a draft to generate?
ExtractionUnstructured documentsStructured fieldsSilent mis-extractionIs there a document to extract from?
AgentsA goal plus tool permissionsExecuted action sequencesCompounding multi-step errorsIs there a bounded sequence with clear rules?

Three Framings That Cost Money

Taxonomy in hand, you can now dismantle the three sentences that route more budget into the 95 percent than any vendor ever did. Each one sounds harmless in a meeting. Each one embeds a species error.

"AI understands our business"

No species understands anything. Prediction recognizes statistical patterns in your history; generation recognizes patterns in language; extraction recognizes patterns in document layouts. Pattern recognition is genuinely powerful, and it is also why every species fails at the edges of its patterns: the customer unlike any previous customer, the contract clause nobody has written before, the PO format from the newly acquired subsidiary. When someone says "it understands our business," translate silently to "it has seen patterns like ours," and then ask the operator's follow-up: seen how many, how recent, and how similar to the cases where we actually lose money? The expensive version of this framing is skipping the verification design because "it understands what we need," which is how a fabricated policy citation reaches a customer wearing your signature block.

"AI decides"

Prediction scores; it does not decide. The distance between "this order scores 82 percent risk of late payment" and "hold this order" is a decision, and a decision needs an owner, a threshold, an exception path, and an audit trail. None of those are properties of a model; all of them are properties of a designed handoff, and designing handoffs is your profession. When a process narrative says "the AI decides which claims to pay," what actually exists is either a threshold somebody set (in which case that somebody made the decision, once, in advance, and should have their name on it) or nobody set it consciously (in which case your process has an unowned decision executing thousands of times, which is not automation, it is abdication). Accountability stays human in every framework this program teaches, not as philosophy but as control design: every AI-touched decision gets a named owner and a review standard, exactly like any other delegated authority in your organization.

"AI will run the process end to end"

This is the fantasy the vendor slide was selling, and the placement map you are about to build is its antidote. Walk any real process step by step and the fantasy dissolves arithmetically: some steps have a pattern to predict, some have a draft to generate, some have a document to extract from, some are physical, some are relational, some are judgment calls your auditors and regulators require a human to own. "End to end" survives only at slide altitude. At step altitude, the honest questions become: which steps, which species, what boundaries, whose sign-off. Gartner's 40 percent cancellation forecast for agentic projects is, at root, a forecast about organizations that bought the slide instead of walking the steps.

AI never takes over a process; it takes positions in one, step by step, one species at a time, under rules you design.

The Artifact: The Four-Species Placement Map

Here is this lesson's deliverable, and it may be the highest-leverage two hours in this level, because it converts you from a spectator of AI proposals into the author of them. The Four-Species Placement Map is your existing process map (swimlane, SIPOC, flowchart, or a numbered list on one page; the notation does not matter) with every step annotated by four tests and a verdict.

For each step, ask the four placement tests in order:

  • The prediction test. Does this step contain a judgment made from pattern experience, with structured history behind it? (Which cases are risky, which will be late, how much will we need.) If yes, prediction is a candidate, and the follow-up question is whether the history is clean and deep enough, which is Chapter 3's territory.
  • The generation test. Does this step produce a composed artifact: an email, a summary, a document section, a response? If yes, generation is a candidate, and the design work is the verification step that ships with it, never optional.
  • The extraction test. Does this step involve reading a document and re-typing its contents into a system? If yes, extraction is a candidate, and the design work is thresholds and sampling audits against silent mis-extraction.
  • The agent test. Is this step actually a short, bounded sequence of steps with explicit rules, mapped exceptions, and low blast radius when wrong? If yes, an agent is a candidate, and the design work is permissions and escalation. If the sequence is long, exception-rich, or high-stakes, the honest verdict is "decompose it and place the species individually."

Then write one of four verdicts on the step: a species name (with the human verification or decision handoff noted beside it), maybe later (a species fits, but a prerequisite is missing: data quality, volume, a stable upstream step), human (no test fires, or the step is judgment, relationship, or accountability that stays human by design), or not AI (the step is deterministic and belongs to ordinary rules-based automation, which is often the correct and much cheaper answer). That last verdict matters more than it looks: a surprising share of "AI opportunities" are actually rules jobs, and putting a probabilistic species on a deterministic step buys you error rates you did not need to own.

Two rules govern the map. First, steps get verdicts, not processes: the moment anyone assigns a verdict to the whole process, you are back on the vendor slide. Second, every species verdict names a human: who verifies the draft, who owns the threshold, who audits the extraction sample, who reviews the agent's escalations. A placement map with no names on it is a wish list.

Walking the Map: Order-to-Cash at Corven Mills

Now the full walkthrough. Corven Mills Distribution is a hypothetical composite: a fictional 480-person industrial-supplies distributor, around 190 million dollars in revenue, roughly 220 orders a day, whose COO has just received the "Autonomous Order-to-Cash" pitch at 280,000 dollars a year. The order-to-cash process owner builds the placement map instead of the business case the vendor offered to write for her. Her process has nine steps; here is the walk, step by step, with her reasoning and the rough stakes.

Step 1: Order intake and entry. About 60 percent of orders arrive as emailed PDF purchase orders that two clerks re-key into the ERP (enterprise resource planning system), roughly four minutes each: call it 8 to 9 clerk-hours a day, near 2,200 hours a year, around 80,000 dollars loaded, plus a 1.5 percent keying-error rate that surfaces downstream as wrong shipments and credit memos. The extraction test fires loudly: documents in, fields out, human confirms on screen. Verdict: extraction, with the clerks moving from typists to verifiers and a weekly sampling audit owned by the order-desk lead.

Step 2: Credit review. Every order from a new account, or one pushing an account past its limit, queues for a credit analyst: about 30 a day, ten minutes each, five analyst-hours daily, and the queue adds a day of cycle time to exactly the orders from the newest customers. Years of payment history sit in the ERP. The prediction test fires: score the queue, let the analyst work the risky tail first and fast-track the clean 70 percent. Verdict: prediction, with the credit manager owning the threshold, reviewing drift quarterly, and keeping the hold decision human. The score ranks the queue; it does not release an order.

Step 3: Pricing and contract-terms validation. Checking negotiated pricing, rebate terms, and special conditions against the contract. It looks like a pattern job until she counts: only about 15 orders a day hit an exception, each one is a different negotiated snowflake, and a mispriced order damages a named relationship. No test fires cleanly; the history is too thin and too heterogeneous to predict from. Verdict: human, revisit in a year if exception volume grows.

Step 4: Inventory allocation and promise dates. A forecasting problem in principle (the prediction test half-fires), but Corven Mills' demand history sits in three systems with inconsistent SKU coding, and the planner's spreadsheet overrides are recorded nowhere. Deploying prediction on that substrate would automate the wrong-but-confident failure mode at scale. Verdict: maybe later, gated on the data-readiness work Chapter 3 covers; she notes Gartner's warning that through 2026, 60 percent of AI projects without AI-ready data will be abandoned, and declines to volunteer.

Step 5: Pick, pack, and ship. Physical work orchestrated by the warehouse management system. No document to extract, no draft to write, no pattern-judgment a model should own from this seat. Verdict: human (with its own automation roadmap that has nothing to do with this taxonomy).

Step 6: Invoice generation. The ERP already generates invoices from shipped orders deterministically. The vendor deck had an AI icon on this step anyway. Putting a probabilistic species on a solved deterministic step adds error surface and audit questions for zero gain. Verdict: not AI, and she writes "rules job, already done" beside it as a small monument to discipline.

Step 7: Collections correspondence. Four collectors spend about 90 minutes a day each composing past-due reminders and payment-plan follow-ups: call it 1,500 hours a year, around 60,000 dollars, and the letters vary wildly in tone and quality. The generation test fires: draft from the account record and dunning stage, collector reviews and sends. Every draft is verified because the failure mode here is a fluently fabricated promise ("as we agreed, the balance is due March 30") that no one agreed to. Verdict: generation, with the collections supervisor owning a weekly quality sample.

Step 8: Cash application. Matching incoming payments to open invoices from remittance advices that arrive as PDFs, spreadsheets, and portal downloads: one and a half FTEs, and the true sequence is extract-then-match-then-post, a bounded chain with clear rules and a mapped exception path. The agent test half-fires, and honestly. But she sequences: prove extraction at step 1 first, then extend the same discipline here, because an agent posting cash wrongly is an action taken, not a draft caught. Verdict: maybe later, explicitly second in line, permissions and escalation to be designed before any pilot.

Step 9: Disputes and deductions. Short-pays, damage claims, negotiated resolutions with named customers. Judgment, relationship, and precedent; the audit trail must show a human weighing the account's history and the customer's temper. Generation may eventually draft the settlement letters, but the resolution itself is the kind of decision that stays owned. Verdict: human.

The finished map reads: three steps placed now (extraction on intake, prediction on credit, generation on collections), two maybe-laters with named gates (allocation, cash application), four steps human or not-AI by design. The three placements together touch roughly 3,700 staff hours and something like 140,000 to 160,000 dollars of annual cost plus a measurable keying-error rate, small enough to baseline honestly, large enough to matter, and each one is a single species on a single step with a named human beside it. The COO takes one look at the map next to the 280,000-dollar "autonomous" pitch and the meeting gets very short. That is what the taxonomy is for: not to slow AI down, but to aim it.

What to Do Monday Morning

You have a taxonomy and a map format. Here is the sequence to make them yours this week.

  1. Pick one process you genuinely own, with 6 to 12 steps, and get its map on one page. If no current map exists, write the numbered step list from memory and validate it with the two people who live in the process; the placement map is only as honest as the step list under it.
  2. Run the four tests on every step (pattern to predict? draft to generate? document to extract from? bounded sequence with clear rules?) and write a verdict on each: species, maybe later, human, or not AI.
  3. Put a name beside every species verdict: who verifies, who owns the threshold, who audits the sample. If you cannot name the human, downgrade the verdict to maybe later, because you have found a missing control, not an opportunity.
  4. Attach rough stakes to the placed steps: hours per year, loaded cost, error rate if you have it. Estimates are fine; the point is knowing whether you are aiming at a 5,000-dollar step or a 150,000-dollar one before anyone writes a business case.
  5. Translate one live "AI" sentence this week. In the next meeting where someone says "AI will handle it," ask which species, on which step, verified by whom. Do it kindly; do it every time. You are installing the question the successful 5 percent ask by habit.
  6. File the map in your readiness portfolio next to Chapter 1's failure-mode checklist. The two artifacts work as a pair: the checklist tells you whether an initiative is being run honestly; the map tells you whether the right species was purchased at all.

Key Takeaways

  • Treat "AI" as four working species, not one thing: prediction (scores, classes, forecasts from structured history), generation (drafts from instructions and context), extraction (structured fields from unstructured documents), and agents (bounded multi-step execution under permissions).
  • Memorize the four failure modes in pairs with their species: wrong-but-confident scores that drift silently, fluent fabrication that reads beautifully, silent mis-extraction that ships bad data in a clean uniform, and agent errors that compound across steps and execute as actions rather than drafts.
  • Kill the three expensive framings on sight: "AI understands our business" (it recognizes patterns and fails at their edges), "AI decides" (it scores; deciding is a designed handoff with a named owner), and "AI will run the process end to end" (it takes positions in a process, step by step).
  • Weigh the agent hype against the arithmetic and the record: per-step errors compound across chains, actions are not drafts, and Gartner expects over 40 percent of agentic AI projects canceled by end of 2027, with only around 130 of thousands of claimed agent vendors judged real.
  • Build the Four-Species Placement Map on any process you own: run the four placement tests per step, write one of four verdicts (species, maybe later, human, not AI), and refuse verdicts at process altitude.
  • Name a human beside every species verdict, who verifies the draft, owns the threshold, audits the extraction sample, or reviews the agent's escalations; a map without names is a wish list, and an unnamed threshold is an unowned decision.
  • Respect the "not AI" verdict: deterministic steps belong to rules-based automation, and placing a probabilistic species on a solved rules job buys error rates and audit burden for zero gain.
  • Expect the Corven Mills shape in your own processes: a minority of steps take a species now, a few wait on prerequisites with named gates, and several stay human by design, which is not a disappointing result but the exact aim that keeps you out of the 95 percent.