Agent Washing: Spotting Rebranded Automation
The renewal notice arrives in March and the operations director reads it twice, because something has changed and it is not the product. The invoice-routing tool her team has run since 2023, the one that moves PDFs from a mailbox into the ERP (enterprise resource planning system) along rules her own analyst configured, is no longer called a workflow suite. It is now "an autonomous AI agent workforce for finance," and the renewal price has grown 38 percent to match the new vocabulary. Same screens. Same rules. Same Tuesday-morning failure when a supplier sends a slightly weird PDF. The only thing that learned anything new this year was the marketing department, and what it learned is that in 2025 and 2026, "agent" became the most profitable word in enterprise software to misuse. This lesson teaches you to take the costume off before the price tag goes on.
The Most Profitable Word in Software
In June 2025, Gartner did something analysts rarely do: it called out an entire market for dressing up. Surveying the flood of vendors claiming "agentic AI," Gartner warned that agent washing, the rebranding of ordinary automation as autonomous agents, was widespread. Of the thousands of vendors claiming agentic capability at the time, only around 130 were judged to be the real thing. Sit with that ratio for a moment. If you took a vendor's "agent" claim at face value in that market, your base rate of being sold a costume was overwhelming. It was not a coin flip; it was closer to buying a lottery ticket where the prize is honesty.
The same Gartner research attached a forecast to the fallout: over 40 percent of agentic AI projects will be canceled by the end of 2027, driven by escalating costs, unclear business value, and inadequate risk controls. A healthy share of those corpses will be mislabeled purchases, projects that were scoped, priced, and governed as agents when the thing in the box was a rules engine with a press release. Meanwhile McKinsey's 2025 State of AI survey found 62 percent of organizations experimenting with or scaling agents, which tells you the demand side of this market is enormous and mostly unable to verify what it is buying. Enormous demand plus unverifiable claims is the exact weather in which washing thrives; it is why the word on the box drifted faster than the software inside it.
None of this means agents are fake as a category. Real agentic systems exist, and Level 3 of this program will teach you to govern them. What it means is that in the current market, the word "agent" on a proposal carries almost no information. The information is in the evidence underneath the word, and this lesson is the instrument for extracting it.
What an Agent Actually Is, and Four Things That Are Not
You met the honest definition in the terminology lesson (Chapter 1.2): an agent plans, uses tools, acts, and adjusts, all under bounded permissions. Here we slow that down into the four capabilities you will test for, because each one is separately fakeable and each one has a separate evidence demand.
An actual agent does four things. It perceives context: it reads the live state of the work (the document, the system records, the history) rather than receiving a pre-formatted input. It plans: given a goal, it composes a sequence of steps itself, and the sequence can differ from case to case because the cases differ. It uses tools: it calls systems (query the ERP, send the request, update the record) as steps inside its own plan, not as a single hardwired endpoint. It adapts: when step three fails or returns something unexpected, it revises the plan and tries another route instead of halting. And all four run inside explicit permissions: what it may read, what it may write, what requires a human. Remove any one of the four and you are looking at something else wearing the name.
What else could it be? Four honest species of automation, each useful, each cheaper, and each sold this year in an agent costume.
| What it actually is | How it works | The tell |
|---|---|---|
| Rules workflow | If-then chains built by hand: "if invoice over $10k, route to controller" | Every path was authored by a person; a case nobody anticipated has no path |
| RPA (robotic process automation) | Screen-level macro replay: recorded clicks and keystrokes against a user interface | Breaks when a button moves; "adaptation" means a consultant re-records the macro |
| Chatbot with function calls | A conversational model that can fire one predefined action per request | One-shot tool use, no planning loop: it answers and acts once, it does not pursue a goal |
| Scheduled script with an LLM step | A cron job or pipeline where one stage calls a language model, such as "summarize, then file" | The sequence is fixed in code; the model decorates the pipeline, it does not direct it |
Notice what the four have in common: in every one of them, the intelligence about how the work should flow lives in a human who configured it in advance. In an agent, some of that intelligence has moved into the system at runtime. That relocation is the entire difference, and it is the thing you must demand to see demonstrated, because it is invisible on a demo screen. A rules workflow and an agent look identical for the eight minutes of a scripted demo. They only diverge when reality goes off script, which is exactly the part of reality demos are built to exclude.
What the Mislabel Actually Costs You
A skeptic might ask: who cares what the vendor calls it, if the software does the job? The answer is that the label sets three prices, and you pay all three.
The premium for the word. Agent positioning commands agent pricing. Autonomy is sold as labor replacement ("a digital employee"), so it is priced against salaries instead of against software, and the delta can be several multiples over the same capability sold as workflow automation. When the capability underneath is a rules engine, you are paying the salary premium for the software price of work. The worked example below puts a dollar figure on this, and the figure is not small.
The governance you inherit for capabilities you did not get. This cost is subtler and usually missed. If your organization believes it has deployed an autonomous agent, it must govern one: approval gates, action logs, rollback paths, permission reviews, an oversight design for a system that acts on its own. That is real work by real committees. Design it for a product that is actually a fixed workflow, and you have built supervision for an employee who never shows up: review meetings for decisions the system cannot make, audit questions it cannot answer, an oversight RACI (responsible, accountable, consulted, informed chart) with a ghost in the R column. You carry agent-grade governance overhead while receiving automation-grade behavior, the worst quadrant available.
The roadmap betrayal. The most expensive cost arrives last. Buyers who suspect the product is not quite agentic are soothed with the promise: "it will learn to handle exceptions over time." An actual agent, or at least a vendor genuinely building one, might eventually make that true. A rules engine can never keep that promise, because handling a novel exception requires exactly the runtime planning and adaptation it lacks by construction; every new exception is another if-then branch a human must author, forever. Organizations plan headcount, process redesign, and business cases around autonomy that is permanently scheduled for next quarter. When it never lands, the project joins Gartner's 40 percent, and the process owner who anchored a staffing plan to a marketing sentence is the one explaining why.
The Five Costumes of Agent Washing
Washed products cluster into five recognizable costumes. Learn the silhouettes and you will spot them across a conference hall.
Costume one: the rebadged RPA suite. The oldest player in this theater. A screen-automation platform, often a very good one, reissues its brochure with "agentic" in the title and its bots renamed "agents." The tell is archaeological: check the product's documentation from two years ago. If the architecture pages describe the same recorder, the same orchestrator, the same selectors, the word changed and the software did not. Ask what the "agent" does when the target application's interface changes; if the answer involves a maintenance contract, you are looking at macros.
Costume two: the chatbot with three buttons. A conversational interface that can trigger a handful of predefined actions: check a status, open a ticket, send a template. Genuinely handy, and one function call away from being honestly described. The tell is the ceiling: ask it to do something that requires two of its actions in sequence with a decision in between, and it either cannot, or it asks you to be the planner. A system where the human supplies the plan and the software supplies single actions is a remote control, not an agent.
Costume three: the workflow builder with an LLM node. A drag-and-drop pipeline tool where one of the twenty node types now calls a language model. The marketing shows the LLM node glowing in the center like a brain. The tell is authorship: who drew the flowchart? If the answer is "your implementation team, during onboarding," then the plan lives in the diagram, not in the model, and the product is a workflow suite with a clever step. Perfectly good products live here; agent prices do not belong here.
Costume four: the "autonomous" tool with a quiet human in the loop. The most deceptive costume, because the outputs look adaptive: odd inputs get handled, exceptions get resolved, the system seems to cope with novelty. It copes because a human, often in a vendor operations center in another time zone, is resolving the queue by hand. The tell is physics: humans have working hours and latency. Ask about processing during weekends and holidays; ask why "autonomous" resolution takes four hours at 2 a.m. and four minutes at 2 p.m. If the service-level agreement quietly says "next business day" for anything unusual, you have found the person behind the curtain. (You will see this exact discovery in the worked example.)
Costume five: the roadmap agent. The product sold today is honestly described automation; the agent is in the deck, arriving next quarter. Next quarter, it is arriving next quarter. The tell is the contract: the price is charged in the present tense while the capability lives in the future tense. The rule from your ROI lesson applies with no modification: you buy what ships today at a price justified by what ships today, and roadmaps are priced at zero.
The Undressing Questions
Five questions strip any costume, and none of them requires a technical background. They work because each targets a capability that cannot be faked in evidence, only in adjectives. Ask them in the meeting, in writing, and in the reference call.
"Show me the plan it made yesterday." A real agent composes plans at runtime, and real engineering teams log those plans, because they could not debug the system otherwise. So ask for yesterday's: the actual sequence of steps the system chose for an actual case, on screen, with timestamps. A washed product has no such log to show, because there was no plan to record; there was a flowchart, and the flowchart is the same for every case. If what you are shown is the configuration diagram rather than a per-case trace, you have your answer.
"What happens when step three fails?" Pick any mid-sequence step and make it fail in your head: the ERP times out, the lookup returns two matches, the document is missing a page. An agent revises: retries another route, gathers what is missing, or escalates with a specific description of what it could not do. Automation halts: an error queue, a ticket, a human. Both are acceptable behaviors; only one of them may be priced as autonomy. Vendors answer this question with their architecture whether they intend to or not.
"What tools can it call, and who granted the permissions?" An agent's action space is a list: the systems it can read, the systems it can write, the operations it may perform, and the named person who approved each grant. A real agent vendor produces this list instantly, because their own security review demanded it before yours did. A washed product's answer collapses to one hardwired integration ("it writes to the ERP") with no permission model, because a fixed pipeline needs none. No list, no agent; and incidentally, no list also means nobody at the vendor has thought about what happens when it writes something wrong.
"Show me it handling an input type it was not configured for." This is the adaptation-versus-configuration test, and it is the one to run live. Bring your own weird case: the supplier invoice in a foreign format, the request that spans two categories, the document with the missing field. Configuration handles what it was configured for; adaptation handles the neighborhood around it. Watch what happens at the boundary. A graceful, specific escalation is honest automation behaving well. A confident wrong answer is worse than either. And a correct answer that arrives after a suspicious delay deserves the costume-four questions above.
"What did it do autonomously at your reference customer last month?" Not in the pilot, not in the demo environment: in production, last month, counted. How many cases end to end without human touch, how many escalated, how many reversed. A real deployment has these numbers because the customer's own governance demands them. If the reference call keeps returning to time saved on drafting rather than actions completed autonomously, the product assists; it does not act. Assistance is valuable. It is also not what the price assumed.
The sin is never the automation; it is the label, and the premium and the governance the label extracts.
The Artifact: The Agent Autopsy Card
This lesson's artifact is a one-page instrument you fill in during or immediately after any "agent" evaluation. It has three blocks. Completed, it tells you what the product is, what it is pretending to be, and what you should pay.
Block one: the four capabilities, each with an evidence demand and a pass line.
- Context. Evidence demand: show the system reading live state (records, history, documents) rather than receiving pre-formatted input. Pass: it pulled the context itself in front of you. Fail: inputs arrive shaped by an upstream template or a human.
- Planning. Evidence demand: yesterday's per-case plan log, with two different cases showing two different step sequences. Pass: inspectable, case-specific plans exist. Fail: one flowchart, authored at onboarding, applied to everything.
- Tool use. Evidence demand: the action-space list (readable systems, writable systems, allowed operations, named permission grantors). Pass: the list exists and the vendor can demonstrate two different tools called within one case. Fail: a single hardwired write, or no permission model at all.
- Adaptation. Evidence demand: a live unconfigured-input test plus the answer to "what happens when step three fails." Pass: revised plan or specific escalation, at machine latency, at any hour. Fail: halt-and-ticket, or resolution latency that tracks human working hours.
Block two: the five costumes and their tells, as a spotting checklist: rebadged RPA (same architecture docs, new adjective), chatbot with three buttons (cannot chain two actions with a decision between), workflow builder with an LLM node (your team drew the flowchart), quiet human in the loop (latency follows office hours; "next business day" for anything odd), roadmap agent (price in the present tense, capability in the future tense). Tick any costume you observed and note the evidence.
Block three: the price implication row. One line, filled in last: given the passes and fails above, does this product price as an agent or as automation? Then write both numbers: the quoted price, and the price of the nearest honestly labeled competitor (a workflow suite, an extraction pipeline, an RPA license) delivering the same demonstrated capability. The gap between those two numbers is what the word "agent" is charging you per year. Every negotiation in Chapter 4 runs on making that gap visible; the next lesson hands you the wider question set that does it across every claim a vendor makes, not just this one.
Score it simply: four passes in block one is an agent, and Level 3 Chapter 4 will teach you the governance it deserves. Zero to two passes is automation in costume: negotiate at automation prices and govern it as what it is. Three passes is the interesting case: an emerging agentic product worth watching, worth piloting, and worth pricing on the three capabilities it has rather than the four it claims.
The Autopsy of AVA: A Worked Example
Here is the card in action, in a composite scenario with realistic numbers. Quillon Ops, a fictional 1,400-person logistics firm, is evaluating "AVA: your autonomous accounts-payable agent." The pitch: AVA receives supplier invoices, validates them, resolves exceptions, and posts them to the ERP without human touch. The price: $240,000 a year, framed as "a fraction of the three AP clerks she replaces." The demo is gorgeous. The head of shared services, who has read this lesson, brings the card.
Planning. "Show me the plan AVA made yesterday." The vendor shares a screen, and it is a handsome seven-step diagram: receive, extract, match, validate, resolve, approve, post. The buyer asks the follow-up: show me two invoices from yesterday that took different paths. Silence, then honesty: the seven steps were configured at onboarding and every invoice walks the same road. There is no per-case plan because there is no planning. Fail.
Adaptation. "What happens when step three fails, when the purchase-order match comes back ambiguous?" The answer, read carefully from the support documentation: the invoice moves to an exception queue and a ticket is created. A ticket is a halt wearing a workflow. Fail.
Tool use. "What tools can AVA call, and who granted the permissions?" One integration: a templated write to the ERP's invoice-entry API. It is a real integration, competently built, with credentials managed properly. It is also a single hardwired endpoint, not an action space. Partial.
Context and the unconfigured input. The buyer runs the live test: a real supplier invoice in a layout AVA was never configured for. It resolves correctly, which is impressive, in four hours, which is strange for software. The buyer asks the physics question: what is resolution latency on weekends? A pause, and then the sentence that undresses the whole product: "anything non-standard received after Friday noon processes next business day." The adaptive layer is a human queue in the vendor's operations office. Costume four, confirmed in one calendar question. Fail.
The card's verdict: one partial, three fails, costume four ticked. AVA is not a fraud; it is a genuinely decent document-extraction-plus-workflow product with a competent ERP integration and a human exception desk, wearing an agent costume priced against clerk salaries. The buyer prices it as what it is: the nearest honestly labeled competitor, an extraction pipeline with workflow routing and no "agent" anywhere in its brochure, quotes $78,000 for comparable capability. Quillon offers $85,000, citing the card line by line, and after two rounds the vendor takes it rather than lose the logo. Annual saving against the costume price: $155,000. Just as valuable: Quillon now designs oversight for an extraction pipeline (accuracy sampling on the extraction step, a service-level agreement with penalties on the human exception queue, clear ownership of the ERP write), not the approval-gate-and-rollback apparatus an autonomous agent would demand. Right price, right governance, no ghost in the RACI.
One more turn of the lesson, and it is the calm one. Quillon did not walk away, because Quillon never actually needed an agent. Recall the four-species lesson: a stable, document-heavy, high-volume process is extraction-and-workflow territory, and a pipeline with one generation step is cheaper, more predictable, and far easier to govern than autonomy would be. Discovering that the product is "just automation" was good news at the right price. The failure mode this lesson exists to prevent is not buying automation; it is buying automation at autonomy prices and then governing a ghost. The autopsy is not a weapon for killing deals. It is an instrument for pricing and governing them correctly, and it saved Quillon $155,000 a year plus the standing weekly meeting an imaginary agent would have convened.
What to Do Monday Morning
The card becomes a skill the first time you run it on a live claim. The sequence:
- Inventory the word. List every product in your current stack or pipeline that carries "agent," "agentic," or "autonomous" in its name or renewal paperwork. Include the ones already installed; washing happens at renewal, not just at purchase.
- Build the card as a one-page document from block one, two, and three above. Ten minutes in a spreadsheet or a document template; it will serve you for years.
- Run block one on the nearest claim. Send the four evidence demands to the vendor in writing: yesterday's plan log, the step-three-failure behavior, the action-space list, and a live slot for your unconfigured input. Written answers are harder to costume than demos.
- Ask the physics question. For any product whose adaptive behavior impresses you, ask for resolution latency by hour of day and day of week. Software does not observe weekends; humans do.
- Price the gap. Find the nearest honestly labeled competitor for the same demonstrated capability and write both numbers on the card. Take the gap, not the adjectives, into the negotiation.
- File the completed card in your readiness portfolio next to your ROI worksheet from the previous lesson. Chapter 4's capstone question set, next lesson, assumes you have both in hand.
Key Takeaways
- Treat the word "agent" as carrying zero information in the current market: Gartner judged only around 130 of the thousands of vendors claiming agentic capability to be real, and predicts over 40 percent of agentic AI projects canceled by end of 2027.
- Test for the four capabilities that define an actual agent: perceiving live context, composing per-case plans, calling multiple tools inside its own plan, and adapting when a step fails, all under explicit, granted permissions.
- Distinguish the four honest species sold in costume: rules workflows, RPA macro replay, chatbots with one-shot function calls, and scheduled scripts with an LLM step; in all four, the flow intelligence was authored by a human in advance.
- Count all three costs of the mislabel: the salary-anchored price premium for the word, the agent-grade governance you build for capabilities you never received, and the roadmap promise ("it will learn to handle exceptions") that a rules engine is constitutionally unable to keep.
- Memorize the five costumes and their tells: rebadged RPA, the chatbot with three buttons, the workflow builder with an LLM node, the quiet human in the loop (latency follows office hours), and the roadmap agent (present-tense price, future-tense capability).
- Ask the five undressing questions in writing: yesterday's plan, the step-three failure, the tool and permission list, the unconfigured-input test, and last month's autonomous production numbers at a reference customer.
- Run the Agent Autopsy Card on every agent claim: four capability pass/fail lines, the costume checklist, and the price implication row that states the annual cost of the word itself.
- Remember that discovering "it is just automation" is often good news: automation is cheaper, more predictable, and easier to govern; the sin is the label and the premium, never the automation.
Skill.re