←
AI Readiness & Process Transformation
Aware · M1 · lesson 1 of 25 · in progress
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Adoption Without Transformation: The Usage Trap

15 min

The slide is titled "AI Momentum" and it is, by any graphic-design standard, gorgeous. A line climbs from the lower left to the upper right: monthly active users, up 340 percent since launch. The transformation lead presents it to the executive committee, the committee nods, and the program is declared a success on the strength of a chart that measures, when you strip it to its bones, how often people opened a window. Six floors down, in the same building, the claims backlog is the same length it was in January. The overtime line is the same. The cost per claim is the same. Nothing that appears on a financial statement has moved at all, and nobody in the room with the beautiful slide knows it, because nothing in the room measures it. This lesson is about that gap: the difference between metrics that measure motion and metrics that measure change, why every incentive in the building conspires to show you the first kind, and how to build the one-page instrument that forces the second kind into the open.

Two Species of Metric, and Why Confusing Them Is Fatal

In the first lesson of this chapter you met MIT's phrase for the most seductive failure pattern in enterprise AI: adoption without transformation. Pilots with rising usage charts and zero business impact made up a large share of the 95 percent that produced no measurable return. We are not going to re-litigate that finding here. We are going to do something more useful: take the phrase apart operationally, so that you can spot the pattern in a dashboard in under a minute and, more importantly, so you can never be the person presenting the gorgeous slide.

Start with a distinction that sounds obvious and is violated in most AI status decks you will ever see. An activity metric measures interaction with a tool: logins, seats provisioned, weekly active users, sessions, prompts submitted, drafts generated, documents "touched by AI." A value metric measures a change in a business outcome: hours redeployed from a process, cycle time cut against a baseline, error rate reduced, cost per transaction lowered, revenue per head raised, a queue cleared faster, a cost line that drops. These are not points on a spectrum. They are different species. One counts motion; the other counts consequences. And the profit-and-loss statement, which is where your CFO lives, responds to exactly one of them.

The confusion between the two is not a niche bookkeeping error. It is arguably the central measurement failure of the enterprise AI era. McKinsey's State of AI research found that 88 percent of organizations now use AI regularly somewhere, yet only about 39 percent can attribute any EBIT impact to it, and among those, most put the impact under 5 percent of earnings. Read that pairing slowly. Use is nearly universal; attributable value is a minority experience, and usually a small one. If usage produced value by itself, those two numbers would be close together. They are not close together. The distance between 88 and 39 is the usage trap, measured at global scale.

A plain-language analogy, because this distinction has to become reflexive. A gym measures swipe-ins at the front desk. Swipe-ins are real data, honestly collected, and they rise every January. But the swipe-in count tells you nothing about whether anyone got stronger, and every gym owner knows members who swipe in, walk the treadmill for eleven minutes while watching a screen, and leave, unchanged, for years. Activity metrics are swipe-ins. Value metrics are what the scale and the blood panel say. An organization that manages its AI program on activity metrics is a gym that congratulates itself on foot traffic while its members' health goes nowhere, except that the gym at least collects membership fees on the delusion, and your organization pays for it.

A Field Guide to the Two Families

You need to recognize both families on sight, because status decks rarely label their metrics honestly. Here is the taxonomy, and the substitution table you will use for the rest of your career: for each activity metric that dominates dashboards, the value metric that should replace it.

Activity metric (what the dashboard shows)Value metric (what should replace it)
Weekly active users, logins, seats provisionedHours redeployed per month from a named process, verified against staffing or timesheet data
Prompts submitted, sessions per userCycle time per unit of work, measured against a pre-launch baseline
Drafts generated, documents producedError or rework rate on outputs that actually reached a customer or a system of record
Documents or tickets "touched by AI"Cost per transaction on the process the AI touches, quarter over quarter
Training sessions completed, champions certifiedThroughput per full-time equivalent on the affected team, before versus after
Self-reported time savings from user surveysVerified time-per-task samples: timestamps or observed measurements on a real work sample

Three things to notice about this table. First, every entry in the left column is easier to collect than its counterpart on the right. The tool's own admin console produces the left column automatically, for free, every day. The right column requires you to instrument a process: capture a baseline before launch, define the unit of work, pull timestamps from the system where work actually lives. That asymmetry of effort is half the explanation for why dashboards look the way they do. People measure what is lying around.

Second, every entry in the left column almost always goes up in the first months of a deployment, regardless of whether the deployment is creating or destroying value. Novelty drives usage. Mandates drive usage. Curiosity, demo-chasing, and one enthusiast generating five drafts per ticket drive usage. A rising left-column chart is compatible with a pilot that is actively making a process slower, and you saw exactly that arithmetic in lesson one, where a tool that cost 90 seconds of pasting to save 60 seconds of writing produced a rising usage chart and a net loss on every single ticket.

Third, and this is the subtle one: the right column can go down in a healthy deployment's early weeks. Cycle time often rises briefly while a team learns a new step, verification is tightened, and the process map is redrawn. An organization managing on the left column sees early success followed by mysterious stagnation. An organization managing on the right column sees an honest J-curve: a dip, then a climb that shows up in a cost line. Only one of those organizations knows what is actually happening.

The honest middle layer: leading indicators done right

Between activity and value there is a legitimate intermediate layer, and you should know how to use it without being fooled by it, because value metrics are lagging by nature. A cost line takes a quarter to move; you need something to steer by in week three. The honest leading indicators are miniature value measurements: verified time-per-task samples (take twenty real tasks, measure them with timestamps before and after, compare); first-pass yield (what fraction of AI outputs are used without substantive correction, checked against edit logs rather than opinion); rework capture (of the outputs that reached the next step, how many bounced back). Each is small, cheap, and pointed at consequence rather than motion.

Contrast these with the most popular pseudo-value metric in the industry: the self-reported time-savings survey. "How much time does the tool save you per week?" is a question that produces answers shaped by loyalty, optimism, sunk cost, and the desire to keep a pleasant tool. People are poor estimators of their own time even when they have no stake in the answer, and here they have a stake. Treat self-reported savings the way you would treat a vendor benchmark: a claim to verify, never a number to book. If a team reports saving two hours a week, run a verified sample. Sometimes it confirms. Frequently the verified figure is a fraction of the reported one, and occasionally it is negative. The gap between reported and verified is itself a diagnostic: the bigger it is, the deeper the organization has fallen into the trap.

The Political Economy of the Usage Chart

If activity metrics were merely a measurement mistake, a memo would fix them. They persist because they are not a mistake; they are an equilibrium. Every party at the table has a reason to prefer the usage chart, and understanding those reasons is what lets you break the equilibrium without making enemies.

The vendor wants renewals. A vendor's revenue depends on the license renewing, and the renewal conversation is easiest when the slide shows growth. Usage grows in almost every deployment for a while, so usage is what the vendor's quarterly business review leads with. This is not fraud; it is selection. The vendor reports the numbers the vendor has, and the vendor has telemetry, not your cost per transaction. If you want the renewal meeting to be about value, you have to bring the value number yourself, because the other side of the table structurally cannot.

The sponsor wants a win. The executive who championed the initiative has reputation staked on it. A value measurement can come back negative; a usage chart almost never does. When quarterly review arrives and the choice is between presenting "monthly active users up 340 percent" and presenting "we have not yet instrumented the process, so impact is unknown," the sponsor's incentive is not subtle. Note what this means: the absence of a baseline is not always an oversight. Sometimes it is a shield, unconsciously or otherwise. A pilot that measures nothing can never be proven to have failed.

The team wants to look modern. For the people using the tool, high engagement numbers signal adaptability in a labor market that prizes it, and no individual contributor benefits from volunteering that the mandated tool costs them time. So usage stays respectable even where enthusiasm has died, propped up by open tabs and minimum-compliance prompts.

And nobody's job description says "demand the value number." That is the vacuum this certification trains you to fill. The finance team assumes the transformation office is measuring impact; the transformation office reports what the vendor dashboard exports; the vendor exports activity. The value question, "what changed in a business number, against what baseline?", belongs to no one by default, which is why the person who asks it calmly and repeatedly acquires influence far out of proportion to their title. You are not attacking anyone when you ask it. You are doing the job that the org chart forgot to assign.

Usage is what a tool does to a dashboard. Value is what it does to a P&L. Never let anyone, including yourself, present the first as evidence of the second.

What Transformation Actually Looks Like in Numbers

So what does the real thing look like? If adoption without transformation is usage charts over an unchanged business, then transformation has a signature too, and it is almost embarrassingly concrete. Transformation looks like a staffing chart that changes shape: the four people who spent mornings on invoice matching now spend them on supplier negotiation, and the timesheet codes prove it. It looks like a queue that clears faster: the claims backlog that sat at nineteen days is at eleven, and the intake volume did not drop, so the difference is real throughput. It looks like a cost line that drops: cost per handled ticket, cost per processed invoice, cost per underwritten policy, moving quarter over quarter in a finance system, not in a survey. It looks like an error rate falling on outputs that reach customers, measured by the complaints, returns, or rework tickets that stopped arriving.

Notice what every item on that list has in common: none of them can be read off the AI tool's own dashboard. They live in the operational and financial systems that surround the work, which is exactly why they are harder to collect and exactly why they are trustworthy. A tool can inflate its own telemetry just by existing. It cannot inflate your general ledger.

Notice also what has to happen upstream for any of these numbers to move. Someone redrew the process so the AI step has clean inputs and a defined output destination. Someone decided what the freed hours would be redeployed into, because hours that are freed but not redeployed evaporate into slack and appear on no financial statement. Someone changed a staffing plan, a service-level agreement, a queue discipline. This is why McKinsey's research keeps landing on the same lever: the high performers, roughly 6 percent of organizations in their survey, are about three times more likely than the rest to have fundamentally redesigned workflows around AI, and workflow redesign shows up among the strongest drivers of EBIT impact. The redesign is not a nice-to-have on top of the tool. The redesign is the mechanism through which usage becomes value. Skip it, and the tool's output pools uselessly at the point of generation, like a fast oven in a kitchen where the tickets are still shouted.

One caution to keep your ambition honest: value metrics move on operational timescales, not demo timescales. A cost per transaction is a quarterly number. Hours redeployed require a staffing decision that may take a planning cycle. This lag is exactly why the dishonest shortcut (report usage now, promise value later) is so tempting and so habit-forming: "later" has a way of receding indefinitely. The honest response to the lag is not to lower the bar; it is the leading-indicator layer from the previous section, with a committed date on which the lagging value number will be read.

A Tale of Two Pilots: A Worked Example

Here is the whole lesson compressed into one company. The company is hypothetical, a composite with realistic figures, but you will recognize it; most large organizations are running both of these pilots right now, and only one of them under this description.

A mid-sized insurer, call it Meridian Mutual, about 2,400 employees, launches two AI pilots in the same quarter. Pilot One is a general-purpose AI assistant rolled out to 600 knowledge workers across marketing, HR, and corporate functions. Pilot Two is narrower and duller: an AI step inside claims handling that assembles the claim file summary and drafts the coverage-position letter for an adjuster to verify.

Pilot One is a phenomenon. By month four it reports 30,000 prompts per month, 78 percent monthly active users, and a highlight reel of testimonials. It wins the company's internal innovation award in November. The steering deck's headline chart is prompt volume, and prompt volume is magnificent. What the deck does not contain is a baseline for any process the assistant touches, because the assistant touches everything and therefore, measurably, nothing. Asked about impact, the program manager cites a survey: users self-report saving "around three hours a week." Nobody verifies a single hour of it against a timesheet, a queue, or a deliverable. Run the arithmetic that nobody in the room runs: 600 users self-reporting three hours a week is roughly 93,000 claimed hours a year, the equivalent of about 45 full-time employees. If that capacity actually existed, somewhere a backlog would be shrinking, an overtime line would be falling, a hiring request would be quietly withdrawn. Finance can find no such trace. The claimed hours are ghosts: real feelings, unverifiable as capacity, redeployed into nothing.

Pilot Two spends its first month doing something that looks like inactivity: measuring. The team pulls six months of timestamps from the claims platform and establishes a baseline: average handle time of 47 minutes per standard claim across 11,000 claims per month, with a first-pass accuracy issue costing rework on about 8 percent of files. They redraw the swimlane so the AI summary step receives a complete file (the upstream document checklist gets fixed first, which takes three unglamorous weeks) and so every AI-drafted letter routes through a named adjuster's verification with a two-minute review standard. Then they turn it on for one claims team, and they read the same timestamps they baselined.

By month five, verified handle time on the pilot team is 38 minutes: 9 minutes saved per claim against a measured baseline. Scaled to the full 11,000 claims a month, that is 99,000 minutes, or 1,650 hours per month of capacity, roughly the workload of ten full-time adjusters. And because the operations director committed in advance to what the hours would become, they do not evaporate: the chronic backlog burns down from nineteen days to eleven over two quarters, the overtime budget line drops by about $31,000 a month, and two open adjuster requisitions are closed unfilled. At a loaded cost of $52 an hour, the redeployed capacity is worth about $85,800 a month, call it a million a year, and the CFO does not have to take anyone's word for it, because every number in the chain (baseline timestamps, current timestamps, claims volume, overtime line, requisitions) lives in a system finance can query.

Now the epilogue, one year later. Pilot One's renewal comes up: roughly $480,000 for the enterprise license. The CFO, newly attentive, asks the avoided question, and the answer is a usage chart and a survey. The renewal is cut to a fraction of the seats, the program manager updates a résumé that says "led award-winning AI adoption program," and the organization files the experience, unfairly but predictably, under "AI didn't really work here." Pilot Two is scaled to all claims teams, its measurement scaffolding goes with it, and its operations director is handed the next two processes and a bigger title. Same company, same year, same underlying model family. The difference was never the AI. The difference was that one pilot was built to produce a value number and the other was built to produce a chart.

The Artifact: The One-Page Value-Metric Worksheet

The instrument that separates the two fates fits on one page, and it is this lesson's deliverable for your readiness portfolio. Six columns. Any AI initiative, proposed or running, gets a row. If a row cannot be completed, that incompleteness is the finding.

Process touched → Baseline number → Value metric → Measurement method → Owner → Review date.

Walk through the columns once, slowly, and then we will fill a real row.

  • Process touched. One named process, specific enough to have a queue, a cycle time, or a cost line. "Claims handling for standard auto claims" qualifies. "Employee productivity" does not; a worksheet row that names no process is the trap in written form.
  • Baseline number. The pre-launch measurement of that process, with its source and date. If the initiative is already running and no baseline was captured, write "none" in this cell and let it be seen; an honest "none" is worth more than a reconstructed guess, and it tells you the first repair.
  • Value metric. The single business number this initiative is supposed to move, chosen from the right-hand column of this lesson's table: hours redeployed, cycle time, error rate, cost per transaction, revenue per head. One metric. An initiative that claims seven value metrics is usually hiding behind the fog of many.
  • Measurement method. Where the number comes from and how it is verified: system timestamps, ledger lines, sampled observation. If this cell says "user survey," go back to the leading-indicators section and upgrade it. The method must be one a skeptical finance partner could rerun without your help.
  • Owner. One name. A person, not a committee, accountable for the value metric moving, which is a different job from being accountable for the tool existing.
  • Review date. The date on which the number will be read and a scale, iterate, or kill decision made. A worksheet row without a review date is a zombie pilot with better paperwork.

Now fill it in for a live example: an accounts-payable team piloting AI-assisted invoice-exception resolution. Process touched: supplier invoice exception handling, the queue of invoices that fail automatic matching. Baseline number: 22 minutes average resolution time per exception, 3,400 exceptions per month, measured from workflow-system timestamps over the four weeks ending March 28. Value metric: average resolution minutes per exception, with monthly hours redeployed as the derived figure (every minute saved is 3,400 minutes, about 57 hours, per month). Measurement method: the same workflow timestamps as the baseline, read monthly, plus a quarterly verified sample of 40 exceptions reviewed end-to-end to confirm the timestamps still mean what they meant. Owner: the AP team lead, by name, in writing. Review date: first business day of each month for the read; day 90 for the pre-committed scale-or-kill decision. Six cells, maybe an hour of work to complete honestly, and this initiative can now never drift into the usage trap unnoticed, because the trap requires exactly the ambiguity these six cells eliminate.

Two usage notes from the field. First, the worksheet is diagnostic before it is managerial: run it against every AI initiative your organization currently has, and count how many rows you can complete. Most first passes complete fewer than half, and the empty cells map your organization's measurement debt more accurately than any maturity assessment you could commission. Second, the worksheet is politically gentle by design. You are not accusing anyone's pilot of failing; you are asking six administrative questions on one page. The equilibrium from the political-economy section survives confrontation easily. It survives a standing one-page form much less well.

What to Do Monday Morning

This lesson converts into practice in a single working session. Here is the sequence.

  1. Draw the worksheet. Six columns on one page: process touched, baseline number, value metric, measurement method, owner, review date. This takes five minutes and outlasts every tool it will ever evaluate.
  2. List every AI initiative you can see, official or shadow, and give each one a row. Do not skip the beloved ones; the beloved ones are where the trap lives.
  3. Fill in what is known and leave the rest visibly blank. The blanks are the report. Count complete rows versus total rows and date the count; that ratio is your organization's usage-trap exposure, and you will want the before picture later.
  4. Sort the flagship initiative's current metrics into the two families using this lesson's table: activity on the left, value on the right. If the right side is empty, you have found your first repair, and it is a measurement repair, not a technology one.
  5. Replace one survey with one verified sample. Pick the most-cited self-reported savings claim and run twenty timestamped tasks against it. Whatever the result, you now own the most credible number in the room.
  6. Put a review date on one dateless pilot. Propose it as housekeeping: "when do we read the number and decide?" Watch carefully who resists a date; resistance to a review date is the usage trap defending itself, and it tells you where to go to work.

Key Takeaways

  • Separate the two species without mercy: activity metrics count motion (logins, seats, sessions, prompts, drafts, AI-touched documents) while value metrics count consequences (hours redeployed, cycle time cut, error rate reduced, cost per transaction, revenue per head), and only value metrics reach a P&L.
  • Distrust the shape of every dashboard by default: activity metrics dominate because they are free to collect, they rise almost regardless of outcomes, and they flatter everyone at the table.
  • Name the equilibrium before you try to break it: the vendor needs renewals, the sponsor needs a win, the team needs to look modern, and nobody's job description demands the value number until you make it yours.
  • Recognize transformation by its physical evidence: a staffing chart changes shape, a queue clears faster, an overtime or cost line drops in a system finance can query, none of which can be read off the tool's own telemetry.
  • Use leading indicators honestly while the lagging value number matures: verified time-per-task samples, first-pass yield from edit logs, and rework capture, never self-reported time savings, which are claims to verify rather than numbers to book.
  • Anchor on the corroborating arithmetic: 88 percent of organizations use AI while only about 39 percent see any EBIT impact, mostly under 5 percent, and McKinsey's high performers are about three times more likely to have fundamentally redesigned workflows, because redesign is the mechanism that turns usage into value.
  • Deploy the one-page value-metric worksheet on every initiative: process touched, baseline number, value metric, measurement method, owner, review date, and treat every cell you cannot fill as a finding.
  • Insist that freed hours become redeployed hours: capacity that is not consciously reassigned to a queue, a staffing plan, or a growth task evaporates into slack and appears on no financial statement, no matter what the survey says.