←
AI Readiness & Process Transformation
Visionary · M3 · lesson 3 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Beyond Efficiency: AI for New Products and Models

15 min

Three years in, the transformation is working. Forty-one use cases live, a verified benefit number that survived finance, and a fifth-in-a-row workflow that reached measured value in nine weeks instead of thirty-one. Then, at the November board session, a director who has been quiet for two years puts down her pen and asks a nine-word question: "This is all cost. Where is the growth?" The transformation lead has an excellent answer for cost per transaction and no answer at all for this. What makes the moment dangerous is not that the question is unfair. It is entirely fair, and it arrives with a trap attached: the instruments that earned her credibility over three years are the wrong instruments for the thing she is now being asked about, and using them anyway is the most common way good growth ideas get strangled by their own sponsors.

The Question That Arrives in Year Three

Every enterprise AI program that survives long enough gets this question, and it is legitimate. Cost programs have a floor: you take forty percent out of an exceptions queue exactly once, and the second forty percent does not exist. A leader whose entire portfolio is efficiency will eventually present smaller numbers with more effort, which is a slow way to lose a mandate. And if AI genuinely changes what an organization can offer, confining it to the back office is a failure of imagination a competitor may not share.

So the answer is not "growth is not my job." It is to go there deliberately, with the right instruments. Here is the trap. The playbook this program spent five levels teaching you (baseline the process, measure the delta, gate the spend, verify the output, kill what fails) rests on a foundation revenue-side work does not have: a stable existing process with measurable current performance. A new AI-enabled offering has no baseline. Its value is not a delta against a known cost; it is a hypothesis about what customers will choose to do. Its failure modes are commercial rather than operational: nobody wants it, or not at your price, or a competitor gets there first. None of those show up in a cycle-time chart.

This produces two symmetrical mistakes, and most leaders make one of them.

The first is applying the operations playbook unmodified. A product idea arrives; the leader, proud of the rigor she has built, asks for the business case: baseline, projected delta, payback period, confidence rating. The team cannot produce one, because the honest answers are "zero, this does not exist yet" and "we do not know within an order of magnitude." So they invent one, and a spreadsheet appears with a total addressable market, an adoption curve and a price point, all guesses wearing the costume of analysis. The idea is then killed for weak numbers (often the good ideas, whose sponsors were honest) or funded on fabricated ones (often the confident ones, whose sponsors were not). Rigor applied to the wrong object becomes theatre.

The second is abandoning discipline entirely. "Product is different, you cannot measure innovation, we have to place bets." True in a limited sense and catastrophic as an operating principle, because it funds enthusiasm. We know where that road goes, because MIT's 2025 GenAI Divide study measured it: enterprise generative AI budgets skewed heavily toward the visible front of the business, sales and marketing, while the clearest measurable returns sat in the unglamorous back office. The same study found 95 percent of enterprise generative AI pilots delivered no measurable profit-and-loss (P&L) return, and only about 5 percent of custom tools crossed from pilot into production. The money went where the story was; the return sat where the structure was. A leader who answers the growth question by loosening the discipline walks straight into the statistic that made this program necessary.

The honest middle is the rest of this lesson: what transfers, what does not, and how to evaluate revenue-side bets with rigor that fits their actual uncertainty.

What Carries Over Unchanged (More Than You Expect)

Start with the reassuring part, because leaders under growth pressure often assume they are starting from nothing. You are not. Four things you already own transfer without modification.

Readiness transfers completely

A customer-facing AI feature has the same readiness anatomy as an internal workflow. It needs data readiness: Gartner's finding that 63 percent of organizations lack or are unsure of AI-ready data practices does not become untrue because the output has a customer's name on it. It needs process readiness: something behind the feature has to fulfil what it promises, and if the promise triggers a manual back-office scramble you have built a demo with a customer attached. It needs people readiness in the places operations leaders forget, specifically support and sales, who will be asked about a capability nobody briefed them on. And it carries governance obligations that are higher, not lower, than internal AI.

That last point is the most common revenue-side failure: a beautiful capability that support cannot service and legal has not reviewed. The demo lands, the launch date is set, and two functions who could have flagged the problem in an hour were never in the room. When the EU Artificial Intelligence Act's transparency obligations for AI-generated content take effect on December 2, 2026, the asymmetry becomes legal rather than merely operational: an internal draft nobody labels is a housekeeping issue, while an unlabelled AI-generated output delivered to a customer in a regulated market is a compliance exposure with a date on it.

Verification architecture transfers

Everything Level 3 taught applies without a change of principle. A customer-visible AI output needs defined quality gates, a sampling regime, an escapes log (the record of defects that got past the gates and reached the recipient), and a named owner for each. What changes is the rate, because an escape costs more when the recipient is a customer rather than a colleague.

Experiment discipline transfers, and it is the right instrument here

Earlier in this chapter you built the cheap, time-boxed, pre-committed experiment: a question stated in advance, the smallest instrument that can answer it, a decision rule written before the data arrives. That is the best-fitting instrument you own for growth work, because revenue-side questions are precisely the kind an experiment answers and a business case cannot.

Kill discipline transfers with more force, not less

The previous lesson made killing projects an enterprise capability rather than an act of courage. Growth bets need it more, because commercial bets fail more often. A low kill rate in a portfolio of revenue-side bets is not evidence of good selection; it is evidence that nothing is being tested hard enough to fail. The shape of a doomed product bet is rarely a loud failure. It is a quiet extension: another quarter, another feature, another segment, because nobody wrote down in advance what would constitute an answer.

What Does Not Transfer, and Why It Matters

Now the harder half. Three things you rely on daily do not survive the crossing, and a leader who does not name them keeps reaching for tools that no longer work.

One: the baseline is gone

On the operations side the baseline founds everything: fourteen minutes per case, 6.2 percent error rate, 4,100 cases a month, measured before you touched anything. The delta is real because both numbers are real. A new offering has no current performance to measure, because there is no current thing. The comparison is not against a measured past but against a counterfactual future: what would have happened had we not done this, which nobody has ever observed. Any "return on investment" (ROI) calculation built on that comparison is not a measurement; it is an assumption stack with a decimal point.

The instruction is unusually specific: refuse to manufacture one, and say so in those words. When the chief financial officer (CFO) asks for the business case on a growth bet, the answer is not a fabricated model and not a shrug. It is this: "I can build you a five-year revenue model. It will be arithmetic on four numbers I invented, and its output will be whatever I choose the adoption rate to be. What I can give you instead is the list of things that must be true for this to work, ranked by how uncertain and how consequential each one is, the cheapest test for the most dangerous one, and what that test costs. That I can defend."

Leaders fear that conversation, and it is almost always the credibility-building move: finance professionals read models built on invented assumptions for a living. It also protects what you spent three years earning, because if you sign a revenue projection that misses by 80 percent, every verified operations number you ever produced gets re-read in that light.

Two: the value mechanism changes

Operations value is cost avoided. Once the delta is verified the value is near-certain, because it depends only on your own organization continuing to do what you have already watched it do. Nobody outside the building has to agree.

Revenue value depends on customers choosing something, and that is uncertain in a way no amount of internal measurement resolves. You can run twenty internal reviews, poll every executive and model the market beautifully without moving one millimetre closer to knowing whether anyone will buy it. The only evidence that counts is customer behaviour, which reorients the whole evaluation approach around one goal: get behavioural evidence as cheaply and as early as possible. Not opinion, which is free and worth what it costs. Behaviour: a signed contract, a paid pilot, a renewal, usage that persists past novelty, a purchase order for something the customer previously paid a person to do. Every euro spent building before that evidence exists is a euro spent on a belief.

Three: the failure profile changes

Operations bets fail slowly and visibly. Extraction accuracy plateaus, the exception rate climbs, and your stage gates catch it because the metric you watch is the metric that matters. You get months of warning. Product bets fail on adoption, on willingness to pay, or on a competitor's move, and can look healthy internally right up until the market answers. The build is on schedule, the demos land, the team is proud, and then it ships and nothing happens. There was no leading indicator inside the building because the failure was never inside the building.

DimensionOperations AI (Levels 1 to 4)Revenue-side AI
BaselineMeasured current performanceNone; comparison is a counterfactual
Value mechanismCost avoided, near-certain once verifiedCustomer choice, uncertain until observed
Evidence that countsInternal measurementCustomer behaviour
Failure signalSlow, visible in your own metricsAbrupt, external, often post-launch
Cost of an errorRework, internalTrust, potentially public
Right early instrumentPilot with a baselineCheapest test of the riskiest belief

The Artifact: The Growth Bet Frame

Here is the lesson's deliverable. The Growth Bet Frame is four questions that replace the business case for revenue-side AI work, each with a defined evidence standard. It fits on one page and is harder to fake than a spreadsheet.

Question 1: What customer problem does this solve, and how do we know it is real?

Not "what does it do." What problem, whose, and what evidence shows it exists independently of our enthusiasm for solving it. Evidence standard: customer conversations with named accounts and dated notes; unprompted demand signals; support-ticket patterns; and the strongest signal of all, things customers already pay someone to do. A problem someone already pays to solve badly is validated; a problem that appears only in a strategy deck is a hypothesis.

One underrated source is sitting in your own systems. Your operations data is product insight. The exceptions your process handles are frequently problems your customers also have. If your team spends 340 hours a month reconciling delivery discrepancies against supplier documentation, your smaller customers probably spend proportionally more on the same problem with worse tools. The program that mapped your processes to find automation candidates has produced, as a by-product, a catalogue of expensive problems in your industry.

Question 2: Why is AI the reason this is now possible?

The mechanism question, sharpened for the revenue side. State what AI makes possible that was not: an economic threshold crossed (unit cost of categorising a document down two orders of magnitude), a capability that did not exist (unstructured input at volume), or a speed that changes the offer (minutes where the manual path took three days).

The filter, stated bluntly: if this offering would have been possible five years ago and nobody built it, AI is not the unlock and the idea needs an entirely different justification. It may still be a good idea, but then it is a market idea wearing an AI label and should be evaluated as one. This question removes most "AI-powered" repackaging, the internal cousin of the agent washing Gartner documented when it predicted over 40 percent of agentic AI projects would be cancelled by the end of 2027.

Question 3: What would we have to believe, and which belief is cheapest to test?

The heart of the frame, and the technique to take away even if you forget the rest. Write the assumption stack: everything that must be true for the bet to pay.

  • Customers want this: the problem is real and ranks high enough to act on.
  • They will pay approximately X, at a price that works for us.
  • We can deliver at cost Y, with unit economics that include verification effort.
  • It will not cannibalise Z, an existing revenue line at better margin.
  • We can support it: service, escalation and account management at the volume implied.
  • We can defend it: a competitor cannot trivially copy it within a quarter.

Then rank each belief by uncertainty multiplied by consequence. Uncertainty: how confident are we, honestly, on a one-to-five scale. Consequence: if this is false, does the bet die or merely get harder. The product gives a ranked list, and the top item is where the next money goes, tested with the cheapest instrument that can produce a real answer.

Two disciplines make this work. First, the cheapest test is almost never a build: it is a pricing conversation, a manual delivery to three customers by hand, a landing page with a real signup, a letter of intent. Second, the test must be capable of returning "no." A survey asking "would you find this valuable?" cannot, because saying yes to a hypothetical costs the respondent nothing. "Here is the price, would you sign a six-month pilot" can.

This replaces the business case at the early stage, and it is more rigorous than one, not less. A business case asserts a conclusion and buries its assumptions in a model; the assumption stack lists them on one page in rank order, riskiest first, each with a named test. Any executive can challenge it without opening a spreadsheet, which is why it is uncomfortable and why it works.

Question 4: What is the cost of being wrong, and can we bound it?

The answer is staged commitment: escalating gates, each with a spending cap and a pre-committed decision, so that being wrong costs the current stage rather than the whole idea.

StageWhat it testsTypical commitmentVerification bar
Concept testProblem reality and price reactionWeeks of a few people, no buildNot applicable, no customer output
Limited pilotDelivery and willingness to pay, real customers, labelledSmall, contracted, fixed termElevated; heavy sampling, named reviewer
Controlled launchRepeatability in one segmentCommitted team, capped scopeHigh; escapes tracked, support trained
General availabilityScale economics and support loadFull commercial commitmentHighest; disclosure, monitoring, audit

The addition specific to customer-facing AI, and the one operations leaders under-weight, is the asymmetric reputational exposure at each stage. An internal AI error costs rework: someone catches it, fixes it, logs it, and the process improves. A customer-facing AI error costs trust, which has no rework path, and it can become public in a way an internal defect never does, through a screenshot, a review, or a regulator's inbox.

The operating rule follows: the verification bar rises as the audience widens. If internal sampling runs at 5 percent of outputs, customer-facing sampling starts at three times that and stays there until the escape rate earns a reduction, stated in the stage plan before launch rather than discovered during an incident.

Operations value is a delta you can verify. Revenue value is a belief the market has to confirm. Never let the second sit inside the same total as the first.

Three Revenue-Side Patterns, and Where They Sit in the Portfolio

Not all growth bets carry the same risk, and treating them as one category is why leaders fund all of them or none. Three patterns, in ascending order of uncertainty.

Pattern 1: AI inside the existing offering

What you already sell gets measurably better: faster turnaround, fewer errors, responsiveness that was previously uneconomic. A distributor gives customers automated delivery-discrepancy resolution inside the existing account portal; a professional services firm cuts report turnaround from eight days to two.

Risk profile: lowest. You have customers, a relationship, a price and a renewal event, so measurement is comparatively clear: retention, expansion and win rates against real history. This is usually the right first move, and leaders skip it because it feels insufficiently visionary.

Pattern 2: AI as a new service line

You sell a capability the organization built for itself: your exceptions-handling method, your document-extraction pipeline, your verification standard, offered to organizations with the same problem and no capacity to build it.

Risk profile: middle, and this is the most underrated path for services, distribution and logistics businesses, for a structural reason: the readiness work already produced the asset. The process map, the standard operating procedure (SOP), the verification standard, the exception taxonomy, the measured performance record, the training material: those Level 3 and Level 4 artifacts are most of a service offering's operating manual, already written and already tested on the hardest customer you have (yourself), with real performance data rather than projections behind them.

The genuine risks are commercial, not technical. Selling is a capability you may not have, external service levels are harder than internal ones, and the internal customer forgave things an external one will not.

Pattern 3: the new business model

The offering changes how you make money: outcome-based pricing where you previously billed hours, a data product built on aggregated flows, a platform position rather than a service position. Risk profile: highest, longest horizon, most ways to be wrong. It most needs the assumption-stack treatment and the smallest possible first commitment, because the assumptions are numerous and correlated: if pricing fails, several others fail with it. Fund the first test, not the model.

The portfolio position and the reporting rule

Revenue bets belong in the deep-bet tier from the Level 4 portfolio lesson: capped as a share of capacity (a common shape is roughly 60 percent quick wins and core workflows, 25 percent scaling, 15 percent deep bets), governed by stage gates, expected to fail more often than the rest. The cap stops growth work from consuming the delivery capacity that produces your verified value; the floor stops it being crowded out every quarter by work with a cleaner business case.

Then the reporting rule, the most important sentence in this section: never let a speculative revenue projection sit in the same total as verified operations value. Report the bets in their own tier, in their own language, describing what has been learned rather than what is projected. The temptation to blend is enormous, because a slide reading "12.4 million of value" beats "3.1 million verified, plus three bets in stage two." Resist it completely. When the bet does not pay, and by construction most will not, a blended number does not lose the speculative portion; it loses the whole number's credibility, including the part you spent three years verifying. That is the confidence discipline from the Level 5 measurement lesson, applied where the temptation to abandon it is greatest.

The Growth Move, and the Product That Answered No Question

Both stories below are illustrative, with hypothetical numbers chosen to be realistic rather than reported.

The growth move: selling the capability you built

Return to the distribution business from earlier in this program. Its exceptions-handling capability is now mature: a documented method, a verification standard with defined sampling, an exception taxonomy refined over eighteen months, and measured performance (backlog down from nine weeks to four, cost per exception down roughly 44 percent). Asked the growth question, the transformation lead does not invent a product. She looks at what the program already built, and notes that smaller distributors have the same problem and no capacity to solve it. Pattern 2. She runs the four questions.

Question 1, the problem. Fourteen dated, documented conversations with smaller distributors, of which eleven describe the same pain in their own words and quantify it roughly. Crucially, three existing customers had already asked informally whether the company would handle their discrepancy queues, the strongest evidence in the set because nobody prompted it. Verdict: real.

Question 2, why now. The manual version of this service was always possible and the margin was always terrible, which is why nobody sold it. The economics work only because categorisation and extraction are automated, taking fully loaded cost per exception from around 11 dollars to roughly 3. A threshold crossed, not a label applied. Verdict: AI is the unlock.

Question 3, the assumption stack. Six beliefs, scored on uncertainty (1 to 5) times consequence (1 to 5): problem is real (2 x 5 = 10, already partly tested), willingness to pay at a viable price (5 x 5 = 25), unit economics hold at external service levels (3 x 5 = 15), no cannibalisation of the core relationship (2 x 3 = 6), support model exists (3 x 4 = 12), defensibility (4 x 2 = 8). Willingness to pay tops the list, and the cheapest test is neither a survey nor a build. It is nine concrete pricing conversations: a specific monthly fee, scope and term, put as a real offer to nine named accounts. Three said yes at the stated price, two countered lower, four declined with reasons worth reading. Cost: roughly three weeks of one person's time, against the most dangerous question in the stack.

Question 4, bounding the loss. A limited pilot with two customers under contract at a deliberately modest price, six-month term, with an explicit non-renewal expectation stated in the sales conversation so that stopping is not a relationship event. Total exposure decided in advance: two people part-time and about 90,000 dollars. The reputational bar is written into the pilot plan: sampling on customer-facing output at three times the internal rate, a named reviewer, a weekly escapes review, and AI involvement disclosed in the service description rather than buried.

Result at nine months. Two pilots ran; one renewed and expanded; one churned for a specific, learnable reason (their exception mix was 70 percent supplier-side rather than delivery-side, so the taxonomy transferred badly), which is a scoping lesson rather than a verdict on the offering. The decision is a controlled launch in one segment, pricing revised upward by about 20 percent and a qualification question added about exception mix. It is reported to the board in the deep-bet tier in its own language: "one bet at stage three of four, one renewal, one instructive churn, cumulative spend 90,000, next gate in Q2." Not one number of it appears inside the verified operations total.

The failure story: the AI product that answered no question

A mid-market software company of roughly 400 people is under pressure from its board and its competitors' press releases to have an AI story. In ten weeks it ships an AI assistant inside its product. It demos beautifully and the launch webinar performs well. Usage peaks in week three. By month four it is under 4 percent of accounts, most of that a handful of power users. The retrospective, which to the company's credit happens, finds four things.

  1. Nobody had established what customer problem it solved. No assumption stack was ever written. The feature answered a competitive question, not a customer one, and competitive necessity is a reason to investigate, never a reason to ship.
  2. Question 2 was never asked. Most of what the assistant did was a conversational wrapper on features customers already reached in two clicks.
  3. Support was never trained and quietly absorbed the load: roughly 180 extra tickets a month, improvised answers about a system nobody had briefed them on, never counted as a cost of the feature because nobody attributed it.
  4. One hallucinated answer to a customer's compliance question required a written apology from an executive and a contract concession. The internal equivalent would have cost an afternoon of rework.

The company's stated conclusion is that "customers are not ready for AI." The accurate conclusion is that it skipped every question this lesson asks, in the one area where skipping them costs most, then attributed the outcome to the market. This is MIT's finding reproduced inside a single product cycle: budget and attention flowed to the customer-facing side, where the story is, while measurable return sat in the back office, where nobody was looking. Hold it alongside the wider record: S&P Global found 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before, and McKinsey found 88 percent of organizations using AI regularly while only about 39 percent could attribute any earnings impact to it. The gap between activity and value is not smaller on the revenue side. It is larger, and louder.

What to Do Monday Morning

You almost certainly have a growth idea on your desk, or one arriving. Run it through the frame before it acquires a spreadsheet.

  1. Write the assumption stack for that idea. One page, six or so beliefs, each stated as a sentence that could be false. Score each on uncertainty times consequence and rank them. If you cannot write it, the idea is not yet an idea.
  2. Test the top belief with a customer conversation, not an analysis. This week, not next quarter. Ask something that can return a no: a real price, a real scope, a real term. Budget for the possibility that three weeks kills the idea, and treat that as the cheapest outcome available.
  3. Check whether your operations capability is itself sellable. List the internal capabilities that are mature, documented and measurably better than the industry norm, then ask whether smaller organizations have the same problem and no capacity to solve it. Pattern 2 usually hides in your own SOP library.
  4. Raise the verification bar on anything customer-facing. Find every AI output reaching a customer today, whether or not you sponsored it. Confirm sampling rate, named reviewer, escapes log and disclosure, set the customer-facing rate at a multiple of the internal one, and put December 2, 2026 in your governance calendar.
  5. Audit your value reporting for contamination. Check your latest value slide line by line for a speculative revenue number sitting inside a verified operations total. If one is there, separate it now, while doing so is a display of rigor rather than a response to an audit.

That closes Chapter 5.2. You now have the full engine: continuous discovery feeding a portfolio, experiments that answer questions cheaply, scaling that survives the enterprise, sunsets that free capacity, and growth bets evaluated with rigor that fits them. What it lacks is the governance and trust structure to run at enterprise scale, across functions, jurisdictions and regulators with dates on their calendars. That is Chapter 5.3.

Key Takeaways

  • Expect the growth question after two or three years of operations wins: efficiency has a floor, so the question is fair, but the operations playbook applied unmodified to a product bet kills good ideas with unanswerable ROI demands.
  • Carry forward what transfers: readiness across data, process and people; the verification architecture; experiment discipline; and kill discipline with more force, because commercial bets fail more often.
  • Recognise that governance obligations rise on the revenue side, since the classic failure is a beautiful capability support cannot service and legal has not reviewed, and the EU AI Act's AI-content transparency duty lands December 2, 2026.
  • Refuse to manufacture a baseline that does not exist, saying so to the CFO in those words and offering a ranked assumption stack and a cheap test instead of a model built on invented numbers.
  • Reorient evidence toward customer behaviour: operations value is cost avoided and near-certain once verified, while revenue value depends on customers choosing, which only a signed contract, a paid pilot or a renewal confirms.
  • Run the four questions of the Growth Bet Frame: is the problem real, is AI genuinely the unlock, what must be true and which belief is cheapest to test, and what is the bounded cost of being wrong.
  • Stage every customer-facing bet through concept test, limited pilot, controlled launch and general availability, raising the verification bar as the audience widens, because a customer-facing error costs trust while an internal one costs rework.
  • Sort growth work into the three patterns, start with AI inside the existing offering, look hard at the new service line because your Level 3 and Level 4 artifacts are most of a service manual already, and never blend a speculative revenue number into your verified total.