←
AI Readiness & Process Transformation
Visionary · M22 · lesson 22 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

The Use-Case Engine: Continuous Discovery to Scaled Value

15 min

The room is booked for two days, the catering is good, and by Thursday afternoon there are forty-one sticky notes on the wall beneath a banner reading AI IDEATION SUMMIT. The energy is genuine. A branch manager who has never been asked anything by head office describes a quoting bottleneck that has annoyed her for six years, and three people photograph her sticky note. Six candidates are selected on Friday, and the transformation lead goes home tired and happy, holding what feels like a year of work. What she is actually holding is five months of work, followed by seven months of an empty pipeline, a delivery team that is not empty, and a January invitation to next year's summit, sent to a room that will be slightly more cynical than this one. The summit was never the problem. It was also never the solution.

The Summit That Refilled Nothing

Follow that firm through the year, because the shape of what happens is the shape of what happens almost everywhere, and the calendar hides it beautifully.

Of the six chosen candidates, four eventually deliver something. Two die in the first eight weeks: one because the data it needed lived in a system the vendor contract did not cover, one because its sponsor moved divisions and nobody inherited it. Normal, healthy attrition. Delivery capacity is four to five concurrent efforts, so through February, March and April the team is comfortably full. In May, two efforts finish within a fortnight of each other and, for the first time since the summit, there is slack.

Here is where the invisible failure happens. The team does not sit idle, because teams never sit idle. They fill themselves. An operations director has been mentioning a contract-summarisation idea in every forum since the summit, where it did not make the cut, and by June it is being worked on: no baseline, no readiness screen, never assessed against anything. In August it fails on a data gap a competent triage conversation would have surfaced in an afternoon. The contracts exist as scanned images of varying quality, in three repositories, with no reliable metadata about which version is current. Four months of a two-person effort gone, and gone in a way that produces a story ("we tried AI on contracts, it did not work") quoted for years by people who never saw the data gap.

By October the pipeline is genuinely empty. Nothing has arrived from anywhere, because there is no anywhere for it to arrive from. Nobody has failed at their job. The transformation lead does the only visible thing available: she books the January summit. Attendance is lower, two functions send delegates instead of leaders, and in the back of the room a manager who contributed a good idea last year, heard nothing for eleven months, and eventually learned by accident that it had been declined decides not to bother with a sticky note this time.

Run the diagnosis as you would on any other business process. This firm did not have a discovery capability that underperformed. It had no discovery process at all: it had an annual event that produced a batch, and a batch is not a flow. A batch depletes on a schedule set by delivery capacity, not by the calendar, and when it depletes, the vacuum fills with whatever is loudest rather than whatever is best. The most expensive consequence is not the idle weeks or even the failed contract effort. It is that the organization now believes discovery is hard and rare, when what it experienced was discovery being absent seven months out of twelve.

Events produce batches. Batches deplete. Only a running engine produces flow, and flow is what a delivery capability eats.

The One Process Your Process People Never Mapped

Chapter 5.1 named discovery as one of the four standing capabilities of a converted operating model, alongside delivery, governance and value tracking, and left its depth for here. Level 4 taught portfolio construction: intake sources, the three-tier mix of quick wins, capability builds and deep bets, correlation flags, capacity realism. That is a planning exercise, run periodically, producing an allocation you defend for a couple of quarters. This lesson is the machine that runs between those planning moments. The portfolio is the snapshot. The engine is the film.

The reframe that organizes everything else is embarrassingly simple, and it is the reason this lesson exists. Your use-case pipeline is a process. It has inputs, stages, queues, handoffs, rework, a constraint, and an output rate. Every discipline you have applied to an order-to-cash process or a claims workflow applies to it without modification: baseline it, instrument it, find the bottleneck, fix the bottleneck, repeat. Yet in most organizations the pipeline is the single process the process people never mapped. Transformation leads who would never accept "we think it takes about a month" from a function they audit will say exactly that about their own intake, and nobody blinks.

The blind spot persists for three reasons worth naming, because naming them breaks them. Discovery feels creative, and creative work is culturally exempt from process discipline. Ideas arrive from people, so measuring their flow feels faintly rude. And the pipeline has no customer complaining about it: the function that submitted an idea nine weeks ago does not escalate, it simply stops submitting. A queue with no complaints and no measurements is not healthy. It is silent, and silence is the most common way a pipeline dies.

The stakes are the stakes of the whole program. MIT's 2025 GenAI Divide study found roughly 95 percent of enterprise generative AI pilots produced no measurable profit-and-loss return, and only about 5 percent of custom tools crossed from pilot into production. McKinsey's State of AI put 88 percent of organizations using AI regularly against only about 39 percent reporting any EBIT impact (earnings before interest and taxes, the profit line the CFO actually watches), with high performers, roughly 6 percent, about three times more likely to fundamentally redesign workflows. S&P Global found 42 percent of companies scrapped most of their AI initiatives in 2025; Gartner expects over 40 percent of agentic AI projects cancelled by end of 2027. Read those numbers as a throughput problem and they change character. The organizations in the 95 percent were not short of ideas; they were short of a system that could carry an idea from a hallway conversation to a scaled, owned, measured operation without it dying of neglect in a stage nobody owned. BCG's 10-20-70 arithmetic (10 percent algorithms, 20 percent technology and data, 70 percent people and process) says where transformation work lives, and a pipeline is 100 percent people and process.

So here is the artifact: the Use-Case Engine. Five stages with entry and exit criteria. A published work-in-progress limit per stage (WIP, borrowed from lean flow: the number of items allowed in a stage at once). A small set of throughput and cycle-time measurements. And standing inputs that keep the front end fed without anyone booking a room. It fits on one page, and once it exists, discovery stops being an event you schedule and becomes an operation you run.

The Five Stages of the Engine

Treat each stage as a workstation with a defined entry gate, a defined exit condition, and a characteristic way of failing. The failure modes matter as much as the criteria, because you will diagnose your own engine by symptom long before you diagnose it by measurement.

Stage one: intake, which is continuous and never a window

Intake is not a form that opens in Q3. It is a set of standing inputs producing candidates whether or not anyone is asking. Level 4 gave you four: the shadow-AI demand map (what people already do with unsanctioned tools, the most honest demand signal in the building), the friction streams from your improvement loops, the pain inventory held by process owners, and executive nominations. A mature engine adds two that only exist once you have a track record.

The first is the replication scan. Take every verified win and screen it against all the other functions for transferability. The invoice-exception triage that worked in finance: does distribution have an analogous exception class? Does service? These are the cheapest high-quality candidates in the enterprise, because the hard parts (the pattern, the prompt discipline, the verification design, the change approach, the governance path) are already paid for. They are also the most consistently missed, for one structural reason: nobody owns the scan. The function that built the win is busy operating it; the function that could reuse it does not know it exists. Assign the scan to a named person, put it in the calendar quarterly, and it takes about half a day.

The second is the capability watch. When a new technical capability becomes viable (longer context, reliable structured extraction from poor-quality scans, an agent pattern that finally holds up under audit), the standing question is not "what could we build with this?" It is "which of our parked and declined candidates just changed status?" That is answerable in an hour against a parked list, and it converts vendor noise into portfolio movement. Without it, a new capability triggers a reactive scramble in which someone invents a fresh use case to justify the excitement, which is how organizations end up with a solution shopping for a problem.

Entry criterion: a stated problem with a mechanism. Not "we should use AI in credit control" but "credit control spends about nine hours a week manually matching remittance advices to open invoices, and a matching step could remove most of that." The mechanism is the sentence saying how the improvement would physically happen. Exit: logged with a source, a submitter, a date, and a one-line mechanism. Characteristic failure: intake as a suggestion box with no criteria. It fills with wishes, the wishes get quietly ignored because there is no basis to act on them, and within two cycles the organization learns that submitting is pointless. A criteria-free suggestion box does not merely fail to produce candidates; it destroys the willingness to submit them.

Stage two: triage, which is fast, cheap and ruthless

Triage happens within days of intake, never weeks, and it is deliberately shallow. Two tests: does the stated mechanism survive five minutes of scrutiny, and does a quick readiness screen against your existing audit scores (data, process maturity, skills, governance) suggest this is plausible within a few quarters? That is it. Triage is not a business case, not a scoring workshop, and not a committee.

Three outcomes, and the middle one makes an engine humane and credible.

  • Advance to shaping, with a named shaper and a rough queue position.
  • Park with a named trigger: the candidate is good but early, and you write down the specific condition that would revive it. "Revive when the vendor-master remediation lands." "Revive when the service platform migration completes in Q3." "Revive when we have a governance pattern for customer-facing generation."
  • Decline with a reason: the mechanism does not hold, the value is trivial, or the process is being retired anyway. Say so in one sentence, to the submitter, by name.

Most candidates are not bad. They are early. A rejection tells a submitter their judgement was wrong; a park with a trigger tells them their judgement was right and the organization is not ready yet, which is a completely different message and relationship. It also creates an asset: a parked list with triggers partially manages itself, because when the vendor-master remediation lands, three candidates raise their hands automatically. Characteristic failure: triage that takes weeks. From the submitter's side a slow triage is indistinguishable from a black hole, and it produces the same behaviour: they stop submitting and start lobbying, which routes demand around your engine and into the ear of whoever will listen.

Stage three: shaping, the stage most engines simply do not have

Shaping is the pre-charter work that turns an advanced candidate into something a delivery pod can start on Monday: baseline established (current cycle time, volume, error rate and cost per unit, measured not estimated), data screened for existence and quality, sponsor confirmed in person rather than assumed, and the hypothesis sharpened into a testable statement with a target and a kill condition.

Almost every struggling engine is missing this stage, and the symptom is unmistakable: unshaped candidates arrive at governance gates and get held. The gate asks for a baseline, there is none, and four weeks evaporate while someone measures what should have been measured before the charter was written. Every one of those holds is charged to delivery capacity, your most expensive resource, to fix a defect a cheap upstream stage would have prevented. Shaping is a buffer in the manufacturing sense: a small stock of ready work that keeps the expensive downstream station running smoothly.

Be honest about its cost, because shaping is not free and it is very commonly the constraint. A shaped candidate represents real analyst effort, and it decays: shape too far ahead and priorities shift, sponsors move, systems change, and you have paid for baselines nobody will use. The discipline is to hold roughly one to two shaped candidates per delivery slot expected to open next quarter, no more. Entry: triage outcome of advance. Exit: a baseline document, a data screen result, a confirmed sponsor, and a draft charter with success and kill criteria. Characteristic failure: either no shaping stage at all (holds at the gate), or shaping treated as a full business case, in which case it becomes a second, slower triage and the queue backs up behind it.

Stage four: delivery, and the WIP limit you have to defend

Delivery is the pod running the method from Levels 2 and 3 under a charter and stage gates: baseline, redesign, controlled pilot, evidence, gate decisions. This lesson does not re-teach it. The engine view adds one thing, and it is the hardest thing here to actually do: a hard cap on concurrent deliveries, set below theoretical capacity, published, and defended.

The counterintuitive lesson comes from flow theory, and leaders reject it on instinct. A system running at 100 percent utilisation has no slack, and in any system with variability (every real programme has it: illness, escalations, a data surprise, a sponsor's reorganisation) queue times at full utilisation do not rise gently. They explode. So a delivery capability with a theoretical capacity of six concurrent efforts should publish a WIP limit of five. The leader who says yes to a sixth because "it is only one more" does not get six deliveries. They get four, later, because every effort now waits on shared people who are context-switching, every gate slips, and cycle time inflates across the whole board.

You will be asked to break that limit roughly monthly, usually by someone senior with a genuinely good idea. Publishing it is what makes it defensible: a limit in your head is an opinion, a limit on a page the executive committee has seen is a commitment. The right answer is never no. It is "yes, and here is the queue: it starts when the pricing effort exits in about five weeks, or sooner if you would like to tell me which current effort to stop." That converts an argument about your judgement into a decision about their priorities, which is where it belongs.

Stage five: scale and handoff, where the exit criterion is acceptance

The final stage is the transition from delivered to operated: the process owner takes it, the quality regime picks it up, the improvement loop starts, value tracking moves from project claim to standing measurement. Lesson 3 of this chapter owns scaling in depth, so take only the engine's contribution here, which is a sharp exit criterion.

The handoff must be accepted, not offered. Accepted means a named owner has agreed in writing, has been equipped (people trained, SOP updated, standard operating procedure being the documented way the work is done, KPI dashboard showing the new measures, exception path defined), and has taken the operating cost into their budget. Offered means the project team sent a handover pack and a calendar invitation.

Engines without this criterion accumulate orphans, and orphans are quietly poisonous. An orphaned use case still runs, still costs licence money, still touches customers, and has nobody accountable for its quality drift. Six months later it produces an error nobody notices for another month, and the resulting story is not "our handoff process failed" but "the AI got it wrong," which taxes every future candidate in the engine. A programme that scales nine use cases and orphans three has not scaled nine; it has scaled six and created three latent incidents.

StageEntry criterionExit criterionCharacteristic failure
IntakeStated problem with a mechanismLogged with source, submitter, dateSuggestion box with no criteria
TriageLogged candidateAdvance, park with trigger, or decline with reasonTakes weeks, reads as a black hole
ShapingTriage outcome: advanceBaseline, data screen, confirmed sponsor, draft charterAbsent (holds at gate) or bloated into a business case
DeliveryCharter approved, slot open under WIP limitGate decision: scale, iterate, or killWIP limit broken, cycle times inflate across the board
Scale and handoffGate decision to scaleHandoff accepted by an equipped, named ownerHandoff offered rather than accepted: orphans

Running the Engine by the Numbers

An engine you cannot read is an engine you cannot tune. Five measurements are enough, and more than five turns the thing into a reporting exercise. Read them the way a production manager reads a line: not one number at a time, but as a picture of where material is moving and where it is sitting.

Throughput

Candidates advanced to scaled and accepted per quarter. This is the engine's output rate and the only number that matters outside your team. Report it as a trailing figure, not a forecast. Two per quarter is a real programme, four is a strong one, and the absolute value matters far less than the trend and the fact that you know it at all.

Cycle time by stage

Elapsed time per candidate per stage, reported as a median so one pathological case does not distort the picture. This is the diagnostic instrument, and its power comes from a split most programmes never make: within each stage, separate waiting time from working time. A candidate sitting in shaping six weeks while a shaper touches it for eight days has one week of work and five weeks of queue. The queue almost always concentrates in shaping or governance gates, and knowing which is the entire difference between hiring people and fixing a process.

State it as a rule, because it is the most useful diagnostic here: if cycle time is dominated by waiting rather than working, adding capacity to the working stages does nothing. If gates meet monthly and a candidate misses one by two days, three weeks of waiting appear in your cycle time and no amount of analyst headcount removes them. The fix is a fortnightly gate slot or an out-of-cycle written approval path, and it costs nothing but a diary change.

WIP against limits

Concurrent items per stage against the published limit. Report the breach, not just the level, because breaches lead the cycle-time inflation you will see two months later. A team at six against a limit of five for three straight weeks has already committed to a bad quarter; the numbers just have not arrived yet.

Pipeline health

The count of candidates at each stage, which is your forecast. An engine with an empty shaping stage will have an idle delivery pod in about six weeks: not an opinion, arithmetic you can act on today. This is the measurement that would have saved the summit firm. In March, with delivery comfortably full, a shaping count of zero was already telling them May would be empty. Nobody was counting, so the warning arrived as a team filling itself with an untriaged idea.

Kill and park rate

The share of intake parked or declined at triage, and the share of deliveries killed or reshaped at a gate. An engine that advances nearly everything does not have triage; it has an in-tray with a rubber stamp. If fewer than a quarter of candidates are parked or declined, either your intake criteria are filtering upstream (fine, check) or your triage is a formality (not fine). Same logic at gates: a stage-gate process that has never killed anything is decoration.

Publishing the engine, which is not the same as being transparent

Now the move that changes the engine's nature rather than merely its visibility. Publish the throughput, the cycle times by stage, the WIP limit, and the delivery slots expected to open next quarter, internally, to every function that feeds you.

Transparency is the weak reason. The strong reason is behavioural. When a function head can see that shaped candidates reach delivery in about six weeks, that unshaped ones wait three times as long, and that three slots open next quarter, something shifts in their planning. They stop lobbying and start preparing. They pull their own baseline before submitting, because the published data shows them exactly what that buys. They time submissions to the slot calendar. Within a few quarters the best-organized functions are handing you pre-shaped candidates with baselines attached, which relieves pressure on your constraint using capacity you do not pay for.

That is the transformation: publication converts the engine from a central service functions queue for into a distributed capability functions participate in. It is a demand-shaping instrument, not a reporting obligation. And a function that shaped its own candidate has already done the sponsor conversation and the baseline work, so the candidate arrives at delivery with an owner who is invested rather than a stakeholder who was consulted, which makes stage five's handoff acceptance easier too.

One caution. Publishing cycle times exposes your constraint, including to people who will use it against you. Publish anyway, with what you are doing about it in the same document. A leader who says "shaping is our bottleneck, median wait 4.5 weeks, here is the fix and the date" is credible. A leader whose numbers are discovered by someone else is on the defensive for a quarter.

The Engine in Year Two: A Worked Example

Take the enterprise you have followed through this level: a 2,400-person business-to-business services and distribution company, six functions, an operating model midway through conversion, delivery capacity of roughly five concurrent efforts. All figures below are hypothetical, chosen to show the shape of a working engine. Do not benchmark against them; build your own baseline.

Intake. Across six standing sources, intake runs at 4 to 7 candidates per month, a deliberately unglamorous rate. Nobody books a room. The highest-yield source is the replication scan: in year two it contributed 5 candidates, of which 3 advanced all the way to scaled, a conversion rate roughly double any other source, at a cost of one person half a day per quarter. The finance invoice-exception pattern transferred to distribution's delivery-discrepancy queue and to service's warranty-claim queue at about 40 percent of the original build effort each, because the verification design, prompt discipline and governance path were already approved. The capability watch contributed 2 candidates, both revivals of parked items rather than new inventions.

Triage. Median 5 business days from logged to outcome, run as a standing 45-minute slot every Tuesday with the transformation lead, a data lead, and the relevant function's analyst. Over four quarters: 31 advanced, 22 parked with named triggers, 18 declined with a reason. Seven parked candidates were later revived, including two that came back automatically when the vendor-master data remediation landed in Q3, exactly as their trigger had anticipated eleven months earlier. Those two cost nothing to hold and nothing to retrieve, and arrived pre-qualified into the shaping queue at precisely the moment the blocking condition cleared. That is a queue managing itself.

The constraint. Cycle-time data through the first two quarters showed shaping at a median of 6 weeks elapsed: roughly 1.5 weeks of actual work and 4.5 weeks waiting for the single transformation analyst who could do it. The instinctive read is a capacity problem and the instinctive fix is a third transformer at, say, 85,000 fully loaded. The data pointed somewhere cheaper. Because the wait was a queue for one specific skill rather than a shortage of total effort, the programme trained two function analysts (finance and service) to shape their own candidates against a published template, at about six days of coaching each. Median shaping elapsed time fell to 3 weeks by Q4, the analyst's load dropped, and two functions gained a capability they kept. The cheaper correct fix was visible only because the waiting-versus-working split had been measured.

WIP. Theoretical capacity was 6 concurrent deliveries; the published limit was 5, and it held, including through one uncomfortable November conversation with a commercial director whose idea was genuinely good and who was offered a January slot with a queue position instead of a sixth concurrent start. He took it, and his effort started on time.

Results. Throughput: 9 use cases scaled and accepted in year two against 4 in year one. Cycle time from intake to scaled: median 22 weeks, down from 31. Both came predominantly from the shaping fix and a move from monthly to fortnightly gate slots, not from any increase in delivery headcount, which was flat.

The demand-shaping effect. The engine's numbers were published quarterly on one page to all six function heads from Q2 onwards. By Q4, three of the six were submitting pre-shaped candidates with baselines attached, unprompted, and those entered shaping at a median of 4 days rather than 3 weeks because most of the work was done. Nobody asked them to. They read the published cycle times, saw what shaping bought them in queue position, and acted in their own interest, which is the only durable form of organizational behaviour change.

Compare that year to the summit firm's, and notice the difference is not talent, budget, or technology. The summit firm generated 41 candidates in two days, converted 4, and burned four months on an untriaged idea. The engine firm generated about 65 across twelve months, triaged all of them within a week, held 22 warm with triggers, and converted 9. One organization ran an event. The other ran a process.

What to Do Monday Morning

You do not need permission or budget for any of the following. Each is a change to how you run something you already run.

  1. Map your pipeline as a five-stage process and find where work actually waits. Take the last 10 to 15 candidates you can reconstruct, date each stage transition, and split each stage's elapsed time into waiting and working. Half a day of archaeology in your own email and meeting notes gives you a baseline. Expect the answer to be shaping and gates, but verify rather than assume: if your constraint sits elsewhere, every fix below aims at the wrong target.
  2. Set a WIP limit below theoretical capacity, write it down, and show it to your sponsor. If you can genuinely run six, publish five. Then rehearse the sentence for the first time someone asks for a sixth: yes, here is the queue position and the date, or tell me which current effort to stop.
  3. Put the replication scan in the calendar as a standing quarterly half-day with one named owner. List every verified win, list every function, and walk the grid asking one question per cell: does this function have an analogous problem with an analogous shape? It is the cheapest source of high-quality candidates you own and it produces nothing until someone owns it.
  4. Introduce park-with-a-trigger as a formal triage outcome this week. Go back through your declined and stalled list, identify the ones that were early rather than wrong, write the specific reviving condition for each, and tell the original submitters. That last step is the whole point: an afternoon of short emails converts old rejections into live relationships.
  5. Publish throughput and cycle time to the functions that feed you, on one page, next quarter. Last quarter's throughput, median cycle time by stage with the waiting split shown, current WIP against limit, candidates at each stage, and slots expected to open. Include your constraint and what you are doing about it. Then watch what the best-organized function does with the information.

A well-run engine will eventually surface candidates that are genuinely uncertain: high potential value, unproven mechanism, no comparable to copy. Feeding those into a delivery pod that expects a charter with a target and a kill condition is a category error, and it is how programmes damage both credibility and budget. That is the next lesson: running real experiments without burning trust or money.

Key Takeaways

  • Reject discovery-as-event: workshops and ideation summits produce a batch that depletes on delivery's schedule, and the vacuum that follows gets filled by whatever is loudest rather than whatever passed triage.
  • Treat your use-case pipeline as a process and apply the discipline you already apply everywhere else: baseline it, instrument it, find the constraint, fix the constraint, repeat.
  • Run the five stages with explicit entry and exit criteria: intake (a stated problem with a mechanism), triage (advance, park, or decline within days), shaping (baseline, data screen, confirmed sponsor, draft charter), delivery (under a hard WIP limit), and scale with an accepted handoff.
  • Add the two mature intake sources: the quarterly replication scan, the cheapest high-quality candidate source you own and worthless until someone owns it, and the capability watch, which asks which parked candidates just changed status.
  • Use park-with-a-trigger as your default middle outcome, because most candidates are early rather than bad, and a written reviving condition turns a rejection into a relationship and a self-managing queue.
  • Publish a WIP limit below theoretical capacity and defend it, remembering that a system at 100 percent utilisation has exploding queue times: allowing six concurrent efforts on five efforts' capacity yields four, later.
  • Split cycle time into waiting and working, because if the time is dominated by waiting then adding capacity to the working stages changes nothing, and the fix is usually a diary change or a trained function analyst rather than a hire.
  • Publish throughput, cycle time and open slots to the functions that feed you, because publication is a demand-shaping instrument that converts a central service into a distributed capability the functions staff themselves.