←
AI Readiness & Process Transformation
Strategic · M7 · lesson 7 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Building the AI Use-Case Portfolio: Value x Readiness

15 min

The spreadsheet has forty-one rows and it is, by every conventional measure, excellent work. Each row is a scored AI use case: value on one axis, feasibility on the other, a weighted composite in column K, sorted descending. The strategist spent five weeks on it and it is defensible line by line. She presents it to the steering committee, and the committee asks a question she cannot answer. "So we do the top eight?" She hesitates, because she can see something the sort order cannot: five of those top eight depend on the same unowned vendor master, three cannot start before the ERP cutover in Q3, all eight sit under one executive sponsor who is rumored to be leaving, and the organization has exactly two and a half people who know how to run a transformation of this kind. The list is right. The list is also not a plan. A list of scored use cases is not a portfolio, any more than a list of stocks is an investment strategy.

The List That Was Not a Portfolio

In Level 2 you learned the process-selection scorecard. It answers one question with rigor: is this particular process worth attacking with AI? Applied to one candidate, it is close to sufficient. Applied to forty-one, it is necessary and nowhere near sufficient, because the enterprise question is different. It is not "is this worth doing?" It is "what mix of bets should we make, given that capacity is finite, that risks correlate, that we need visible proof before we can afford ambition, and that we get roughly one budget cycle to prove this program is not the last one?"

Those four constraints turn a list into a portfolio problem, and each breaks the sort order differently.

  • Capacity is finite and smaller than you think. Not budget capacity: human capacity, the number of people who can actually run the baseline-redesign-pilot-measure method from Levels 2 and 3. In most 2,000-person companies that is between one and four, and it is the real ceiling regardless of budget.
  • Risks correlate. Five pilots on the same unowned vendor master do not carry five independent risks. They carry one risk, five times. A portfolio diversified on the value axis can be catastrophically concentrated on the dependency axis, and the scorecard cannot see it because it scores each case alone.
  • Evidence precedes ambition. A program that cannot show one verified, baselined win in its first two quarters is usually not killed. It is slowly defunded in the third, by a CFO (chief financial officer) who stopped asking questions because she stopped expecting answers.
  • Composition changes what a program becomes. A portfolio of only quick wins builds no capability and runs out of easy processes. A portfolio of only deep bets produces no evidence before the budget cycle ends. Neither failure appears in any individual case's score.

And one composition failure is so common that it belongs at the top of the lesson rather than buried in the failure story: the hero project. One enormous flagship initiative, eighteen months long, seven figures deep, carrying the entire program's credibility inside it. That is not a portfolio. It is a bet with a press release attached. Hold the shape in mind, because everything the portfolio discipline does is designed to make hero projects structurally impossible.

A list of scored use cases is not a portfolio, any more than a list of stocks is an investment strategy. The portfolio is the mix, the correlation, and the capacity it has to fit inside.

Your artifact is the Portfolio Allocation Sheet: one page carrying the three-tier mix (quick wins, capability builds, deep bets) with target percentages of capacity, each surviving candidate's value and readiness scores, its dependency tags, and its correlation flags. It shows the shape of your program rather than its wish list, and it makes a hero project look, at a glance, exactly as reckless as it is.

Pipeline Intake: Where Candidates Come From Determines Their Quality

Most enterprises get their candidate list the same way: someone emails a request for AI ideas, and back come forty-one items of wildly uneven quality, dominated by whoever writes the longest emails. A portfolio inherits the biases of its intake, so intake gets designed deliberately, across four sources ranked by the evidence value of what they produce.

Source 1: the shadow-AI demand map

Your People Readiness Atlas found that 71 percent of employees are already using AI tools for real work, most of it unsanctioned. Every enterprise treats this as a governance problem. It is also the highest-quality candidate source you will ever have, for a reason that is almost embarrassingly simple: people do not sneak tools that do not work.

An employee who quietly pastes supplier emails into a consumer chatbot every morning, at some personal risk, has run an unfunded, months-long pilot and reached a verdict: the task is AI-tractable, the output is good enough to be worth the effort, and the pain was severe enough to justify going around policy. Those are three findings your formal process would spend eleven weeks and a consultant to produce. This is pre-validated demand, and the usual organizational response is to send a note about acceptable use and ignore it. Harvest it instead: for each shadow-use cluster, record the task, the volume, the tool, the workaround, and the reason the sanctioned path was avoided.

Source 2: the Process Maturity Grid's build-ready functions

Your grid classified six functions: two BUILD-READY, three REMEDIATE-FIRST, one GREENFIELD. Build-ready functions are terrain that can hold value: documented processes, measurable baselines, owned data, a manager who knows their own cycle times. Sweep those two systematically, because a mediocre use case in build-ready terrain out-delivers an excellent one in terrain that cannot support it.

Source 3: the pain inventory

The Level 2 process inventory already carries cost and error data across functions. Filter it for high-cost, high-error, high-volume processes and pull the top rows regardless of function. This source yields strong value cases with often weak readiness, which is exactly what makes it useful: those candidates populate the capability-build and deep-bet tiers, and they tell you what the remediation backlog is for.

Source 4: executive nominations

The CEO went to a conference. She came back with an idea. It is going on the list.

Do not fight it, and do not roll your eyes at it. Fighting costs political capital you need for harder arguments. Instead, do what the scoring apparatus was built to let you do: put the nomination through the same intake and scoring as everything else, in public, using criteria and weights declared before anyone knew what the CEO would nominate. That is the constitutional function of a pre-declared scorecard. When the criteria were set before the nomination existed, they can decline it without anyone having to. You are not saying no to the CEO; the readiness score is saying not yet, and you are the person offering the remediation path that would change it.

Intake hygiene: the mechanism field

One field on the intake form does more filtering than the entire scoring model that follows it. Every candidate must be stated as a hypothesis with three parts: the process it touches, the constraint that currently limits that process, and the expected mechanism by which AI relieves that constraint.

One sentence: "AI categorization cuts invoice-exception cycle time because the queue, not the reading, is the constraint." That names a measurable outcome, names the bottleneck, and makes a falsifiable claim about the causal path. If the constraint were reading speed, the expected gain would be smaller. If it were an approval that sits with one person on Tuesdays, AI would relieve nothing, and the honest answer would be to fix the approval.

Now try writing that sentence for "we should use AI in HR," or "an AI assistant for the whole company," or "AI-powered supply chain intelligence." You cannot: there is no process, no named constraint, no mechanism, only a category and a wish. Candidates that cannot state a mechanism are not candidates. They are aspirations, and aspirations do not enter a portfolio, because a portfolio is a set of commitments about how capacity converts into outcomes.

Enforce the field and a meaningful fraction of intake dies before scoring at near-zero cost, the cheapest filtering available. It also teaches the organization, one returned form at a time, what an AI use case actually is: after two cycles the submissions improve, because people start hunting for a mechanism before they submit.

The Two-Axis Score at Portfolio Scale: Value x Readiness

Survivors get scored on two axes, one to ten each. The value axis is the Level 2 instrument, essentially unchanged: annualized hours or dollars at stake, error and rework cost, strategic weight, customer or compliance exposure, under weights declared before you looked at any candidate.

The second axis is where Chapter 4.1 pays for itself. At single-process scale you derived feasibility case by case: looked at the data, checked the documentation, asked whether users would adopt it. That costs days per candidate. At forty-one candidates it costs a quarter, and by the time you finish, the first ones are stale.

So you do not re-derive readiness. You import it. Each candidate inherits its readiness components from audits you already ran:

  • Process readiness from the Process Maturity Grid: the candidate's function already carries a maturity score, and its BUILD-READY, REMEDIATE-FIRST, or GREENFIELD classification travels with every candidate inside it.
  • Data readiness from the Critical Dataset Register: look up the specific datasets the candidate consumes and take their scored condition (ownership, completeness, accuracy, accessibility, structure). If those datasets carry a remediation line in the 22-line backlog, the candidate's data readiness is capped until the line closes.
  • People readiness from the People Readiness Atlas: capability level in that function, presence or absence of champions, current change load, honest adoption prognosis.
  • Governance readiness from the framework's governance dimension: whether this use-case class has an approved control pattern, a human-gate design, and a decision-rights answer, or would be the first of its kind.

Say this out loud in the room, because it is the moment eleven weeks of audit stop being a cost and become an asset: the enterprise readiness audit is scoring infrastructure. It was expensive, it was slow, several people asked what it was for, and this is what it was for. Every future candidate now gets a readiness score in twenty minutes instead of four days, because the terrain is already surveyed. Over a three-year program, that is probably the larger half of the audit's return.

Declare the threshold in advance too. A workable default: a candidate enters only if value is at least 5, readiness at least 5, and the sum at least 12. Below the line it is not rejected forever; it is returned with the failing readiness component named and the remediation that would change the answer. "Not yet, and here is what yet requires" is a different message from "no," and it is the one that keeps your pipeline full next quarter.

The Three Tiers, Defined by Purpose Rather Than Size

Here is where good scoring still produces bad portfolios. The team sorts by composite score and funds top-down, yielding a portfolio shaped entirely by whatever the scoring happened to favor. The discipline is to allocate into three tiers first and let each tier compete internally. The tiers are not defined by project size or duration. They are defined by what each is for.

Tier 1: quick wins (purpose: evidence and momentum)

High readiness, modest value, an eight to fourteen week cycle, located in build-ready functions, tightly scoped, properly baselined, and chosen partly for visibility to the audience that controls next year's budget.

That last clause makes people uncomfortable, so be honest about it without being cynical. The quick win has a political function, and pretending otherwise is naive rather than principled. A program that cannot show a verified win in two quarters loses its budget in the third. That is not hypothetical: it is the mechanism behind the year in which 42 percent of companies scrapped most of their AI initiatives (S&P Global, 2025). Many of those did not fail because AI did not work for them. They failed because nothing finished in time to be shown.

The crucial point is that a quick win's evidence is real evidence: pre-pilot baseline, pre-committed success and kill criteria, honest measurement, documented decision. You are not manufacturing a demo; you are choosing, among genuinely valuable candidates, the ones whose proof arrives early enough to keep the program alive. Choosing the order in which you produce truth is strategy. Choosing which truths to tell is spin. Tier criteria: readiness 7 or above, a baseline that exists or can be captured in two weeks, one function's cooperation rather than three, a cycle under fourteen weeks, and an outcome expressible in one sentence a non-specialist believes.

Tier 2: capability builds (purpose: raising the terrain)

The items that produce no direct profit-and-loss impact and unblock everything else: high-beneficiary lines from the remediation backlog, data-ownership assignments, the access fast-path, the grounding infrastructure that lets a retrieval-augmented generation (RAG) pattern cite your own documents, the role-based training program, the control scaffolding your first agent will need.

This tier is simultaneously the line most likely to be cut and the most expensive to skip, and the reason is a category error in presentation. Called "overhead" or "foundational investment," it competes against use cases carrying dollar figures and loses every time, because it has no advocate, no visible outcome, and no demo. So stop presenting it as overhead. Fund capability builds as portfolio items with beneficiary arithmetic attached. Vendor-master ownership and deduplication is not a data-hygiene chore costing 34 effort-days. It is the prerequisite for seven of your eleven surviving candidates, which means 34 effort-days unblocks roughly two-thirds of the portfolio's value, and each of those seven carries a capped data-readiness score until it is done. Written that way it stops being a cost line and becomes the highest-leverage item on the page.

This tier is where BCG's 10-20-70 rule gets funded or ignored. If 10 percent of value comes from algorithms, 20 percent from technology and data, and 70 percent from people and process, a portfolio with no capability-build line has funded the 10 percent and hoped the 70 percent shows up by itself. It does not. McKinsey sharpens the point: 88 percent of organizations use AI regularly, only about 39 percent report any EBIT (earnings before interest and taxes) impact, and high performers are roughly three times more likely to have fundamentally redesigned workflows. Workflow redesign is capability, and it lives here.

Tier 3: deep bets (purpose: asymmetric upside)

The redesign that changes a function's economics. The revenue-side experiment. The greenfield function with no process to improve, which therefore has to be designed. Long, genuinely uncertain, expensive to be wrong about, and where the 95 percent lives: MIT's GenAI Divide found roughly 95 percent of enterprise generative AI pilots produced no measurable profit-and-loss return, and the deployments that consumed the most money and credibility were disproportionately the ambitious ones.

The response is not to avoid deep bets; a portfolio without them is a cost-reduction program that plateaus in eighteen months. It is to cap them at a deliberate share of capacity and stage-gate them hard: a discovery phase with a real go or no-go decision before any build capacity is committed, a named kill criterion, and a charter statement that a documented kill is a successful outcome in this tier. A later lesson develops the gating fully. Here, what matters is that the cap is a number you wrote down rather than a feeling you had. This is also the one tier that may admit a candidate below the readiness floor, because its purpose is to buy an option rather than to deliver now, and it may do so only through a capped discovery phase whose explicit job is to raise that readiness score before any build capacity is committed.

The target mix: a policy declared before allocation

Before you assign a single candidate to a tier, declare the target mix as a percentage of capacity, in the same spirit as your scoring weights: in advance, in writing, with reasoning attached, so allocation becomes policy rather than a series of arguments you must win one at a time. An illustrative starting allocation for an organization coming out of a scrapped-project year, where credibility is the scarcest resource:

TierTarget share of capacityWhy this share
Quick winsroughly 50%Credibility deficit. The program needs two or three verified, baselined wins inside two quarters or the budget conversation goes badly. Half of capacity buys three.
Capability buildsroughly 30%The audit found terrain is the constraint. Under-fund this and every later quarter's quick wins get harder, not easier, and the portfolio runs out of easy processes by year two.
Deep betsroughly 20%Enough to keep one real option alive and to signal strategy rather than a cost program. Small enough that failure is survivable and an overrun cannot eat the proof engine.

These percentages are a judgment, not a law, and the judgment moves with program maturity and trust. A high-trust organization with three years of delivered wins and an executive team that has watched the method work does not need half its capacity proving it can finish things; it might run 25 percent quick wins, 30 percent capability builds, and 45 percent deep bets, because its scarce resource is no longer credibility but ambition, and a failed bet costs it a lesson rather than a defunding. An organization that just lost a seven-figure hero project might deliberately run 65 percent quick wins for a year until it has rebuilt the right to be trusted with anything longer. The rule: the mix follows the evidence bank. Declare the current mix, declare the condition under which it shifts, and revisit it on the same cadence as the heat map.

The Three Disciplines No Single Scorecard Can See

Everything so far could in principle be done candidate by candidate. The next three checks cannot: they are properties of the set, not of any member, and they are why the Portfolio Allocation Sheet is one page rather than a folder of scorecards.

Discipline 1: correlation flags

Tag every surviving candidate with its dependencies in five categories: dataset (which registered datasets it consumes), system (which platforms it touches, and any freeze, upgrade, or migration window on them), vendor (which external platform it depends on, including shared ones nobody thinks of as a dependency), function (whose operational attention it consumes, and how much change that function is already absorbing), and sponsor (which executive owns it politically).

Then do not read the tags per candidate. Read them as concentration across the set. "Seven of our twelve candidates depend on the vendor master" is one sentence doing two jobs. It is the strongest argument for funding the vendor-master remediation line, because it makes the beneficiary arithmetic visible. And it is a warning, because one data failure now stalls two-thirds of the program simultaneously, in the same month, in front of the same steering committee.

This is the concentration finding from your data audit doing a different job. In the audit it was a diagnosis; in the portfolio it is an allocation input, naming what to fund first and where the single point of failure sits before you build it. A portfolio diversified by function, value, and tier can still be concentrated on one dataset, one vendor, or one person, and no sort by composite score will ever reveal it.

Discipline 2: capacity realism, and the failure mode nobody names

Level 3's replication playbook taught the transformer-hours arithmetic: the constraint on spreading a method is not enthusiasm, it is the number of people who can run it. At portfolio scale that becomes the hard ceiling on allocation.

Count honestly. How many people can independently baseline a process, redesign the workflow, define human gates, run a pilot with pre-committed kill criteria, measure against the baseline, and write the honest verdict? Not how many are interested, and not how many attended training: how many have done it and could do it again unsupervised. In a 2,400-person company that has run one successful transformation, the answer is often two, or two and a half counting the analyst who did half of one under supervision.

Convert to concurrency strictly. A trained transformer cannot run four efforts at once, because the method has phases of different intensity: discovery and baselining are heavy, the pilot's mid-phase is lighter, measurement and write-up are heavy again. With staggered phases, two transformers support roughly three concurrent efforts. Not nine. Three, and only if the phases are offset so two heavy ones never land in the same fortnight.

Now the part worth memorizing, because it is why capacity overruns go undetected for two quarters. A portfolio that exceeds its capacity does not fail visibly. It fails as universal slowness. Nothing is cancelled or escalated. Every effort simply takes 60 percent longer, every baseline lands late, every measurement window slides, every steering update reads "on track, minor delays," and at the end of two quarters there are seven efforts at 70 percent completion and zero verified wins. Zero verified wins is indistinguishable, from the CFO's chair, from a program that did nothing. Then the budget conversation happens. Name that failure mode in your own organization and use the name out loud, because named failure modes get noticed early and unnamed ones do not. Universal slowness: everything moving, nothing finishing.

The last check takes five minutes and prevents a failure that has nothing to do with technology. Count funded items per executive sponsor. If a majority sit under one, the portfolio has a single political failure mode, and executives change jobs on a timescale shorter than deep bets. When that sponsor leaves, the portfolio leaves: the successor inherits initiatives they did not choose, defended by a strategist they did not hire, and their fastest route to their own agenda is a review that quietly descopes the lot. The fix is not to abandon your strongest sponsor but to ensure at least two functions and two sponsors hold something they would personally defend, so the program has more than one root in the ground. If that forces you to swap a slightly stronger candidate for a slightly weaker one in a second function, take the trade: you are insuring the whole portfolio for a few points of composite score.

The Artifact: A Portfolio Allocation Sheet, Worked End to End

Build the whole thing for the enterprise you have been auditing all level: 2,400 people, business-to-business services and distribution, six functions, one scrapped-initiative year behind it, four audit artifacts and a funded heat map in hand. All figures are hypothetical, shown to illustrate the arithmetic rather than as benchmarks.

Intake: 19 candidates in, 11 out

A three-week intake window produces 19 candidates: 7 from the shadow-AI demand map (the clusters behind the atlas's 71 percent finding: supplier email triage, contract-clause lookup, ticket summarization, RFP response drafting, meeting-note extraction, spreadsheet reconciliation, policy question answering), 5 from the maturity grid's two build-ready functions (Finance Operations and Customer Operations), 4 from the pain inventory, and 3 executive nominations.

The mechanism field kills 4 before scoring. Two are shadow-map entries that turn out to be personal productivity habits with no process attached (a manager summarizing his own reading; nobody else's cycle time changes). Two are pain-inventory items whose constraint, once written down, is an approval bottleneck and a staffing gap, neither of which AI relieves. Those two are referred to process improvement without AI, a genuinely good outcome the field produced for free.

Fifteen are scored against the pre-declared threshold. Four fall below it: two of the three executive nominations, one grid candidate, one pain-inventory item. The nominations are declined by the criteria, in the meeting, with the failing component named and a remediation path offered: "demand forecasting scores 3 on data readiness because the demand history lives in three systems with no owner; that is register line 14, roughly 40 effort-days, and closing it moves this candidate to a 7." The CEO's response, in practice, is rarely resistance. It is "then why isn't line 14 funded?" Which is exactly the conversation you wanted.

Eleven survive: 5 from the shadow map, 4 from the grid, 1 from the pain inventory, 1 executive nomination. Note the conversion: the shadow map supplied 37 percent of intake and 45 percent of survivors, the best ratio of any source, which is the empirical case for treating unsanctioned use as a demand signal and not only as a policy violation.

Capacity and declared mix

Honest capacity: 2.5 trained transformers. Over a 26-week two-quarter window that is 65 transformer-weeks, supporting 3 concurrent delivery efforts plus 1 discovery track at staggered phases. Declared mix, written before allocation: 50 percent quick wins, 30 percent capability builds, 20 percent deep bets.

The sheet

TierItemSourceValueReadinessCorrelation flagsCapacity
Quick winContract-renewal notification triage (Customer Ops): replication of the invoice-exception methodShadow map78Data: contract table. Sponsor: COO12 wk
Quick winSupplier-statement reconciliation exceptions (Finance Ops)Grid68Data: vendor master. System: ERP Q3 window. Sponsor: CFO11 wk
Quick winRemittance-advice matching (Finance Ops)Exec nomination67Data: vendor master. Sponsor: CFO9 wk
CapabilityVendor-master ownership and deduplication (register line 3)Registerenablingn/aUnblocks 7 of 11 candidates9 wk
CapabilityData-access fast path (3 weeks to 4 days)Atlasenablingn/aUnblocks all tiers4 wk
CapabilityRole-based enablement program, delivered through the 41 championsAtlasenablingn/aSeeds the 2 champion-empty functions7 wk
Deep betSales function discovery then redesign (greenfield), gate at week 6Grid94No data owner. Sponsor: CRO. Single-sponsor risk13 wk

Seven items funded, 65 transformer-weeks allocated exactly: 32 weeks to quick wins (49 percent), 20 to capability builds (31 percent), 13 to the deep bet (20 percent). The declared mix held, which is the point of declaring it: when the chief revenue officer argued for a second sales initiative, the answer was arithmetic rather than opinion.

Four survivors are parked with written reasons, because a parked candidate with a reason is a promise and one without is a grievance: procurement spend classification (Q3 ERP window, sequenced into next cycle), HR onboarding answer assistant (pending policy-corpus remediation, register line 9), quality-claim summarization (capacity), inside-sales quote drafting (capacity and sponsor concentration).

Reading the concentration

  • Vendor master: 7 of 11 survivors. The capability build costing 9 transformer-weeks and producing no profit-and-loss impact of its own is now visibly the insurance policy on two-thirds of the portfolio. Nobody argues to cut it once the count is on the page. That is beneficiary arithmetic doing in one line what a business case could not do in four slides.
  • Q3 ERP window: 4 of 11. Four candidates collide with the freeze the People Readiness Atlas already flagged as a change-load risk. One is inside the funded set and needs its phases offset; the others are parked or sequenced. That is a sequencing problem, and sequencing is the next lesson.
  • Sponsor spread: 5 of 11 sat under the CFO before allocation. The funded set has two under the CFO, one under the COO, three enterprise-wide capability builds, and one under the CRO: three sponsors with real skin in it, thin but survivable. The check changed the allocation, swapping one CFO-sponsored candidate for a Customer Operations candidate scoring a point lower, deliberately, to buy the second root.
  • Capacity: 65 available, 65 allocated. The portfolio fits with no slack, which is itself a finding: the first unplanned event consumes a buffer that does not exist. The mitigation is written on the sheet as a pre-agreed descope order naming which item slips first.

That last line is a small discipline with outsized value. Decide the descope order while everyone is calm and the portfolio is theoretical. Decide it in month four, during the crisis, and it will be decided by whoever is loudest.

The Failure Story: The Hero Project

Consider the alternative, which is what most enterprises actually do. A mid-sized manufacturer's board wanted an AI strategy. What it approved was one project: a 4.2 million dollar "AI-powered supply chain intelligence platform," eighteen months, one vendor, one enormous integration touching four systems, and a business case built on a projected 9 percent reduction in inventory carrying cost. It was, everyone agreed, the transformative one. It was also the only one. No quick-win tier, because quick wins looked unambitious beside it. No capability-build tier, because the vendor's proposal included a data workstream and everyone assumed that covered it. No other bets, because the capital was committed.

Note the shape before the story unfolds: the program's entire credibility now lives inside a single artifact with an eighteen-month reveal, and the first honest datapoint about whether this organization can convert AI into value arrives in month eighteen. It arrived early, and badly.

  • Month 14: the data foundation surfaces. The supplier master has 31 percent duplication, the item master has three competing hierarchies, and lead-time history was never captured usably. No Critical Dataset Register was ever built, so this discovery happens during integration testing rather than during planning. Remediation is scoped at 7 months and 900,000 dollars, unbudgeted. Gartner's finding that 63 percent of organizations lack AI-ready data practices, and that through 2026 some 60 percent of AI projects without AI-ready data will be abandoned, is not an abstraction in that room. It is the change order on the table.
  • Month 19: the executive sponsor takes a job elsewhere. His successor inherits a project she did not choose, 4.2 million dollars spent, no interim evidence of any kind, and an unbudgeted remediation. Her rational move is a review. Sponsor spread: one. The portfolio had a single political root and it was pulled.
  • Month 22: the platform is quietly descoped to a dashboard. The dashboard is fine. It shows inventory positions on a screen where previously they were in a report. Nobody schedules a post-mortem, because a post-mortem would require someone to say the words out loud.

The organizational conclusion, stated in corridors rather than documents, is "AI doesn't work for us." The company attempts nothing again for two years.

Price the counterfactual, because that is where the real loss sits. For roughly a third of the money, that organization could have run a dozen eight-week efforts across its build-ready functions. Perhaps four would have produced nothing, which is what kill criteria are for. Perhaps six would have produced modest, verified wins. Perhaps two would have found something worth building a deep bet around, chosen on evidence rather than a vendor's slide. And it would have built what it needed most and now does not have: people who know how to do this, a register of its own data, a measurement habit, and an evidence bank. Instead it bought a dashboard and a two-year moratorium.

The case against the hero project in one sentence: it converts an entire strategy into a single point of failure, at the exact moment when the organization has the least evidence about its own capability. MIT's divide is instructive in a way the headline hides. The 5 percent that crossed from pilot to production were not the most ambitious deployments. They picked one process, integrated deeply, redesigned the surrounding work, and measured. That is a quick-win shape, repeated. Ambition earned its place afterward, on top of evidence.

The Portfolio Allocation Sheet makes the hero project nearly impossible, and that is its quiet best feature. Put a 4.2 million dollar, eighteen-month, single-vendor, single-sponsor item on that page and every column indicts it: it consumes the deep-bet cap several times over, it leaves the quick-win tier empty, its correlation flags concentrate every category at once, its capacity requirement exceeds the honest transformer count, and the sponsor-spread check returns one. You do not have to win the argument. You have to show the page.

What to Do Monday Morning

You can produce a first Portfolio Allocation Sheet in about a week if the audit artifacts exist, and in about three weeks if you gather readiness inputs as you go.

  1. Run intake with the mechanism field enforced. Require one sentence per candidate: process, constraint, mechanism. Return everything that cannot produce it, with a note explaining what the field wants. Expect to lose a fifth to a quarter of the list at no cost, and expect two returns to be process problems worth fixing without AI.
  2. Add the sources you are not using. Most lists are entirely executive nominations and pain inventory. Pull the shadow-AI demand clusters from your People Readiness Atlas and sweep your build-ready functions deliberately. If you have no atlas, ask three managers what tools their teams actually use, and write down what you hear without reacting to it.
  3. Import readiness instead of re-deriving it. Pull process readiness from the maturity grid, data readiness from the register, people readiness from the atlas, governance readiness from the framework. If a candidate's dataset carries an open remediation line, cap its data-readiness score until that line closes, and say so on the sheet.
  4. Tag dependencies and read the concentration. Five categories: dataset, system, vendor, function, sponsor. Then count across the set and write the two or three concentration sentences that fall out at the top of the sheet. Those sentences are your capability-build business case and your risk register in the same breath.
  5. Compute honest transformer capacity. Name the individuals who could run the full method unsupervised, multiply by weeks in the planning window, convert to concurrent efforts at roughly 1.2 per transformer with staggered phases. Resist rounding up: exceeding capacity produces universal slowness, invisible until the quarter ends.
  6. Declare your three-tier target mix before allocating a single project. Write the percentages, the one-line reasoning for each, and the condition under which the mix shifts toward deep bets. Then allocate inside the tiers, let each tier compete internally, and park the rest with written reasons and named unblocking conditions.
  7. Run the sponsor-spread check and set the descope order last. Count items per sponsor and make sure at least two executives are genuinely invested. Then decide, in writing and in advance, which item slips first when something goes wrong. Something will go wrong.

Key Takeaways

  • Distinguish the scorecard's question from the portfolio's: the Level 2 instrument answers "is this one worth doing?", while the portfolio answers what mix of bets to make under finite capacity, correlated risk, and a credibility deadline.
  • Design intake across four ranked sources, treating the shadow-AI demand map as the highest-evidence one because people do not sneak tools that do not work, and scoring executive nominations against pre-declared criteria so the criteria can decline them gracefully.
  • Enforce the mechanism field on every candidate (process, constraint, expected mechanism in one falsifiable sentence), because candidates that cannot state a mechanism are aspirations, and the field filters a meaningful fraction of intake at almost zero cost.
  • Import readiness from your audits rather than re-deriving it per case: the maturity grid, dataset register, people atlas, and governance dimension turn the audit into permanent scoring infrastructure that prices any future candidate in twenty minutes.
  • Allocate into three tiers defined by purpose: quick wins buy real evidence and political runway, capability builds raise the terrain and must be funded with beneficiary arithmetic rather than sold as overhead, and deep bets carry asymmetric upside under a hard capacity cap and a stage gate.
  • Declare the target mix as policy before allocating (roughly 50/30/20 coming out of a scrapped-project year, shifting toward bets as the evidence bank grows) so allocation becomes arithmetic rather than an argument won case by case.
  • Read dependency tags as concentration across the whole set, since portfolio risk is correlated risk: "seven of eleven depend on the vendor master" is at once the strongest argument for the remediation line and the warning that one data failure stalls two-thirds of the program.
  • Count transformer capacity honestly and refuse the hero project structurally: an over-committed portfolio fails invisibly as universal slowness (seven efforts at 70 percent, zero verified wins), and one enormous flagship initiative is not a portfolio but a single point of failure with a press release, indicted by every column of the Allocation Sheet.

The portfolio says what you are betting on and in what proportions. It does not say when, or in what order, or how to sequence around a Q3 ERP freeze, a change-fatigued function, and a capability build that must finish before three quick wins can start. That is the sequencing problem, and it is the next lesson.