←
AI Readiness & Process Transformation
Strategic · M20 · lesson 20 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

The Business Case That Survives Finance

15 min

The person who decides whether your program lives is not the chief financial officer (CFO). It is the analyst two desks down from her, twenty-nine years old, who maintains a spreadsheet nobody asked him to build. It has one row per AI business case the company approved in the last three years, and four columns: what was promised, what was spent, what was delivered, and who signed. Eleven rows. In the "delivered" column, nine cells say some version of "not separately measurable." He does not editorialize. He does not need to. When your wave-one case arrives in his inbox, he will open it beside that spreadsheet and know within ninety seconds whether yours is row twelve or something new. This lesson is about being something new, and the strange discovery at the center of it: the case that admits the most uncertainty is the one that gets funded.

Why Finance Rejects AI Cases (It Is Not Hostility)

Ask a room of transformation leads why finance blocks AI investment and you get folklore: finance does not understand the technology, finance is conservative, finance thinks in quarters while AI pays back in years. That folklore is comfortable and almost entirely wrong, and believing it will cost you a funding cycle. Finance funds uncertain, long-horizon, technically opaque things constantly: warehouse automation with eight-year paybacks, enterprise resource planning migrations whose benefits are famously hard to isolate, insurance captives whose economics almost nobody in the building can explain. Finance is not allergic to uncertainty. It is allergic to undisclosed uncertainty, a completely different allergy, and the distinction is the whole lesson.

Here is what actually happens. Your case lands on a desk that has processed three years of AI proposals and is pattern-matched in seconds against the shape of the ones that failed. That shape is remarkably consistent across companies and industries: a single-point benefit projected forward as a hockey stick, costs limited to the license fee and an implementation quote, no measured baseline for the process being improved, no explanation of the mechanism by which the number is supposed to change, and a payback calculated as though adoption were total from week one. Not one of those is a technology problem. Every one is a document problem. The proposals failed the same way, so the reader learned to reject the shape, and your genuinely good idea is being judged on its silhouette.

The evidence behind that reflex is not folklore either. S&P Global found 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before. MIT's GenAI Divide study found roughly 95 percent of enterprise generative AI pilots produced no measurable profit-and-loss (P&L) return. McKinsey's 2025 State of AI survey found that although 88 percent of organizations use AI somewhere, only about 39 percent can attribute any earnings-before-interest-and-taxes (EBIT) impact to it, and most of those report it at under 5 percent. Gartner expects over 40 percent of agentic AI projects to be canceled by the end of 2027, and names escalating costs as a leading cancellation driver alongside unclear business value.

Sit with that last one, because it is the most useful of the four. Projects are not canceled mainly because the technology disappointed. They are canceled because the cost line kept climbing past what the business case said it would be, and once a program is repeatedly asking for more money than it forecast, its credibility dies before its results arrive. That is a business-case failure wearing a technical costume. The single highest-leverage thing you can do for your program's survival is not to improve the technology selection. It is to write a cost section that does not have to be revised upward in month five.

If your organization lived through its own scrap year, and most have, add a local layer to the national statistics. Somebody in your building signed those eleven rows and defended them at a mid-year review. The memory in the room when you present is not abstract skepticism; it is a specific person remembering a specific bad afternoon. Your case is not competing against a neutral prior. It is competing against a scar.

Finance does not reject your AI case because it is uncertain. It rejects it because it is shaped exactly like the ones that failed, and the shape is the only thing a busy reader can evaluate in ninety seconds.

Which sets up the counterintuitive move this lesson turns on. The instinct, facing a skeptical audience, is to project confidence: round the benefits up, smooth the ramp, keep the cost section tight so the payback looks fast. That instinct is exactly backwards. The reader's problem is not that your numbers are too small. It is that they cannot tell whether your numbers were built or wished. Every range with a stated basis, every cost line other people omit, every honest ramp, is evidence of construction. Admitted uncertainty is not weakness here; it is the fingerprint of a real analysis, and it is the only thing that distinguishes you from row twelve.

Note what has changed since Level 1. There you learned to read a vendor's return-on-investment (ROI) math skeptically: demand full-cost accounting, hunt the missing verification-labor line, treat "up to 40 percent faster" as a marketing artifact rather than a forecast. That skill now points at your own keyboard. You are the one making claims, and the fastest way to lose a finance audience is to demand rigor of vendors that you did not apply to yourself.

The Artifact: The Wave Business Case

The artifact this lesson gives you is the Wave Business Case: one page plus a workbook, in seven fixed sections, each defensible line by line. It is scoped to a wave, meaning the two to four portfolio items you intend to fund and run together in one sequenced block. That scoping matters: a single use case is too small to justify the shared capability spend it depends on, and the whole program is too large for anyone to interrogate. The wave is the unit at which the economics resolve. The order of the seven sections is not cosmetic either. It is a persuasion sequence built for a reader who decides in thirty seconds whether to read carefully or skim, and each section answers the objection the previous one raises.

1. The Baseline: what this costs today, measured

The case opens with the baseline, not with the opportunity. Not with market context, not with the chief executive's AI ambition, not with a competitor's press release. It opens with a measured statement of what the target processes cost the company today, expressed as an annual run rate, with the measurement method named and the source cited by version: "Baseline Pack v2.1, invoice-exception handling and supplier-onboarding, measured across eleven weeks, sample method described in tab 2."

Do this and you have already separated yourself from the overwhelming majority of what crosses that desk. Most AI cases contain no measured baseline at all. They contain an estimate of the benefit and a silence where the current state should be, which is why their benefits can never be verified afterward: there was nothing to verify against. When a case opens with a measured baseline and a method, the reader's posture changes physically, because the document has just demonstrated that somebody went and looked.

Cite the baseline by version, the way you would cite a document control number in a standard operating procedure (SOP). Versioning signals that the number has a history and can be re-derived, and that when the post-implementation review happens in fourteen months the comparison will be against a fixed known thing rather than a remembered one. If your baseline changed between v1.0 and v2.1, say so and say why in the workbook. An analyst who finds a documented baseline revision trusts you more, not less.

2. The Mechanism: why the number changes

One paragraph. Not a technology description, not an architecture diagram, not a capability list. A causal statement of what constrains the process today and how the redesign attacks that constraint. For example: "The constraint in invoice exceptions is not reading speed; it is the two-day queue created because exceptions batch to a Tuesday and Thursday review. The redesign attacks the queue by classifying and routing on arrival, which removes the batching, and the classification step is where AI does the work. Reading speed improvements, which is what the vendor demonstrates, are worth roughly nothing here."

That paragraph does more work than any other hundred words in the document. Finance funds mechanisms, not technologies. A mechanism can be argued with, checked, and disproven, which is precisely why it is credible. A technology cannot: "we will apply AI to invoice processing" is not a claim, it is a hope with a vendor attached. If you cannot state the mechanism in one paragraph without using the word "leverage," you do not yet understand your own case well enough to defend it, and the meeting will discover that faster than you will.

There is a private diagnostic hiding here. Try stating the mechanism for each item on your roadmap. Items whose mechanism comes out crisp in two sentences are usually your real candidates; items whose mechanism comes out as a paragraph of adjectives are usually items where somebody fell in love with a capability and went looking for a process to attach it to. That is the paving-the-cow-path failure, caught before it costs anything.

3. Benefits as Ranges, with the Basis of Each Stated

Never a point estimate. Three numbers, each with its basis written next to it, in one line of plain language:

  • Low: the conservative case, derived from your organization's own measured precedent applied at its most pessimistic. Basis: "the invoice-exception redesign delivered a measured net 428,000 annual run-rate reduction; the low case assumes this wave achieves the same per-transaction effect on only the two highest-volume exception types, which is 58 percent of volume."
  • Expected: the precedent scaled by the wave's actual volume, adjusted for the differences you can name. Basis: "same per-transaction effect across all exception types, discounted 15 percent for supplier-onboarding's lower structural similarity to the precedent."
  • High: the expected case plus second-order effects, and every second-order effect labeled speculative in the document itself. Basis: "expected case plus reduction in late-payment penalties and the early-payment discount capture that becomes possible below a three-day cycle. Speculative: not modeled in payback, shown for completeness."

Why ranges beat points is worth being blunt about. A point estimate is a promise. The moment you write "this will deliver 165,000 in annual savings," you have created the number against which you will be judged, and the judgment will be binary. A range with stated bases is not a promise; it is an analysis, and it invites the reader into the reasoning rather than daring them to test the result. Ranges also make you honest with yourself, because writing the basis for the low case forces you to confront the future where things go mediocre rather than badly, which is the most likely future and the one nobody models.

Then the rule that gives the range its teeth, stated in the document in plain words: the low end must still clear the hurdle. If your case only works at the expected or high number, it is not a case, it is a hope with three decimal places. Put it in the text: "At the low case, payback is 26 months, inside the 30-month hurdle. We are asking you to fund something that works even if we are substantially wrong." No sentence buys more credibility per word, and it costs nothing if your case is real.

4. Full Costs: the section that earns the trust

This is where cases are won, because it is the section every failed case got wrong in the same way. Most AI business cases contain two cost lines: the license and the implementation quote. Your case enumerates the lines the others omit, and enumerating them is the clearest signal you can send that this document was built by somebody who has done this before.

  • Licenses and consumption. Separate the fixed seat or platform fee from the variable consumption charge, and state the variability: cost per case at the measured pilot rate, the range observed, and what drives the top of that range (retries, long documents, higher-volume months). Consumption billing is the line that surprises companies in month five, and that surprise is what Gartner's escalating-cost cancellation driver looks like from inside a building. Model it as a band, and say what happens to the case at the top of the band.
  • Integration and provisioning. Connectors, identity and access provisioning, environment setup, the data plumbing the redesign requires, and the internal engineering hours those consume at a loaded rate, not a notional one.
  • The verification tax. The ongoing human cost of checking AI output at the designed human gates. Price it from the gate arithmetic you already did: gate volume per month, average review minutes per item, loaded hourly cost, sampling percentage for items that pass without full review. This is the line most cases pretend is free, and pretending it is free is exactly what you learned to catch vendors doing in Level 1. Counting it against yourself, in a named line, is the credibility move a CFO's analyst notices immediately and remembers for years.
  • Change and training. Priced per role from your People Readiness Atlas, not as a single training budget. Forty people needing four hours is a different number from four people needing forty, and the second is usually where the real cost lives. Include supervisor time, SOP rewrites, and the productivity dip during the ramp.
  • Transformer and program-management hours. Your own time, your process designers' time, workshop facilitation, steering-committee overhead. Internal hours are not free just because they are already on payroll; they are the scarcest resource in the program, and pricing them at a loaded rate turns "we can absorb it" into an actual capacity decision.
  • The ongoing quality regime. The monitoring rota, sampling hours, drift checks, quarterly re-baselining. This is run-rate, it is permanent, and it is the difference between a control that exists on a slide and a control that exists on a calendar.

Then the distinction that determines which budget your program hits: one-time versus run-rate, separated and totaled separately. One-time costs (integration, initial training, configuration) may be capitalized or come from a transformation fund; run-rate costs (licenses, consumption, verification, monitoring) hit an operating budget somebody owns forever, and that person needs to know their number before the meeting, not after. Cases that blend the two into one "total investment" figure are not simplified, they are unusable, and a finance reader has to disassemble them before they can begin. Do the disassembly yourself and you have saved your reader the only work they were going to have to do by hand.

5. The Adoption Ramp, Modeled

Benefits do not begin in week one. Everybody knows this, and almost no business case reflects it, because reflecting it makes the payback period longer and the author is afraid of the longer number. This is the single most common place where an otherwise honest case quietly lies.

State the ramp as a curve justified from your own pilot history. An illustrative shape, and yours should be yours: 0 percent of modeled benefit in weeks 1 to 4 (configuration, parallel running, nobody has changed how they work yet), 40 percent in weeks 5 to 10 (early adopters converted, the rest still shadow-processing), 80 percent in weeks 11 to 16 (SOP fully switched, exceptions still routed manually by the cautious), steady state thereafter. Under it, one sentence: "This shape is taken from the invoice-exception rollout, where the measured curve reached steady state in week 15."

Here is what modeling the ramp does, and it is not what you fear. It turns a fantasy payback of four months into an honest payback of nine. The honest nine survives; the fantasy four gets audited to death in month five, when the mid-year review compares actuals to a forecast that assumed instant adoption and finds a gap the size of a scandal. The fantasy four is not a faster case. It is a slower case with a hidden failure scheduled into it. You are choosing between defending nine months now and explaining a 300 percent variance later, to an audience that will not be in a generous mood.

A modeled ramp also converts future failure into future normality. When week 7 shows 38 percent of modeled benefit, that is not a disappointment requiring explanation; it is the plan, and you can say so in one line of a status report. Programs die from unexplained variance more often than from bad results, and a ramp curve is a variance-absorption device.

6. The Downside and the Kill Condition

This is the section no traditional business case has, and the one every post-scrap-year CFO is actually looking for. It answers, in numbers: what do we lose if this fails, and what stops the loss?

Three components. First, maximum exposure before the next gate, as a single number: "Maximum exposure before the wave-1 evidence gate is 140,000, comprising 95,000 of one-time spend committed by that date and 45,000 of run-rate incurred during the ramp." Second, the pre-committed kill condition, a testable statement with a date and a metric, straight from your decision-memo discipline: "If, at the week-16 gate, net measured cycle-time reduction on exception handling is below 20 percent against Baseline Pack v2.1, wave 1 stops and wave 2 does not start." Third, the salvage list, itemized: what survives a kill and what its residual value is. The cleaned vendor master. The documented exception taxonomy. The gate design and its trained reviewers. The integration connector, reusable by three other items.

Now understand what those three components do to the decision in the room. A case without a capped downside asks the committee to make a bet: commit the full amount, find out in a year. A case with a capped, pre-committed downside and an itemized salvage list asks them to buy an option: spend 140,000 to learn whether the remaining 400,000 is worth spending, with roughly a third of the 140,000 recoverable as reusable assets even in the failure case. Committees that will not approve bets approve options routinely, because an option has a known worst case and a bet does not. That structural difference, not eloquence and not enthusiasm, is what gets deep bets funded in a post-scrap-year organization.

There is a second effect that pays off much later. By pre-committing the kill condition in the funding document, you have made a future kill a fulfilled commitment rather than an admitted failure. The strategist who kills a wave at a pre-declared gate is executing the plan the committee approved; the strategist who kills a wave that had no kill condition is confessing. The difference in career consequence is enormous, and it is decided at the moment of funding, not at the moment of killing.

7. The Evidence Chain

Every number on the page carries a source. Not footnotes cluttering the page, but a workbook behind it in which every figure is traceable: the Baseline Pack by version, the audit register entry, your own prior Delta Table from the precedent project (with its confidence words intact, including the ones that say "indicative" rather than "measured"), vendor quotes with the assumptions they were quoted under, and the loaded-rate table behind every hour priced anywhere in the case.

Two rules make the evidence chain work. First, no number appears on the page that does not exist in the workbook, including the ones that seem too obvious to source. Second, the workbook goes to the CFO's analyst a week before the meeting, every time, forever. Not the summary. The workbook, with the tabs unhidden.

That second rule feels dangerous the first time and never again. Practitioners resist it because they imagine the analyst hunting for weaknesses to ambush them with. What happens is the opposite: the analyst finds the weaknesses in private, emails you about them, and you fix or explain them before the meeting. The committee then watches the analyst nod, and that nod is worth more than your entire presentation. Make the pre-brief standard practice for every case you submit, so that on the one occasion when your numbers are shakier than you would like, the pre-brief is not a signal of unusual anxiety but simply what you always do.

Wave 1 in Numbers: A Worked Example

All figures here are hypothetical and illustrative, built to show the shape of a defensible case rather than to be borrowed as benchmarks. The setting is the enterprise storyline: an eleven-item portfolio sequenced into waves, one proven precedent (the invoice-exception redesign, measured at a net 428,000 annual run-rate reduction), and a finance function that has watched an earlier program fail.

Section 1, baseline. Wave 1 covers two processes: invoice-exception handling at the second site and supplier onboarding. Combined measured annual run rate: 610,000, from Baseline Pack v2.1, measured over eleven weeks, method in tab 2, with two known limitations disclosed (the seasonal December volume spike is excluded; supplier-onboarding rework hours are estimated from a four-week sample rather than a full census).

Section 2, mechanism. Stated separately per item. Exceptions: the constraint is the batching queue, attacked by classify-and-route on arrival. Onboarding: the constraint is the three-way document check that stalls waiting on a human to reconcile inconsistent vendor records, attacked by pre-reconciliation against the cleaned vendor master, which is why this item depends on the vendor-master capability build.

Section 3, benefits as ranges. Low 95,000 per year, expected 165,000, high 240,000. Bases: low assumes precedent-equivalent effect on the two highest-volume exception types only and no onboarding benefit at all in year one; expected assumes precedent-equivalent effect across exception types plus a 15 percent discounted onboarding effect; high adds early-payment discount capture and reduced supplier-onboarding abandonment, both labeled speculative and both excluded from the payback calculation.

Section 4, full costs. One-time 218,000, run-rate 74,000 per year, split as follows.

Cost lineOne-timeRun-rate (annual)
Platform license (18 seats)021,000
Consumption at metered rate (band: 8,000 to 19,000)012,000
Integration, connectors, provisioning84,0000
Verification tax (gate hours, priced)031,000
Change and training, priced per role46,0000
Transformer and program-management hours72,0000
Quality regime (monitoring rota, sampling)16,00010,000
Total218,00074,000

The verification tax line deserves its own sentence, because it is the line the analyst will circle: 31,000 per year is 1,240 gate reviews per month at an average 4.5 minutes each, plus a 10 percent full-review sample of auto-passed items, at a loaded rate of 46 per hour, itemized in workbook tab 6. A case that prices its own oversight this precisely is making a statement about what kind of document it is.

Section 5, adoption ramp. 0 percent of modeled benefit in weeks 1 to 4, 40 percent in weeks 5 to 10, 80 percent in weeks 11 to 16, 100 percent from week 17, justified by the measured week-15 steady state of the invoice-exception precedent. Applied to the expected case, this delays roughly 47,000 of first-year benefit.

Payback. At the expected case, 14 months. At the low case, 26 months, inside the organization's 30-month hurdle. The document says so explicitly: this wave clears the bar even if the conservative case is what happens.

Section 6, downside. Maximum exposure before the week-16 evidence gate: 140,000. Kill condition: net measured cycle-time reduction on exception handling below 20 percent against Baseline Pack v2.1 at the week-16 gate stops wave 1 and blocks wave 2. Salvage list: the cleaned vendor master (required by three other portfolio items, residual value roughly 40,000 of avoided rework), the exception taxonomy and its labeled corpus, the gate design and eleven trained reviewers, and the integration connector.

The Analyst Pre-Brief, and the Number You Revise Down

The workbook goes out on a Tuesday, nine days before the committee. The analyst replies on Thursday with two challenges, which is a good outcome, because two specific challenges means he read it.

Challenge one: the consumption estimate's variance. "Your 12,000 sits near the bottom of your own 8,000 to 19,000 band. What happens to payback at 19,000?" The answer comes straight from the workbook: at the top of the band, expected-case payback moves from 14 months to 15, and low-case from 26 to 28, still inside the hurdle. The band is disclosed on the page; the sensitivity is tab 9. Answered in one email, and it strengthens you, because the sensitivity existed before he asked.

Challenge two: the training-hours assumption. "You have 46,000 of change and training against 18 seats. Show me the per-role build." Tab 4 shows it: three roles, different hour loads, supervisor time included. But looking at it with a stranger's eyes, you notice something you had been slightly generous about. The supervisor coaching hours were priced for a full sixteen weeks when the ramp says the cautious-routing period ends at week 16 and coaching intensity drops sharply after week 10. Honest correction: 46,000 becomes 41,000.

You revise it down. Voluntarily. Before anyone made you. Then you email the analyst a one-line note saying what you changed and why, and in the meeting he mentions it unprompted: "They found an overstatement in their own training line and corrected it before submission."

Understand what that costs and buys. It costs 5,000 of padding you were never going to spend. It buys a room that now believes the other numbers, and, more durably, the benefit of the doubt on your next four cases, including the one where a number turns out wrong for a reason nobody could have foreseen. Credibility in a finance relationship is not built by being right. It is built by being caught being honest when you had the option not to be.

The Language Discipline

The prose in a business case is load-bearing, and three constructions will cost you more than any number. "Transformational" and its family (game-changing, step-change, revolutionary) tell a finance reader that the author had no number here and reached for an adjective. "Up to X percent" is the most recognizable vendor construction in the language: it discloses the best case and hides the distribution, and using it in your own document says you adopted the vendor's rhetoric wholesale. Unlabeled vendor benchmarks ("industry leaders see 40 percent reductions") are worse than nothing, because the first thing a good analyst asks is which industry, which leaders, measured how, and you will not know.

Apply to your own prose exactly the skepticism you apply to vendor decks. The replacement sentence is not glamorous and is worth more than any adjective available: "We estimate, based on our own measured pilot, with these caveats." Every clause works. "We estimate" claims the right register. "Our own measured pilot" names the evidence and its provenance. "With these caveats" promises that weaknesses are disclosed rather than discovered. Write that sentence, or its equivalent, wherever you are tempted to write "significant."

The same discipline governs confidence words. Carry them across from your measurement practice: a measured number says measured, an estimated number says estimated, an indicative number says indicative, and no editing pass may ever upgrade one to another for the sake of the executive summary. That upgrade is the most common integrity failure in business-case writing, and it usually happens on a Thursday night, in good faith, in the name of clarity.

The Hockey Stick and the Process Debt It Leaves

This failure story deserves sympathy rather than mockery, because the person in it did nothing his organization had taught him was wrong. All figures are illustrative.

A program lead at a mid-sized manufacturer builds a case for a six-use-case AI program during a quarter when the board has publicly committed to AI leadership. He is thorough by the standards he knows: he interviews process owners, collects vendor quotes, builds a clean deck. The case projects 2.1 million in annual savings across document processing, quality-report drafting, supplier-email triage, maintenance-log summarization, purchase-order matching, and internal knowledge search. Benefits begin in month two, because that is when the vendor says the platform will be configured. The cost side is the license total plus the implementation quote, which is what the vendor's own ROI template asked for. Payback: seven months. It is approved with very little challenge, which is the first thing that should have worried everybody and worried nobody.

By month eight, the picture is this. Measured savings, to the extent anybody can measure them without a baseline, are around 180,000 annualized, and half of that is contested. Two of the six use cases are still in configuration. Adoption in the three live ones sits somewhere between 30 and 60 percent, and nobody can say precisely because usage reporting counts sessions rather than completed work. Underneath it all sits the omission that turned a disappointing program into a damaging one: the verification staffing was never budgeted. The human gates on document processing and purchase-order matching were designed, correctly, into the workflow. But no hours were funded for them, so reviews are performed by people who already had full jobs, which means they are performed fast, which means the gates under-run their own review standard. Quality problems follow in month six. A supplier is paid twice. A quality report goes out with a fabricated equipment reference number in it.

The mid-year review does not just cut the program. Cutting a program is normal and survivable. What the CFO does instead is institute a rule: all AI investment proposals now require Finance co-authorship, with a mandatory analytical review before submission to any committee.

Run the arithmetic on that rule, because it is the real cost of the story. Co-authorship adds roughly nine weeks to every AI business case from that day forward, and the company runs maybe fifteen AI proposals a year across its functions. That is a permanent tax of over a hundred weeks of aggregate delay annually, paid by people who had nothing to do with the original case, on proposals individually more rigorous than the one that caused the rule. Good proposals are slowed exactly as much as bad ones.

This is process debt, and it is the hockey stick's true legacy. One inflated business case does not just waste its own money. It taxes the organization's future speed, permanently, through a control that will outlive everyone involved. When you are tempted to round a benefit up or leave the verification line out because the payback looks better without it, the honest accounting of that temptation is not "I might be a bit wrong." It is "I might be the reason this company takes nine extra weeks to fund anything, forever."

Notice also which repair the CFO chose. Not better technology evaluation, not more pilots: she demanded that Finance be in the room while the numbers were built, which is a precise diagnosis. The failure was in the construction of the case, not the choice of tool. The Wave Business Case is what a program lead writes to make that rule unnecessary, or, where the rule already exists, to make co-authorship a formality rather than an excavation.

The Portfolio-Level Case

The Wave Business Case is one of two documents you owe. The other is the portfolio-level case, and it is not the sum of the item cases. Summing them produces a number that is both wrong and unpersuasive: wrong because shared capability spend gets counted twice or not at all, unpersuasive because it presents eleven independent bets rather than one designed program. The portfolio case answers a different question: what are the economics of the whole program, including the parts no single item would ever justify?

Allocating the capability builds

The hard problem is shared infrastructure. The vendor-master cleanup, say 180,000, is required by seven of the eleven items and is a prerequisite for two. No individual item's case can carry it: charged whole to supplier onboarding, that item's payback goes from 16 months to well past the hurdle and the item dies, taking the capability and six other dependent items with it. There are two defensible conventions, and knowing which to use in which room is a real skill.

  • Allocation: charge each of the seven beneficiaries one-seventh of the 180,000, roughly 25,700 apiece, and show each item's payback with its allocated share included. This lands well with a finance culture that thinks in fully-loaded unit economics and with organizations that charge shared services out to business units. No item looks artificially cheap. Its weakness is fragility: if two beneficiary items are later dropped, the remaining five inherit a larger share and their cases have to be redone.
  • Program overhead: hold the 180,000 at the program level as a capability build, with the beneficiary arithmetic shown explicitly ("prerequisite for items 3 and 7, materially improves items 1, 4, 6, 9, and 11; the sum of those items' expected annual benefit is 840,000"). This lands well with a committee that thinks in program terms and with a CFO who wants the capability decision made once, on its own merits, rather than smuggled into seven item cases. Its weakness is that overhead lines attract scrutiny and look like a slush fund unless the beneficiary arithmetic is airtight.

Pick one, state which convention you used and why in a single sentence, and be able to produce the other view on request in the workbook. The sentence matters more than the choice. An analyst who sees a stated convention concludes you thought about it; one who sees an unexplained shared-cost treatment starts looking for what else was moved around quietly.

The reserve

Ask for a reserve, name it, and size it as a percentage of program spend rather than hiding contingency inside individual line items. Padding every line by 15 percent is the amateur approach: it inflates every number, makes every line less defensible, and gets caught. A named reserve, say 12 percent of program spend, with a stated rule for releasing it (who approves, against what evidence), is a normal instrument that capital committees see constantly in construction and information-technology programs and know how to evaluate. It also gives you somewhere honest to put the thing you actually fear, which in AI programs is integration effort and data remediation discovered in flight.

The portfolio hurdle

This is the most important argument in the portfolio case, and the one that has to be made in advance or it cannot be made at all. The program clears its bar if the portfolio does, not if every item does. Write it into the funding document plainly: we expect to fund eleven items, we expect roughly two to three to be killed at their evidence gates, kill conditions and maximum exposures are pre-committed per item, and the portfolio still clears the 30-month hurdle with those kills assumed. Then model it that way, with the assumed kill rate baked in as the base case rather than as a downside scenario.

Two things happen when you do this. First, you become able to fund deep bets at all. An item with a genuinely uncertain mechanism and a large potential payoff can never clear an every-item-must-succeed standard; it can comfortably sit inside a portfolio that has priced expected failures. Second, and this is the part that pays off eighteen months later, you have made future kills survivable. The kill you execute at a gate in month nine is not a surprise, not a scandal, and not a referendum on the program. It is the portfolio behaving exactly as the funding document said it would.

Say the hard version out loud to the committee, once, at funding: a portfolio in which nothing fails was underambitious. That sounds like a slogan and is actually a technical claim about selection. If every item cleared, you selected only for certainty, which means you systematically skipped the items where the largest value was, because those are exactly the items nobody can be certain about in advance. A CFO who has lived through a scrap year will not find this reckless; she will find it the first honest thing anybody has said about an AI portfolio in three years. And it is the argument you need intact when the first documented kill lands, later in this chapter.

One more piece of arithmetic belongs in the portfolio case, and it is the bridge to what comes next. Look back at the wave-1 cost table and notice which lines are largest: integration at 84,000 and program-management hours at 72,000, both of which are consequences of a sourcing decision that was made before the case was written. Whether you build the connector, buy a platform that ships it, or partner with someone who owns it changes the biggest lines in every case you will ever write, and it changes them at portfolio scale, not one item at a time. That decision is next.

What to Do Monday Morning

Take the live case you have in flight, the one sitting in a deck waiting for a committee date, and rebuild it in this order. Expect a working day, and expect the payback number to get worse.

  1. Move the baseline to the top and the mechanism to second. Open with the measured current-state run rate, its method, and its version number. Then write the mechanism, one paragraph per item, no vendor language. If you cannot write the mechanism, stop: you have found a problem worth more than the rest of this exercise.
  2. Enumerate the six omitted cost lines and price your verification tax. Licenses and consumption as a band, integration and provisioning, verification tax, change and training per role, transformer and program-management hours, ongoing quality regime. Price the verification tax explicitly from gate volume, review minutes, and loaded rate. Separate one-time from run-rate and total them separately.
  3. Convert every point benefit into low, expected, and high, each with its basis in one sentence. Then test the rule that matters: does the low case still clear your organization's hurdle? If it does not, you do not have a case yet; you have a candidate for the next wave or for a smaller first gate.
  4. Model the adoption ramp from your own pilot history. Write the curve, write the one-sentence justification, and recompute payback with the ramp applied. Whatever the new number is, that is the number you will defend, and it is now defensible.
  5. Write the downside section. Maximum exposure before the next gate, the pre-committed kill condition with its date and metric, and the itemized salvage list with residual values. This section is usually 120 words and it changes the character of the decision from a bet to an option.
  6. Send the workbook to the CFO's analyst a week before the meeting. Unhide the tabs, invite the challenges, and when one exposes a number you were generous about, revise it down before anyone makes you, and say so.

Key Takeaways

  • Diagnose the real reason finance rejects AI cases: not hostility to uncertainty but recognition of a failed shape, single-point benefits, license-only costs, no baseline, no mechanism, and a payback assuming instant adoption.
  • Build every case in the Wave Business Case's seven sections in order: baseline, mechanism, ranged benefits, full costs, modeled adoption ramp, capped downside with kill condition, and evidence chain.
  • Open with a measured baseline cited by version, because a case that opens with evidence of having gone and looked has already separated itself from most of what reaches a CFO's desk.
  • State the mechanism in one paragraph, since finance funds mechanisms it can argue with, not technologies it must take on faith, and an item whose mechanism will not compress is usually an item in love with a capability.
  • Replace point estimates with low, expected, and high ranges carrying stated bases, and enforce the rule that the low end must still clear the hurdle or the case is a hope.
  • Price the verification tax, change spend, internal hours, consumption variability, and the ongoing quality regime in your own case, separating one-time from run-rate, because counting your own costs the way you demand vendors count theirs is the fastest credibility you can buy.
  • Model the adoption ramp honestly, since a nine-month payback survives a mid-year review while a four-month fantasy gets audited to death, and remember that one inflated case leaves process debt that taxes every future proposal in the organization.
  • Argue the portfolio hurdle in advance, with an expected kill rate baked into the base case, because that argument is what permits deep bets to be funded and what makes documented kills survivable later.