Reporting AI ROI and Risk to the C-Suite
The slot says 20 minutes and it is the fourth item on a nine-item agenda. The chief executive officer (CEO) has been in the building since seven and has a board call after lunch. The chief financial officer (CFO) has a laptop open, with a spreadsheet on it that is not yours. The general counsel is reading something on her phone. You prepared for six days: the scorecard, the quadrants, the trend arrows, the confidence words, twenty-two pages of workbook behind them, every number defensible. And in about ninety seconds you will discover the thing nobody teaches strategists: this room is not going to evaluate your evidence. It is going to construct a story about your program out of whatever you put in front of it, in the first four minutes, using its own arithmetic and its own fears. If you did not supply the frame, it comes from somewhere else: the loudest anecdote in the room, the department head who had a bad experience in March, or the vendor pitch the CEO sat through last Thursday. This lesson is about supplying the frame.
The Meeting Where Programs Are Quietly Defunded
Programs rarely die by decision. Ask anyone who has watched a serious initiative disappear and they will struggle to name the meeting where it was cancelled, because usually there wasn't one. What happened instead was an accumulation of quarterly reviews in which the program was present but not legible: numbers shown, heads nodded, nobody leaving with a sentence they could repeat. Then a budget cycle arrived, someone needed money, and the program was on the list of items whose value nobody could articulate under time pressure. That is a failure to have been understood, collected quarterly and cashed in once.
The previous lesson built the instrument: the AI Value Scorecard, four quadrants (efficiency, quality, adoption health, risk posture), each metric carrying a trend and a confidence word, designed so a program's honest position fits on one page. It is necessary and not sufficient, because a scorecard answers "what is true?" and the quarterly review is asking three other questions, held privately by three different people, which it will answer with or without your help.
Here is the intelligent mistake almost every strategist makes. They treat the report as an information problem: get the numbers right, present them clearly, and the correct conclusion follows. What they have walked into is an interpretation problem. Executives are not short of information; they are short of time and long on pattern recognition. Each is running a fast, mostly unconscious sort: is this working, is it worth what it costs, is it going to hurt us? They finish that sort in the first few minutes, then hold the conclusion loosely for a quarter and firmly after two.
Nobody leaves that room without a story about your program. The only variable is whether you are the one who wrote it.
You are also reporting into a market that has made up its mind about your category. Your executives read the same coverage everyone else did: MIT's 95 percent of enterprise generative AI pilots with no measurable profit-and-loss return, S&P Global's 42 percent of companies scrapping most of their AI initiatives in 2025 (up from 17 percent), Gartner's forecast that over 40 percent of agentic AI projects will be cancelled by the end of 2027 with unclear business value among the leading reasons, and McKinsey's split screen: 88 percent of organizations using AI regularly, only about 39 percent able to attribute any earnings impact to it. That is your audience's prior. They expect your program to be one of those, not from hostility but from base rates, so nothing you say in the first four minutes is heard neutrally.
The report's job, then, is not to convey the scorecard. It is to make the program's honest position legible enough, in the audience's own arithmetic, that continuing to fund it is rational rather than an act of faith. Faith is a renewable resource for about three quarters. Rationality renews indefinitely.
Three People, Three Questions
Run the diagnostic that shapes everything: who is in the room, and what is each of them holding? Three audiences, distinguished not by seniority but by the question each one cannot stop asking.
The CEO: "Is this working, and is it fast enough?"
The chief executive holds a competitive clock. Their fear is not wasted money; it is that the company is being outpaced while its AI program produces respectable quarterly percentages. What they need is story: what the organization can now do that it could not a year ago. A CEO does not retain "cycle time down 47 percent." A CEO retains "the exceptions that used to sit for nine days now clear in four, so we stopped losing the suppliers who called twice a week." One of those can be repeated to a board member in a corridor.
The CFO: "Is the money real, and what is it costing?"
The chief financial officer holds arithmetic and a suspicion: that your benefits are gross and your costs partial. Their fear is signing off on a number that turns out to be a hope with a decimal point. A CFO who can see the shape of your uncertainty will fund you through a bad quarter. A CFO who suspects you are hiding the shape will not fund you through a good one.
The risk-holder: "What could hurt us, and who is watching?"
This is the chief operating officer (COO), the general counsel, the chief information security officer (CISO), or the chair of the board's risk committee, sometimes all four. They hold the headline that has not been written yet, and their fear is learning of a problem at the same moment as the regulator, the journalist, or the customer. They are not asking you to promise nothing will go wrong, but whether, when something does, the organization hears it from you first.
| Reader | What satisfies the question | What reads as evasion |
|---|---|---|
| CEO | A concrete before-and-after, capability built, an honest pace statement | Percentages without a scene; "on track" with no definition of the track |
| CFO | Net value with confidence composition, full costs, avoided spend, the breaking assumption | Gross benefits, cumulative totals, costs limited to licences |
| Risk-holder | Incidents with detection latency, controls firing, artifact currency against dates, concentration | A green box labelled "no issues"; risk discussed only when asked |
One report serves all three only if it is organized by those questions rather than by the program's internal shape. Which brings the near-universal structural error: reporting by workstream. The pack opens with a data section, then portfolio, then change and adoption, then governance, because that is how the program is organized. It is a perfectly rational document that maps to nobody's question. The CFO sits through the data update waiting for a number. The risk-holder waits through the portfolio slide for the incident count. The CEO, who wanted to know whether the company is moving, receives four progress bars. Everyone is served eventually, in fragments and out of order, and the room assembles its own story from the pieces. Structure is not cosmetics: the order of your blocks decides which question gets answered while attention is highest.
The Artifact: The Quarterly AI Report
The Quarterly AI Report is five blocks in a fixed order, each written for a named reader, in the same sequence every quarter, forever. The rigidity is the point: a room that has seen the same five blocks three quarters running knows where to look and starts deciding. Novelty in a governance document is not creativity, it is friction.
Block 1: The Position
Three sentences at the top of page one, written the way you would say it privately to a peer you respect: measurable value in two of five functions, a quarter behind plan in the rest for reasons you understand, value real but smaller than the original case assumed, largest exposure named.
The discipline is honesty calibration, with a mechanical test: the position must be one an insider would recognize as accurate. The room contains insiders, including the sponsor whose function was slower than hoped and the analyst who has read your workbook more carefully than you have. If your position line is a shade sunnier than the truth those people carry, you have not fooled anyone. You have told the room your reporting is promotional, and everything downstream is discounted at a rate you cannot see.
Block 2: Value, in CFO arithmetic
Roll up the efficiency and quality quadrants into money, and observe five rules.
- Show the CLEAR-only subtotal beside the total. The confidence vocabulary from Level 3 (CLEAR when a delta survives every control and its mechanism is named, PROBABLE when the window limits it, SUGGESTIVE when the volume is too thin to weigh) belongs in the boardroom, not only in the measurement pack. A total of 412,000 with a CLEAR-only subtotal of 260,000 tells a CFO what they need: the number you would defend under audit, and the number including everything you believe. Present only the larger figure and the discount is applied by someone else, at a rate you do not choose.
- Show full costs, including the lines nobody volunteers. Technology and licences, but also change and training, the verification tax (the priced human hours spent checking AI output at designed gates), and program overhead. A CFO knows those costs exist. The only question is whether you do.
- Report what you have avoided. Your stage-gate log records the items you killed and, computed by finance from each item's approved case, the counterfactual spend those kills prevented. A program reporting 340,000 of avoided spend beside 412,000 of realized value is reporting a portfolio; one reporting only the 412,000 is reporting a wish, because portfolios have losses and wishes do not.
- Volunteer one number downward. Every quarter, find the softest assumption in your own pack and revise it against yourself before anyone asks. It is the cheapest credibility available. A strategist who corrects their own optimism in public is one whose remaining numbers get believed.
- Name the assumption that would break the projection. "This holds if the two remediate-first functions start redesign in Q4; a one-quarter slip takes roughly a fifth off next year." You have told the room where to look, which is what an advisor does.
Block 3: Momentum, in CEO story
One concrete before-and-after, told as a scene. Not a chart: a scene with a person in it. The invoice exception that used to sit at nine days at p90 (the 90th percentile, the boundary of the slowest tenth of items, the ones that age in a queue while a supplier calls twice a week) now clears in four, and that supplier has stopped calling. Thirty seconds, and it is the sentence the CEO uses when a board member asks what the AI program has done. Give them a percentage and they repeat nothing, which is the same as reporting nothing.
The capability built. Trained people, working method, reusable artifacts: the baseline pack, the redesign method, the gate charter, the champion network. These outlive any individual tool, and a CEO deciding whether the company is building an advantage or renting one needs them counted. Twenty-two people who can run a baseline is a balance-sheet fact dressed as a training statistic.
The honest pace statement. Where the program is slower than hoped, and why: dependencies, the enterprise resource planning (ERP) upgrade window that blocks integration until March. The rule underneath is worth memorizing: a CEO told the truth about pace once will accept it thereafter; a CEO surprised by pace will intervene. Intervention almost always takes the form of a reorganization, a new sponsor, or an external review, none of which speed anything up.
Block 4: Risk posture, in the risk-holder's frame
Annotate the risk quadrant for a reader who thinks in exposure.
- Incidents, with detection latency. Not just how many, but how long between occurrence and detection. Six days is a number a risk committee can act on; "one incident, resolved" is not.
- Guardrail events, framed as controls working. Four cases where a designed control stopped something is good news reported as good news. Programs hide these on an instinct that any event sounds bad, throwing away their best evidence that the control layer is real.
- Artifact currency against the calendar. Fairness checks, model documentation, data protection assessments: which are current, which overdue, by how many days, against which obligation. The EU AI Act calendar is dated (general-purpose AI obligations since August 2, 2025, AI-content transparency from December 2, 2026, high-risk Annex III from December 2, 2027), so compliance readiness is a schedule, not a posture.
- Vendor concentration against a stated tolerance. Not "we use three vendors" but "vendor A now carries 61 percent of AI-touched volume against a stated tolerance of 50, flagged amber, exit estimate last refreshed in Q1."
Then the section that does more for your standing than anything else in the document: "What would worry me most", written by you, in the first person, one short paragraph naming your own top worry and what you are doing about it. Not the risk register: yours, the thing you think about on the drive home.
It looks like a flourish and is not. Every executive holds a mental model of you, and by default it is advocate: the person whose job is to make this program look fundable, whose assessments are discounted accordingly. The first-person worry paragraph is the only device that reliably converts you to advisor, because advocates do not volunteer their own fears. It pays its dividend on the quarters when the news is good: an advisor's good news is believed at face value, an advocate's is examined.
Block 5: The ask and the decision
End with something the room must actually do: continue as planned, fund wave 3 at a stated amount, resolve the concentration question by naming who owns it and by when. Price the alternatives, so the room chooses between costed options rather than reacting to a request. A report without an ask trains the room to treat the meeting as informational, and informational meetings get shortened, then delegated, then dropped "since the pack is circulated anyway." Put an ask on the last page every quarter, even "confirm the risk committee owns the concentration decision." The ask makes the meeting a decision forum, and decision forums survive.
The Preparation Craft: What Happens Before the Room
The document is maybe half the work. The other half happens in the ten days before the meeting, and it is where experienced strategists spend their effort.
The pre-brief circuit
You learned the principle in Level 2 as the pre-wire and in Level 3 as the hostile-analyst read: no ambushes, in either direction. At C-suite altitude it becomes a fixed circuit, run every quarter:
- The CFO's analyst gets the full workbook a week early. The highest-return hour you will spend. The analyst finds your weak assumption in private, emails you, and you fix or explain it before the meeting. In the meeting, the analyst nods, and that nod is worth more than any slide you own.
- The risk-holder gets block 4 early, with the incident detail behind it. A general counsel who first learns of an incident in a room containing the CEO will remember the format of that discovery longer than the incident.
- The sponsor is never surprised by anything in the pack. If a function is named as slow, its executive knew a week ago, agreed the wording or registered disagreement, and is not learning it in front of their peers.
The circuit costs perhaps three hours a quarter and converts the meeting from a performance into a confirmation: every hard conversation already had in private, with the person who owns it.
Anticipatory answering
List the five hardest questions the room could ask, the ones you hope nobody asks. "Why is the CLEAR subtotal only 63 percent of the total?" "What happens if the two slow functions never start?" "If vendor A doubled its price at renewal, what would we do?"
Then answer three of them inside the document, in the block where they belong, before anyone asks. This exploits a real asymmetry: a challenge already addressed in writing reads as competence; the same challenge answered live reads as defense. Identical content, opposite effect, because written in advance signals you saw it coming while answered live signals you are holding a position. Two sentences each, and the room's questions improve, because you have moved the conversation past the obvious objections into the ones worth executive time.
The bad-quarter protocol
Sooner or later the numbers are poor: adoption fell, a pilot failed, value came in at half the projection. The protocol is the opposite of instinct.
- Report early, not at the meeting. The moment you know, the sponsor knows, and the CFO and risk-holder know within days. Bad news travelling at the speed of the calendar looks like concealment even when it is not.
- Lead with the diagnosis and the response, not the number. "Adoption in claims fell to 34 percent because the redesign shipped without the exception path, and the fix is scoped at four weeks" is a program working. "Adoption fell to 34 percent" is a program failing. Same fact.
- Bring the decision you are recommending. Kill it, re-scope it, extend it with a condition, or accept a lower run rate. A bad number with no recommendation invites the room to invent one, and rooms invent blunt instruments.
The return sounds like sentiment and is not: the credibility earned by reporting one bad quarter well exceeds the credibility of three good ones. Three good quarters are consistent with a program working and equally consistent with a program selecting its evidence. One bad quarter, reported early with a diagnosis and a recommendation, is consistent with only one of those. And the surprise, not the number, is what costs you: never let a bad quarter arrive as news in the room.
Language discipline
- No "transformational." Nor "game-changing," "unlocking," or "journey." Executives have heard these attached to every failed initiative of the last decade, and adjectives that outrun evidence make the evidence sound weaker.
- No vendor benchmarks presented as your results. If a figure comes from a supplier's case study, say so in the same sentence or leave it out. One conflation, discovered, costs you the whole pack.
- No percentages without denominators. "Adoption is at 71 percent" means nothing until the room knows 71 percent of what, measured how, among how many people.
- No cumulative-savings number that only ever rises. The specific banned construction. A running total since inception can never decrease, never reflects a bad quarter, and never contains a loss, which is why it feels so good on a slide. Any CFO who has seen one knows what it is: an instrument built to be immune to reality. Report the quarter and the run rate.
Worked Example: Norvik Group's Q3 Report
Norvik Group, the mid-market industrial distributor whose scrap year opened this level, is four quarters into a governed program. Every figure below is illustrative.
Block 1, the position. "The program is delivering measurable value in two functions, invoice exceptions and service triage, and is a quarter behind plan in the two functions we sequenced remediate-first, for reasons we predicted and can evidence. Realized value this year is 412,000 against a full cost of 760,000, so we are not yet net positive and did not expect to be until Q1. Our largest open exposure is a vendor concentration we chose deliberately and have not repriced since Q1."
Block 2, value. Realized value year to date: 412,000 total, of which 260,000 is CLEAR, the remaining 152,000 PROBABLE with its two soft assumptions listed. Full costs: 470,000 technology (licences, consumption, integration) and 290,000 change (training, champion time, the verification tax at the invoice gate, program overhead). Avoided spend: 340,000, computed by finance from the approved cases of the two items killed at gate this year. And one number volunteered downward: the analyst pre-brief found the redeployment assumption behind 38,000 of efficiency value credited freed hours at a rate the function had not actually redeployed, so it was revised to 24,000 before the pack was finalized, with a one-line note saying so. The block closes with the breaking assumption: "the full-year projection of 900,000 assumes the two remediate-first functions begin redesign in Q4; a one-quarter slip reduces it to roughly 720,000."
Block 3, momentum. The scene: "An invoice exception that sat at nine days at p90 in January now clears in four. The practical meaning is that the supplier who used to call our accounts team twice a week to chase aged items has not called since May." Capability built: 22 people trained to run a baseline, a redesign method used in three functions, a gate charter with four quarters of decisions logged, a champion network of eleven. The pace statement, unhedged: "We are a quarter behind in claims and procurement. The cause is the data remediation we sequenced first and the ERP upgrade window, which blocks the procurement integration until March. We knew both; we did not accelerate them; we do not recommend trying."
Block 4, risk posture. One incident, a mis-scored batch in service triage, detected six days after occurrence, root cause a stale grounding pack, with detection latency now itself tracked against a two-day target. Four guardrail events: three where the confidence threshold correctly refused to auto-process, one where the escalation rule caught a document type outside scope. One overdue artifact, named: the fairness check on the service triage model, due June 30, now 41 days late, owner named, scheduled for September 12. Vendor A amber: 61 percent of AI-touched volume against a 50 percent tolerance, exit estimate last refreshed in Q1. And the first-person paragraph, printed as written: "What would worry me most is the concentration. We are deepening our position with vendor A faster than we are refreshing our estimate of what leaving would cost, which means our exit number is quietly becoming fictional. I have asked for a refreshed exit estimate by October, and I would like the risk committee, not me, to decide the tolerance."
Block 5, the ask. Approve wave 3 funding at 310,000, with the alternative priced (deferring one quarter saves 310,000 this year and pushes roughly 240,000 of benefit into next year). And decide the concentration question: accept 61 percent with a documented rationale, or fund a second-source pilot at an estimated 85,000.
What happened in the room
The meeting took 22 minutes of its 20-minute slot. The CFO's opening challenge, which was going to be whether the 412,000 was gross or net of the redeployment assumption, was already answered in block 2, including the downward revision; she said "I saw that, thank you" and moved to the full-year projection, where the breaking assumption was stated, so the conversation went straight to whether Q4 redesign starts were realistic. The CEO asked about the supplier who stopped calling, then whether the method could reach the branch network by year end, which is a growth question and the best signal a chief executive can give you. The general counsel had read block 4 three days earlier and had one comment about the overdue fairness check.
The concentration question was taken up by the risk committee as its own item, exactly what the first-person worry paragraph was engineered to produce: a risk owned by a body with authority to decide it rather than by the strategist who noticed it. Wave 3 was approved. Total elapsed time on a program spending three quarters of a million dollars: 22 minutes, because everything hard had already happened in private.
The Failure Story: The Good-News Quarterly
A different program, a competent strategist, three consecutive excellent quarterly reports. Each opens with a cumulative-savings figure that has risen every quarter: 180,000, then 430,000, then 690,000. Each carries a portfolio slide where every item is green or amber-trending-green. Each is well made and honestly intended; nobody is lying. There is simply no risk section, because the risk quadrant held nothing dramatic and writing up minor incidents felt like manufacturing anxiety. Adoption figures appear without denominators, and two pilots that quietly stopped go unmentioned because they were never formally launched. The executives are pleased and the program is praised at an all-hands.
In quarter four, two things surface within a fortnight. A model incident in a customer-facing process is found by a customer complaint rather than by monitoring, and turns out to have been running for five weeks. And a required assessment tied to a dated regulatory obligation is two months overdue with no owner assigned. Both are individually manageable: the incident affected a modest number of cases and is fixable in weeks, and the artifact is a documentation gap, not a breach. Had either appeared in a report that routinely carried such material, the response would have been proportionate.
That is not what happens. The executives' reaction is not calibrated to the events at all. It is calibrated to what they have just realized about the reporting, and the CEO says the sentence every strategist should fear, flatly, in a meeting that was supposed to be about the incident:
"What else has not been on these pages?"
Nothing was concealed, and every individual number in three quarters of reporting was accurate. What the room has discovered is that the reporting was selecting: the document had a shape that admitted good news and no place to put anything else, so those three quarters carry far less information than everyone had assumed. Every previous report is discounted in a single moment.
The strategist's answer is, as it happens, honest and nearly complete. It does not matter, because it cannot be verified quickly enough to save the mandate: verification would mean re-examining three quarters of work, and no executive team has that appetite. They take the cheap alternative: an external review, a reporting standard imposed from outside, a budget reduced pending confidence, a sponsor who becomes less available. The program is not cancelled. It is supervised, which is slower and, for the person running it, considerably worse.
The lesson generalizes past AI entirely. A report that never contains bad news is not reassuring. It is uninformative, and the room eventually notices. These rooms are staffed by people who have run programs for decades and know the base rate of things going wrong; a document that contradicts it is not read as excellence for long. Put the risk section in while it is boring, which is the only time you can install it for free.
Reports inform decisions, and decisions need somewhere to land: a body that can take the concentration question, decide it inside a month, and own the consequence. That body is the governance committee, and building one that decides rather than discusses is next.
What to Do Monday Morning
- Re-cut your last quarterly pack into the five blocks. Re-sort every element of the deck you most recently presented into position, value, momentum, risk, ask. Two things become obvious: how much of the pack served no reader's question, and which block you have almost nothing for. Usually it is risk; sometimes it is the ask.
- Compute your avoided-spend number from the gate log. Pull every item killed or rejected at a gate this year and ask finance, not yourself, to compute the counterfactual from each approved case. Even a conservative figure changes what your report is.
- Write the "what would worry me most" paragraph in first person. Your actual top worry, what you are doing about it, what you want the room to decide. Write it before you talk yourself out of it. If it feels slightly uncomfortable to leave in, it is the right paragraph.
- Send the workbook to the CFO's analyst a week early. This quarter and every quarter after, until it is simply how your program works. Add a one-line note asking what looks soft, because a request for a challenge is easier to answer than an invitation to review.
- Put a specific ask on the last page. A decision, a funding request, or an ownership assignment, with the alternatives priced. If you have nothing to ask, ask the room to confirm ownership of your top risk. A meeting with a decision in it stays on the agenda.
Key Takeaways
- Treat the quarterly review as an interpretation problem, not an information problem: the room builds a story from whatever it is given in the first four minutes, and if you do not supply the frame, the loudest anecdote or the last vendor pitch does.
- Structure the report by the three audience questions rather than by workstream: is this working and fast enough (CEO), is the money real and what does it cost (CFO), what could hurt us and who is watching (risk-holder).
- Build the Quarterly AI Report in five fixed blocks every quarter: the three-sentence position, value in CFO arithmetic, momentum in CEO story, risk posture in the risk-holder's frame, and an explicit ask with priced alternatives.
- Report value the way finance reads it: the CLEAR-only subtotal beside the total, full costs including change and verification, avoided spend computed by finance from the gate log, one number volunteered downward, and the assumption that would break the projection.
- Write the "what would worry me most" paragraph in the first person every quarter, because naming your own top worry converts you from advocate to advisor, and an advisor's good news is believed at face value.
- Run the pre-brief circuit without exception: workbook to the CFO's analyst a week early, the risk block to the risk-holder separately, and no sponsor ever surprised by a word in the pack.
- Answer three of the five hardest possible questions inside the document, since a challenge addressed in writing reads as competence while the identical challenge answered live reads as defense.
- Follow the bad-quarter protocol (report early, lead with diagnosis and response, bring a recommendation), because one bad quarter reported well buys more credibility than three good ones, and a report that never contains bad news is uninformative rather than reassuring.
Skill.re