The Readiness Flywheel: Making Readiness Continuous
Two programs, side by side, both four years old, both respectable. The first has completed nine AI use cases, and the ninth took eleven weeks from idea to measured value. The second has completed fourteen, and the fourteenth took thirty weeks, almost exactly what the fourth took. Same industry, similar budgets, comparable talent, and one difference nobody in either organization can name without help: the first built assets while it built projects, the second built only projects. The word usually reached for here is "flywheel," used so loosely it has stopped meaning anything. This lesson makes it mechanical and inspectable: five assets, five couplings that feed them, five leaks that drain them, five observables that prove the thing is turning. By the end you will be able to say, with evidence rather than hope, whether your next use case will be cheaper than your last.
What a Flywheel Actually Is, and Why Most Programs Do Not Have One
Strip the metaphor down to a testable definition, because the metaphor is where the thinking usually stops. A flywheel exists when the output of each cycle becomes an input that makes the next cycle cheaper, faster, or better, without additional investment. Every clause is load-bearing. Output, not activity: something must persist after the project closes. Becomes an input: it must be picked up by the next cycle, which requires a mechanism, not a hope. Cheaper, faster, or better: the improvement must show up in a number somebody already tracks. Without additional investment: if the second cycle is only faster because you spent more on it, that is not compounding, that is purchasing.
Apply that definition honestly to the average enterprise AI program and it fails on the second clause almost every time. The assets get produced. They simply never get picked up. Use case seven produces a beautiful baseline of the claims-intake process, and it lives in a project folder that use case eleven's team never learns exists. Use case three's delivery lead builds an excellent assessment-to-charter template, and it lives on her laptop, so use case eight's lead builds his own. Use case five remediates a customer master file scoped precisely to its own needs, so use case nine remediates the same file again from a slightly different angle. Use case two trains eleven people in the redesign method, all eleven go back to day jobs where the method is never used, and by the time use case ten needs those skills the skills have decayed.
Every cycle produced real value and every cycle produced a genuine, reusable asset. None of them connected. This is the single most expensive pattern in enterprise AI, and it is invisible because each individual project looks fine: nobody fails an audit for it, nobody gets a bad review for it. The only symptom is that the organization keeps paying full price for capability it already bought.
A flywheel is not built by producing more assets. Every active program already produces them. It is built by installing the couplings that make one cycle's output the next cycle's input.
This reframing is the whole lesson. The five assets exist by default in any program genuinely doing work. Only the couplings are optional. That is why programs run for four years with competent people and produce no compounding: nobody decided against the couplings, they were simply nobody's job. A coupling is never glamorous. It is a filing requirement, a named owner, a staffing rule, a prioritization key, a reporting habit. Five small, boring, specific mechanisms are the difference between a program that gets cheaper every year and one that pays the same tuition forever.
The research backdrop makes the stakes concrete. MIT's 2025 GenAI Divide study found roughly 95 percent of enterprise generative AI pilots delivering no measurable profit-and-loss (P&L) return, with only about 5 percent of custom tools crossing from pilot into production. McKinsey's State of AI found 88 percent of organizations using AI regularly while only about 39 percent could attribute any EBIT (earnings before interest and taxes) impact. The usual diagnosis is that the tools disappoint. The diagnosis this lesson offers is narrower: most organizations never reached their fifth attempt carrying the advantages their first four should have bought. They ran a fifth pilot with a first pilot's cost structure, and a first pilot's cost structure fails 95 percent of the time.
The Five Compounding Assets
Five assets compound in an AI transformation program. For each, three things matter and only three: the coupling (the mechanism that carries it forward), the leakage (the way it escapes), and the measure (the observable proving it compounds). If you cannot name all three for an asset, that asset is not compounding in your organization, whatever the strategy deck says.
Asset 1: Evidence
Every engagement produces evidence: a process inventory, a baseline pack (pre-change measurements of cycle time, volume, error rate, cost per unit), a delta table showing what actually moved, and failure data showing what did not work and why. This is the most obviously reusable material your program creates, and the most reliably wasted.
The coupling: evidence is filed in a form the next assessment can query. Not archived, not "available on request," but structured so a new assessment can ask "do we already have an instrumented baseline for invoice exception handling?" and get an answer in an afternoon. Practically: the process inventory, baseline pack, and delta table are enterprise assets with a home, an index, and a custodian, not project artifacts that die with the charter. The AI-native lesson made this argument at process level; here it is the same argument with a bigger denominator.
The leakage: evidence that lives in a project folder, a closed workspace, or in the transformer's head. The head is the sneakiest version, because it feels like retention. It is not. It is a single point of failure with a notice period.
The measure: the share of new assessments that reuse an existing baseline rather than building one from scratch. This number starts at zero by definition and should climb every year. If it is flat after three years, your evidence is leaking.
Asset 2: Method
Every engagement improves the method: assessment templates, redesign canvases, prompt patterns that survived contact with real work, verification checklists, stage-gate specifications with their entry criteria. People learn things in delivery that nobody could have written in advance, and those learnings are perishable.
The coupling: a versioned method library with a single named owner whose actual job includes folding each engagement's improvements back in. Versioned means a number that changes and a history showing what changed and why. Named owner means one person, not a community of practice, because a community of practice owns nothing. The mechanism is a short close-out step in every engagement: what did we change about how we work, and has it reached the library?
The leakage: every team keeping its own copy. This is the most common failure in mature-looking programs, because local copies feel efficient and produce divergence with no learning. Six teams with six template sets do not produce six times the improvement. They produce zero, because no version accumulates the others' lessons.
The measure: the method's version history, and specifically whether the last three engagements contributed to it. Two questions, answerable in five minutes, that tell you more about a program's health than most dashboards.
Asset 3: People
Every use case trains people. Business participants learn what AI can and cannot do in their own process, analysts learn the redesign method, a few discover they are unusually good at this. Capability is created in every engagement whether you intend it or not.
The coupling: two mechanisms together. First, the deliberate progression from the literacy ladder (champion to applied user to specialist to practitioner, with named criteria at each step) so learning has somewhere to go. Second, and more powerful, a staffing rule: people who learn the method in one engagement are deliberately placed in the next. Every engagement includes at least one person being developed alongside someone who already has the method. That is an apprenticeship model, and it is how methods have always propagated in professions that take craft seriously.
The leakage: the participant who returns entirely to their day job. Capability decays fast when unused, and the decay is silent. Eighteen months later the organization genuinely believes it has forty trained people and functionally has six.
The measure: redesign density (how many people can independently lead a process redesign with an AI step in it) tracked over time, alongside the internal-promotion share of practitioner roles. If every practitioner is an external hire, your people asset is not compounding, it is being purchased.
Asset 4: Foundations
Every use case builds foundation: data remediation on the sources it touches, access fast-paths through security review, grounding infrastructure (retrieval pipelines that let a model answer from your content rather than its training), and control scaffolding such as logging, evaluation harnesses, and monitoring. All of it is reusable and most of it is built as if it were not.
The coupling: remediation is prioritized by beneficiary count rather than by requester. This is the data-program rule from Level 4 and the single most powerful coupling in the entire flywheel, so it deserves the plainest possible statement. Fix a data domain because the project that asked loudest needs it and you have bought a one-project asset at one-project cost. Fix it because eleven identified downstream candidates need it and you have bought an eleven-project asset at roughly the same cost. Same money, same team, an order of magnitude difference in what the organization owns afterwards. The mechanism is a prioritization key: every candidate domain in the remediation backlog carries a count of the use cases, current and pipeline, that depend on it, and the count drives the order.
The leakage: per-project point fixes that nobody generalizes. Watch for the sentence "we only fix what this use case needs," which sounds like admirable scope discipline and is, at portfolio level, a decision to pay for the same repair repeatedly.
The measure: the ratio of new use cases consuming already-certified data to those remediating their own. In most programs this ratio is where the largest single time saving is hiding.
Asset 5: Trust and Permission
The fifth asset is the one leaders treat as a mood rather than a balance sheet. Every honest delivery, every documented kill, every accurate quarterly report earns credibility, and credibility has a direct operational effect: it lowers the cost of the next approval. Approval cost is real cost: weeks of calendar time, three rounds of committee, a business case rewritten twice, a delivery team idling while paperwork circulates.
The coupling: the deliberate practice of reporting kills and bad quarters, drawn from the culture and enterprise-reporting disciplines. This is what converts trust from a mood into an asset, because a program that only ever reports wins provides no information: its reports cannot be distinguished from marketing, so each new claim gets re-examined from scratch. A program that has published two kills and one honest bad quarter has demonstrated that its numbers survive contact with bad news, and its next number is believed on sight.
The leakage: a single inflated number. Trust is the only one of the five assets that can be drained to zero by one event, and this asymmetry is precisely why the measurement disciplines in this program are non-negotiable rather than aspirational. One optimistic figure that later collapses resets the balance and reinstates the challenge cycle for years.
The measure: approval cycle time for a new charter, and the depth of challenge it receives. Both fall visibly in a trusted program, not because scrutiny is abandoned but because it becomes targeted. The committee stops re-litigating your measurement method and starts asking about the one thing genuinely uncertain in this case.
The Artifact: The Flywheel Map
Here is the lesson's deliverable, and it fits on one page. The Flywheel Map lists the five compounding assets and names, for each, the coupling that feeds it, the leakage that drains it, and the observable that proves it is turning. Filled in honestly, it is the most direct diagnostic of long-run program health available, because it does not ask what you have built. It asks what carries forward.
| Asset | Coupling (the mechanism) | Leakage (where it escapes) | Observable (proof it turns) |
|---|---|---|---|
| Evidence | Baselines, inventories, and delta tables filed in a queryable enterprise store with a custodian | Evidence in project folders or in the transformer's head | Share of new assessments reusing an existing baseline |
| Method | Versioned method library with a named owner and a close-out contribution step | Every team keeping its own template copy | Version history, and whether the last three engagements contributed |
| People | Literacy ladder progression plus a staffing rule placing each engagement's learners in the next | Participants returning entirely to day jobs, capability decaying silently | Redesign density over time; internal-promotion share of practitioner roles |
| Foundations | Remediation prioritized by beneficiary count, not by requester | Per-project point fixes nobody generalizes | Ratio of use cases consuming certified data to those remediating their own |
| Trust | Deliberate reporting of kills and bad quarters alongside wins | One inflated number, which resets the balance | Charter approval cycle time and depth of challenge |
Two rules for using it. First, fill in the leakage column before the coupling column: leakage is observable today, couplings are often aspirational, and a map filled in from the aspiration side describes a program that does not exist. Second, treat any blank observable as a red flag, not an administrative gap. An asset with no measure is an asset nobody is accountable for, and in eighteen months it will be an asset nobody has.
The Leakage Audit
The Flywheel Map's second column drives a standing exercise: for each asset, where is value escaping right now? Run it as a working session, twice a year, with the people who actually deliver rather than the people who report on delivery. Five leaks account for the overwhelming majority of what you will find.
- Evidence in project folders. Ask a delivery lead to produce the baseline pack for a use case completed two years ago, and time the retrieval. More than a day, or needing one specific person, means the evidence is not filed, it is stored.
- Method in personal copies. Ask three delivery leads to send the assessment template they used last. If the three are different versions and the differences never reached the library, the method is diverging rather than accumulating.
- People un-redeployed. Take the participant list from your last three engagements and ask where each person is now. The proportion in unrelated day jobs with no follow-on placement is your decay rate.
- Remediation un-generalized. Ask the data team how the last five remediation items were prioritized. If the answer names projects rather than beneficiary counts, the most powerful coupling in the flywheel is not installed.
- Trust spent on an optimistic number. Read your last four quarterly reports and count the bad news. Zero is not a good score: it means either an implausible run of luck or a reporting habit quietly optimizing for comfort.
Leakage is normal and continuous, which is why the audit is periodic rather than a one-time fix. Couplings decay under organizational pressure exactly as machinery decays under load: people leave, priorities shift, a busy quarter makes a close-out step feel optional. A program that audits twice a year and finds two leaks each time is healthier than one that ran a perfect audit once, in year one, and never looked again.
The Startup Problem, and What Breaks a Turning Flywheel
Now the honest part, because a lesson selling compounding without naming its cost would be exactly the optimism this program spends its time dismantling.
The first two turns feel like pure cost
A flywheel is heavy and starts slowly. On the first turn every coupling is a tax: the filing requirement adds days to a project with no successor yet; the method library takes real time to populate for a method used once; the staffing rule puts an inexperienced person into a critical engagement; beneficiary-count prioritization means fixing data your loudest sponsor did not ask for; publishing a kill costs political capital before the trust balance has any deposits in it. On the second turn the benefits are small and hard to see. On the third and fourth the curve bends. Programs are cancelled between the first and second turn almost universally, by a sponsor who was never told this shape was expected.
So the leader's obligation is to say it, in advance, in the funding conversation. The multi-year investment logic from earlier in this level exists for exactly this: a three-year shape in which year one buys foundations and evidence, year two shows unit-cost improvement, and year three shows portfolio value that could not have been bought directly. A sponsor who agreed to that shape reads a slow second turn as the plan. A sponsor who did not reads it as failure, and is not wrong to.
Two practical moves follow. Instrument the couplings early, from use case one, even though the numbers will be embarrassing: baseline reuse of zero percent in year one is not a bad result, it is the origin point of a trend line, and you cannot show a trend you did not start measuring. And report the leading indicators before the lagging ones arrive. Time-to-value and cost-per-use-case lag, moving in year two or three; baseline reuse, method contributions, certified-data consumption, and internal practitioner share lead, moving in months. Reporting the leading ones makes compounding visible while it is still too small to see in outcomes, which is precisely the window in which programs get killed.
What breaks a turning flywheel
Three things reliably stop a flywheel that was working, and all three are survivable if you can see them coming.
- A reorganization that disperses the people without moving the method. Restructures are usually not yours to prevent. What you control is whether the method library, the evidence store, and the custodian roles survive the redraw with named owners in the new structure. People scatter; if the method is institutional they carry it with them and the flywheel slows rather than stops. If the method lived in those people, the reorganization is an extinction event.
- An inflated claim that spends the trust asset. Usually not a lie: usually a number reported at its most flattering framing under pressure, later walked back. One of these can undo three years of approval-cycle improvement, because trust is asymmetric, earned in increments and spent in lumps.
- A period of neglect where the couplings quietly stop being enforced. The close-out step gets skipped in a busy quarter. The method owner leaves and the role sits vacant. The staffing rule gets waived because this engagement is too important for a learner. None of these is a decision anybody would defend if asked, which is exactly why they happen without anybody being asked.
The last one carries the section's most important structural point: a flywheel does not coast. Physical flywheels store energy and decelerate against friction; organizational ones do the same, and the friction is turnover, reprioritization, and busyness. Maintenance is not new investment. It is enforcement of couplings that already exist, which is cheap, unglamorous, and almost never on anybody's objectives. Put it on somebody's.
Worked Example: Three Years of a Turning Flywheel
The following numbers are illustrative, built to show the shape of compounding rather than to promise it. They follow the enterprise storyline used through this level: a mid-sized organization, a central AI function, a portfolio that grew from one use case to roughly thirty over three years.
Evidence. Year one baseline reuse was zero percent by definition: there was nothing to reuse. By year three, 64 percent of new assessments drew on an existing instrumented baseline rather than building one, and a typical assessment fell from six weeks to two and a half. The mechanism was not clever: a filing requirement in the engagement close-out and a named custodian for the evidence store.
Method. The library reached version 4.2, with contributions from 11 of the last 12 engagements. The highest-value component was the assessment-to-charter template, credited with about a week saved per use case, mostly by ending the argument about what a charter should contain. A week per use case across roughly a dozen engagements a year is a meaningful number produced by a document nobody would put in a board deck.
People. Redesign density went from 2 people to 9. Seven of the nine were promoted internally rather than hired. The named mechanism was the staffing rule: every engagement includes one person being developed, no exceptions, including the high-profile ones. The exception request comes up every time, and refusing it is most of the work.
Foundations. In year one, 8 percent of new use cases consumed already-certified data; by year three, 61 percent did. This was the largest single time saving in the portfolio, and it traces to one policy change at the end of year one: the remediation backlog was reordered by beneficiary count rather than by requester. The domains that got fixed were not the ones with the loudest sponsors, they were the ones sitting under the most pipeline.
Trust. Charter approval went from five weeks with heavy, wide-ranging challenge to nine days with targeted challenge. Asked what bought that, the program lead gave a specific and slightly uncomfortable answer: two publicly documented kills and one honest bad quarter. The committee had watched the program report against itself and concluded it did not need to.
The composite observable. Time-to-value, measured from charter approval to first verified delta in a business number, went from 31 weeks on the first use case to 9 weeks on the most recent. The discipline is in how that was reported: attributed across the five couplings (foundations and evidence carrying most of it, method and people the rest, trust compressing the approval segment at the front), not to any tool or platform. A leader who says "we got three times faster because of our new platform" has told a story that collapses at the next vendor change. A leader who says "we got three times faster because 61 percent of use cases now start on certified data and 64 percent start on an existing baseline" has told a story that survives.
And one leak, found in the year-two audit. Three engagements' method improvements had never reached the library, because the owner role sat vacant for a quarter after a promotion and nobody noticed. The gap was quantified (roughly two months of rework across following engagements), the role was backfilled, and the close-out step moved from a project checklist into the gate criteria so a missing contribution blocked closure. This is what a healthy audit finding looks like: unglamorous, specific, fixed in a week, and quantified so the next person tempted to leave the role vacant knows the price.
The Failure Story: Fourteen Projects and Zero Assets
A large insurer ran an AI program for four years with a competent central team, real executive support, and no scandal. Fourteen use cases were completed, several genuine successes: a claims triage assistant that measurably cut handling time, a document extraction workflow that removed an underwriting backlog, a policy-comparison tool agents actually liked. Nobody would have called it a failing program.
Then an internal review asked one question: why did the fourteenth use case take as long as the fourth? Not longer, which would suggest decay. The same, which suggests something stranger: four years of learning had bought nothing measurable in delivery efficiency.
The review found five broken couplings, one per asset, and each of them was locally reasonable.
- Evidence: baselines lived in project SharePoint sites that nobody indexed. Each site was tidy. There was no way to search across them, so nobody did.
- Method: each delivery lead kept a personal template set. All four leads were good; their templates had diverged into four dialects and none had accumulated the others' improvements.
- People: participants returned to their functions with no follow-on placement. The program had trained an estimated 130 people and could not name 10 who had used the method twice.
- Foundations: remediation was scoped per project by explicit design, and the team was proud of the discipline: "we only fix what this use case needs." Three use cases had remediated overlapping slices of the same policyholder data.
- Trust: reporting had always emphasized wins, so every new approval was negotiated from scratch. No kill had ever been published; two initiatives were quietly wound down and described as "transitioned."
The team was not lazy and not unskilled. Every individual decision had a defensible rationale, and the five together produced an organization that had built fourteen projects and zero assets. The review's most uncomfortable sentence, the one that circulated afterwards, was that the company had paid for the same learning fourteen times.
Notice how the diagnosis fits the wider evidence. S&P Global found 42 percent of companies scrapping most of their AI initiatives in 2025, up from 17 percent the year before. Gartner found 63 percent of organizations lacking or unsure of AI-ready data practices. McKinsey found the roughly 6 percent of high performers about three times more likely to fundamentally redesign workflows, and BCG's 10-20-70 rule (10 percent algorithms, 20 percent technology and data, 70 percent people and process) puts the weight of the work exactly where couplings live. None of these findings is about model quality. All are about whether an organization converts one engagement's work into the next engagement's starting position.
What to Do Monday Morning
Five moves, in order, none of them requiring new budget.
- Draw the Flywheel Map for your own program. One page, five rows. For each asset, name the coupling that exists today in mechanism terms (who does what, at what moment). If you cannot name a mechanism, write "none," because "we encourage sharing" is none.
- Run the leakage audit on your most recent engagement. Where is its baseline pack filed, and can a stranger find it? Which template improvements reached the library? Where are its participants now? What did its data remediation fix, and who else can use that fix? What did its reporting say about what did not work? Write the five answers down; that page is your starting diagnosis.
- Make remediation prioritization beneficiary-count-driven if it is not already. This is one change to a backlog ordering rule and it is the highest-leverage single action in this lesson. Add a column to the remediation backlog counting current and pipeline use cases per domain, and sort by it. Expect an argument with whoever has the loudest project; have the count ready.
- Install the staffing rule. Every engagement includes one person being developed, decided before the engagement is staffed, not after. Write it into the resourcing template so waiving it requires a named approver rather than a shrug.
- Start reporting two compounding indicators. Baseline reuse rate and internal practitioner share, on the standing report, starting at whatever they are today even if zero. You are not reporting achievement, you are opening a trend line, and the trend line protects the program through its slow second turn.
Do these five and you have not launched a readiness initiative, which is the point. Readiness stops being a project the moment the couplings exist, because every piece of work the organization was going to do anyway now leaves something behind that makes the next piece cheaper. That is what "making readiness continuous" means: not a permanent programme office, but mechanisms that turn ordinary delivery into accumulating capability.
Which leaves one thing unfinished, and it is the hardest thing in the transformation. A flywheel that turns well changes what the work is, which changes what the jobs are, which puts a leader in a room with a person whose role is about to be different and who deserves the truth about it. That conversation closes this chapter, and it cannot be delegated to a mechanism.
Key Takeaways
- Define the flywheel precisely: it exists only when each cycle's output becomes an input that makes the next cycle cheaper, faster, or better without additional investment, and most programs fail that test at the "becomes an input" clause.
- Recognise that the five assets (evidence, method, people, foundations, trust) are produced by default in any active program, so the leader's real job is installing couplings, not building assets.
- Treat each coupling as a small, specific, unglamorous mechanism: a filing requirement, a named owner, a staffing rule, a prioritization key, a reporting habit.
- Prioritise data remediation by beneficiary count rather than by requester, the single most powerful coupling available, because it converts a one-project cost into an N-project asset at the same price.
- Protect trust as the one asset a single inflated number can drain to zero, and build it deliberately by publishing kills and bad quarters alongside wins.
- Run the leakage audit twice a year against the five standard leaks (evidence in folders, method in personal copies, people un-redeployed, remediation un-generalized, trust spent on optimism), because leakage is continuous.
- Name the startup problem out loud when funding is agreed, instrument the couplings from use case one, and report leading indicators such as baseline reuse and internal practitioner share before the lagging outcome numbers arrive.
- Remember that a flywheel decelerates rather than coasts: reorganisation, an inflated claim, and quiet non-enforcement stop a turning one, and maintenance means enforcing existing couplings, not funding new initiatives.
Skill.re