The Readiness Maturity Model: Where Organizations Get Stuck
The consultants presented on a Thursday. Eleven weeks of interviews, a survey sent to four hundred people, a workshop with the executive committee, and a final slide with a needle on a dial: "You are at 2.4." Behind it sat forty pages of dimension scores and a target state described as "AI-enabled at scale," for $180,000. The transformation director asked the only question that mattered, politely: "What do we do differently on Monday?" The room offered strategy refresh, capability building, and executive alignment, which are not things you do on Monday, only things you say about Monday. Two years later that deck sits in a drive nobody opens and the organization is still, by any honest reading, at 2.4. This lesson builds the version that would have been worth the money: a maturity model organized not around what each stage looks like, but around how organizations get stuck at it.
Why Most Maturity Models Are Decorative
Maturity models are the most-purchased and least-used instrument in enterprise transformation. The standard model has five stages with aspirational adjectives (ad hoc, emerging, defined, managed, optimized), a weighted questionnaire, and an output that is a number with a decimal point. It describes destinations beautifully, says nothing about movement, and invites the wrong follow-up question, "how do we get to 3?" What a leader needs is not another description of good, which this program has supplied for a hundred and twenty-three lessons, but a diagnosis: of the many true things that could be improved, which single structural blockage is preventing this organization from converting effort into capability?
That changes the design completely. Organizations do not progress smoothly: they move in a lurch, then sit, sometimes for years, spending real money while the underlying capability does not change. The sitting is the phenomenon, and it deserves a precise name: the stall. A stall is not a shortage of ambition or talent but a structural condition in which the organization's current mode of working cannot produce the next capability, no matter how much more of it is done. Effort applied to a stall is consumed by the stall.
So the stall names the intervention and the stage does not. Being told you are at stage two is not actionable, because stage two organizations do many things correctly. Being told that your method lives in three people's heads, so your fourth use case will cost what your first did, tells you what to build.
A maturity model that names your stage flatters you. A maturity model that names your stall employs you.
So this model is built inside out. Each stage gets a short description, because the description is not the content, then four things that are: the stall (how organizations get stuck there), the mechanism (why, stated as structure rather than as a failing, since blaming motivation produces exhortation and exhortation does not move structure), the unlock (the intervention that changes the structure), and the proof-of-movement observable. Stages are defined by capability, not ambition or spend: an organization with a chatbot deployed to twenty thousand employees, an AI strategy, and a chief AI officer can sit squarely at stage one, and many do.
The Five Stages and Their Stalls
Read each stall first. If one describes your last eighteen months with uncomfortable accuracy, you have located yourself faster than any questionnaire will.
Stage 1: Experimenting
Individuals and departments use AI, mostly unofficially: scattered licenses, tools on departmental cards, perhaps a flagship initiative with executive attention. What there is not is method: no standard way to select a process, no baseline discipline, no gate criteria, no owner accountable for a business number. Work happens; capability does not accumulate.
The stall: the pilot graveyard. A steady stream of promising demonstrations, each impressive in the room, none producing measurable value. The graveyard fills and never empties, because nothing in it is formally buried: pilots do not end, they fade, get paused, or go permanently "under review."
The mechanism. Nothing has a baseline, so nothing can be proven or disproven. That sounds like a measurement inconvenience and is the whole trap: without a pre-change measurement of cycle time, volume, error rate, and cost per unit, a pilot cannot produce a verdict. No verdict means no learning: success teaches nothing because it cannot be attributed, failure teaches nothing because it can be explained away as bad timing or the wrong vendor. So the cycle repeats indefinitely, and it feels like progress because the activity is real. This is where MIT's 95 percent lives: 95 percent of enterprise generative AI pilots delivering no measurable profit-and-loss return, roughly 5 percent of custom tools crossing from pilot to production. It is where S&P Global's 42 percent of companies scrapping most AI initiatives in 2025 comes from, up from 17 percent a year earlier, and where McKinsey's gap sits: 88 percent of organizations using AI regularly, only about 39 percent able to attribute any EBIT (earnings before interest and taxes) impact.
The unlock: one process measured before anything is done to it. Not a strategy, not a governance framework, not a platform decision. One baseline, captured before intervention: typically two to four weeks of an analyst's time, the cheapest intervention in the model and the only one that changes what the organization can know.
The proof of movement: a decision made on evidence rather than enthusiasm, including at least one documented kill. Not "we have baselines," since baselines can be collected and ignored, but a decision with a date, a measured delta, and a named decision-maker. The kill matters more than the scale-up: a program that only decides "yes" has not shown its evidence has teeth.
Stage 2: Piloting With Discipline
Use cases now have charters with pre-committed success and kill criteria, baselines before change, stage gates with real entry criteria, and verified deltas afterwards. A skeptical CFO can be shown a before-and-after that survives inspection.
The stall: the one-hero-project trap and the capacity ceiling. The organization has proved it can do this once and cannot do it repeatedly. Ask for the fourth use case and the honest answer is that the same three people would have to run it, and they are busy. Throughput plateaus at what those people can personally carry, and executive appetite does not raise the ceiling.
The mechanism. Everything depends on specific people and nothing is an asset. The method lives in heads, the evidence in project folders, so each new use case starts near zero. This is the flywheel failure exactly: assets are produced by every engagement and never picked up by the next, so there is no coupling between cycles. Such a program is not compounding, it is repeating, and repeating at fixed unit cost has a ceiling defined by headcount. One delivery lead's departure can drop an organization back to stage one in a fortnight, and nobody notices for a quarter.
The unlock: install the couplings, starting with two. A versioned method library with a single named owner whose job includes folding each engagement's improvements back in, so templates, checklists, and gate specifications accumulate rather than fork. And a staffing rule: people who learn the method in one engagement are deliberately placed in the next, alongside someone who already has it. That is apprenticeship, and both couplings are cheap relative to effect, because they change the unit economics of every future engagement rather than the outcome of one.
The proof of movement: the second and third use cases measurably faster than the first, with the difference attributable to reuse. The attribution clause makes this honest. Faster because the process was simpler does not count. Faster because the baseline already existed and two of the five people had done it before does.
Stage 3: Program
There is now a portfolio rather than a collection: candidates scored on value and readiness, foundations sequenced deliberately, a governance function with real gates, a change workstream, quarterly reporting that shows wins and kills. It looks like serious transformation, and often is.
The stall: the permanence problem. An excellent program that remains a program: dependent on its sponsor, funded annually, re-justified every planning cycle, unable to survive a reorganization or a well-earned promotion. Everyone involved knows this and nobody says it, because saying it sounds like disloyalty.
The mechanism. The four standing capabilities that make it work (discovery, delivery, governance, value tracking) are staffed by people on loan from home functions and funded from project money. Capabilities funded from project money must compete with projects, and they lose, not because leaders are short-sighted but because the comparison format is unfavorable. A project promises a specific return by a specific date. A method library promises that the eleventh use case will be cheaper than the seventh: true, compounding, and unpersuasive in a spreadsheet with a twelve-month horizon. So the compounding investments get trimmed first, quietly, every year. Then the sponsor moves, and the structure is revealed to have rested on one person's political capital.
The unlock: base-budget conversion of the standing capabilities, with named permanent owners. Move the four capabilities into base budget, so removing them requires an explicit decision rather than a passive non-renewal, and give each a permanent named owner rather than a rotating chair. The test: if the current sponsor left tomorrow, would these four still have budget lines next planning cycle? If the answer needs a hopeful tone of voice, the conversion has not happened.
The proof of movement: the program survives a leadership change or a reorganization with its cadences intact. Not in name: with gate reviews still on schedule, the value report still published with kills in it, the method library still versioned, the pipeline still refreshed. This is the one observable you cannot manufacture on demand, which is what makes it trustworthy.
Stage 4: Operating Model
Capabilities are standing and permanently funded, a use-case engine runs continuous discovery through to scaled value, and the flywheel turns: evidence, method, people, foundations, and trust each measurably compound. Evidence infrastructure exists (logging, evaluation, monitoring, an audit trail per AI-touched decision) and literacy is deliberate rather than incidental. Roughly, this is McKinsey's high-performer population, the approximately 6 percent who are about three times more likely than others to have fundamentally redesigned workflows rather than bolting AI onto existing ones.
The stall: the estate problem. The organization has become very good at adding and never learned to subtract. Every quarter adds systems, integrations, retrieval pipelines, dashboards, model versions, vendors, and control obligations. None of it leaves. Maintenance, support, monitoring, and governance overhead grow until they consume the capacity that produced the growth. Throughput falls despite growing headcount, the most demoralizing pattern in the model.
The mechanism. Gates face forward: the governance apparatus scrutinizes what enters the estate and nothing re-examines what is already running. A system that passed its gate in year one is never re-asked whether it still earns its keep, whether its grounding corpus has gone stale, or whether the process it supports still exists. The estate accumulates dead weight steadily, each item carrying recurring cost in licenses, support, incident response, model updates, and compliance evidence. Gartner's forecast that over 40 percent of agentic AI projects will be canceled by the end of 2027 is, in part, this stall arriving on schedule for a cohort that built fast and never built the subtraction reflex.
The unlock: the sunset review and the reallocation reflex. A calendared re-examination of every running AI system against the evidence standard it had to meet to be born: is the delta still real, is the usage still there, is the risk profile unchanged, does anyone still own it? Paired with moving freed capacity to the pipeline immediately, and with the leader behaviors that make stopping safe: crediting whoever proposed the retirement, never treating a sunset as an admission that the original decision was wrong, and reporting retirements in the same quarterly pack as launches.
The proof of movement: throughput rising while the estate's size is deliberately managed, and at least one running system retired on evidence. Both clauses required. Retiring things while throughput falls is decline with better paperwork; rising throughput with an unbounded estate is the stall dressed as success, and it reverses within about two years.
Stage 5: AI-Native
Process redesign with an AI step in it is a normal capability rather than a project. Absorption speed (how quickly a new capability moves from available to embedded in real work with measured effect) is a competitive measure, because the differentiator is no longer access to models, which everyone has, but the rate at which the organization converts access into changed work. Readiness is continuous rather than periodic: there is no assessment season, because the instrumentation runs all the time.
The stall: complacency and drift. Stage five is not an endpoint, because presenting it as one would be the same flattery every other model offers. An organization here is genuinely good, and its risk is the risk of the genuinely good: the disciplines become rituals performed without their reasons. The gate review still happens but its entry criteria have quietly softened, the value report is still published but nobody has audited its attribution logic in two years, and the couplings decay a little each quarter while maturity is assumed rather than measured. Drift is invisible from inside because nothing breaks; it surfaces as an incident, eighteen months later, in the function everyone thought was fine.
The mechanism. Success removes the felt pressure that created the discipline. Every discipline in this program was born from a pain: baselines from unprovable pilots, gates from zombie projects, verification from a fabricated figure that reached a customer, sunset reviews from a bloated estate. When the pain has been absent for two years, the discipline looks like overhead to people who never experienced it, and those people are now the majority. Nobody decides to abandon it; it is optimized away in reasonable increments by capable people in good faith.
The unlock: measured maintenance. At stage five the leader's job changes from building to proving, continuously: a calendar of deliberate tests that can fail. The leakage audit (are the couplings still carrying assets forward, or has each team quietly forked its own templates again?). The retrieval drill (pick five AI-touched decisions at random and reconstruct each from the evidence layer within an hour). The fire drill (simulate an incident, run the escalation, time it). The sunset review on its calendar. The culture test moments (does someone raising an inconvenient measurement still get thanked in public?). All are cheap and all are skippable, which is why they must be scheduled and reported like any other control.
The proof of movement: the drills are run, they occasionally fail, and the failures are fixed. Drills that always pass are theatre: a clean sheet across four quarters does not demonstrate maturity, it demonstrates that the drills are too easy. The observable is the fix log: what did the last retrieval drill fail to reconstruct, and what changed as a result?
Three Honest Properties Generic Models Get Wrong
Stages are not uniform across an organization. An enterprise does not have a maturity stage. Its functions do, and they routinely sit two stages apart: finance at stage four while legal is at stage one. Averaging these manufactures a fictional organization that exists nowhere and hides both the exemplar you should copy and the gap you should close. Worse, for anything cross-functional the enterprise's real stage is closer to its weakest participating function than its strongest, because the immature partner sets the pace of every gate and verification step. This is what the process-maturity grid taught at Level 4: measure by function, report the distribution, never the average. A histogram of six functions across five stages tells a leader where to send capacity; a 2.4 tells them nothing.
Skipping does not work, but acceleration does. An organization cannot buy its way from stage one to stage four. The stalls are structural and each unlock builds the capability the next stage requires: base-budget conversion of a method library is meaningless with no method to put in it, and a sunset review needs an evidence standard built three stages earlier. This is BCG's 10-20-70 arithmetic (10 percent algorithms, 20 percent technology and data, 70 percent people and process) restated as a sequence: the 70 percent cannot be purchased, because each layer rests on the one below. But the transit is far faster with the method than without it, which is the honest version of this program's promise, because the expensive part was never the doing, it was the years spent applying effort to the wrong layer. The promise is not "skip to stage four," it is "do not spend eighteen months discovering that your fourth pilot failed for the same structural reason as your first."
The model is a diagnostic, not a scorecard. Violating this property is fatal to the instrument, and it usually happens within a month. The temptation is obvious: five stages and twelve functions, so you publish a quarterly grid of ratings with a target that everyone reaches stage three by year end. It feels like accountability. It is also the fastest way to destroy the information the model exists to provide, by the mechanism the cross-site variance lesson warned about: ratings depend on self-reported evidence, so once they are a published comparison, the incentive to report favorably overwhelms the incentive to report accurately. You have built a league table of reporting optimism. The useful use is private: a leader locating their own organization honestly to decide where the next increment of capacity goes.
The Artifact: The Self-Location Test
The model comes with a companion instrument, and this is the lesson's deliverable. The Self-Location Test is eight questions a leader answers honestly in about an hour, alone, with no survey and no consultants. Each is keyed to a stage's proof-of-movement observable rather than its description, because descriptions invite generous self-assessment and observables do not. A "yes" must be backed by evidence a skeptic could inspect: a document, a date, a named person, a number.
- Verdicts. In the last twelve months, has any AI initiative been scaled, changed, or stopped on a measured before-and-after comparison, including at least one documented kill?
- Reuse. Were your two most recent use cases measurably faster to value than your first, with the difference attributable to reused assets rather than easier problems?
- Method as asset. Does the method exist as a versioned library with one named owner, updated by the last three engagements, rather than in specific people's heads?
- Permanence. Are discovery, delivery, governance, and value tracking staffed by permanent named owners and funded from base budget rather than project money?
- Survival. Has the program come through a leadership change or reorganization in the last eighteen months with its cadences (gate reviews, value reporting, pipeline refresh) intact?
- Subtraction. In the last twelve months, has a running AI system been retired on evidence through a sunset review, and was the freed capacity reallocated deliberately?
- Throughput versus estate. Is the number of use cases reaching measured value per quarter rising while the estate is deliberately managed rather than simply growing?
- Drills. Are the maintenance drills (leakage audit, retrieval drill, fire drill, sunset review, culture tests) run on a calendar, do they sometimes fail, and is there a log of what got fixed?
How to score it. Question 1 proves exit from stage one; 2 and 3 from stage two; 4 and 5 from stage three; 6 and 7 from stage four; 8 proves you are holding stage five rather than assuming it. Your stage is the highest level for which you answer yes to every question at that level and below. Stop at the first no: it is your stall.
The correction that makes it honest. Leaders systematically over-locate their own organizations by roughly one stage, and the cause is not vanity: they answer from their best example. The flagship function is vivid, recently reviewed, and the thing they personally spend time on, so it becomes the mental sample. The correction is one instruction, and applying it is most of this instrument's value: answer every question from the median function, not the flagship.
Worked Example: One Enterprise Located Honestly
The model applied across the arc of the enterprise storyline this level has followed. All figures are illustrative.
Year zero: stage one. Six functions (finance shared services, customer operations, supply chain, human resources, legal, field service), five unambiguously stage one: local licenses, a few enthusiasts, two vendor pilots that faded rather than concluded, no baselines. One division inside finance shared services was arguably stage two, with charters, gate criteria, and a real baseline on invoice-exception handling (41,000 exceptions a year, 22 minutes average handling time, 11 percent rework). The correct enterprise location was nevertheless stage one, because nothing that division built was shared, findable, or reusable. Capability that cannot travel is not enterprise capability.
The unlock and its proof. That baseline produced a verdict: a document-classification pilot everyone expected to scale was measured at a 1.4 minute per-exception saving against a pre-committed gate threshold of 6 minutes, and was killed at gate two in writing, with the delta table attached. That avoided roughly $310,000 of planned scale-up spend, produced the enterprise's first evidence-based decision, and took about three weeks of analyst time to make possible.
Year one: couplings, to stage three. Both stage-two unlocks were installed: a versioned method library with a named owner, and a staffing rule placing trained people into later engagements. The evidence was arithmetic: the first use case took 34 weeks from charter to verified delta, the second 21, the third 16, with attribution documented rather than assumed (5 weeks saved by reusing an existing baseline, 4 by not rebuilding templates, 4 by two of five team members having done it before, the rest from easier scope and labelled as such). By year end: a portfolio of nine scored candidates, a governance function, a change workstream, quarterly reporting.
Year two: base-budget conversion, to stage four. The permanence stall arrived on cue: the method library owner's time was cut from full-time to 40 percent because a delivery project needed the capacity. The intervention was the operating-model move: 11.5 full-time equivalents (FTE) across the four capabilities converted from project funding to base budget, each with a permanent named owner. The proof arrived unplanned nine months later, when the sponsor left and two functions merged in a reorganization. Gate reviews continued on schedule; the value report was published on time with two kills in it. Cadences intact.
Year three: the honest self-location. Running the test from the median function rather than the flagship: questions 1 through 5, yes with evidence. Question 6, yes, since a retrieval assistant in customer operations was retired after a sunset review found usage at 4 percent of eligible users and a grounding corpus fourteen months stale, freeing roughly 0.8 FTE of support load reallocated to the pipeline. Question 7, no: throughput had risen from 4 to 6 use cases reaching measured value per quarter, but the estate had grown from 19 to 31 systems, so "deliberately managed" was not defensible. Question 8, no: the drills existed on paper and only the fire drill had been run.
The honest location was stage four overall, stalled at the estate problem, with two functions (legal and field service) still at stage two and AI-native markers present in exactly two functions, making any enterprise claim of stage five premature. The flagship reading would have been stage five. The year-end retrospective is the point of the exercise: locating the organization one stage lower than the flagship suggested was the most useful analytical move of the year, because it redirected roughly a third of the next year's discretionary capacity away from an "AI-native enterprise" narrative toward the estate and the two lagging functions. The narrative would have cost a year. The diagnosis cost an hour.
The Failure Story: The Model as a Scorecard
A large organization adopts a five-stage maturity model and does the thing that feels most like leadership: it publishes function-level ratings quarterly, targeting stage three for all fourteen functions by year end. Cycle one: the average is 2.1. Cycle two: 2.4. Cycle three: 2.8. The trend line is beautiful and the underlying practice has not changed, because ratings depend on self-reported evidence. They do not measure maturity; they measure how a function chooses to describe itself in a document its leadership reads.
Two events fix the outcome. First, a function that genuinely regressed, having lost its only two trained practitioners to internal moves, reports an improvement from 2 to 3, justified by a new tool deployment and a training session. Nobody is lying in any way they would recognize; the rubric has room, and the room gets used in the direction the incentive points. Second, and decisively, the one function reporting honestly and low (1.5, down from 2) is asked to present its remediation plan at a leadership meeting. The session is not hostile; it does not need to be. Every other function head watches, and honest reporting ends that afternoon.
Eighteen months later the enterprise believes it is at stage three across the board. Then an incident: a customer-facing summarization tool in a function rated 3.2 has been producing incorrect entitlement figures for five months, with no logging, no named owner, no baseline, and no gate record, because a capable team deployed it under a general platform approval. That is stage-one practice inside a function reporting stage three, and remediation, customer credits, and manual re-review cost roughly $1.2 million, plus the more expensive loss of confidence in every number the program had produced. The autopsy is worth memorizing: the model was accurate, the measurement was social, and a diagnostic used as a target stops being a diagnostic within two reporting cycles.
What to Do Monday Morning
A map is only useful once you put a pin in it. One hour, alone, this week.
- List your functions and find the median. Order every function that touches AI work by capability, as honestly as you can, and identify the middle one. That function, not your flagship, is your organization for the next half hour.
- Run the Self-Location Test from the median. Eight questions, each answered yes only where you can point to a document, a date, a name, or a number.
- Name your stall in one sentence and check it against the mechanism. Write "we are stuck because..." then read the mechanism paragraph for that stage. If it describes your last eighteen months, you have located yourself correctly. If not, you are probably one stage too high; re-run from the median.
- Choose the single unlock for your stage and nothing else. One baseline, or the method library and staffing rule, or base-budget conversion, or the sunset review, or the drill calendar. The most common failure after reading a maturity model is starting on all five layers at once, spreading capacity thin enough that no stall clears.
- Define the observable that will prove you moved, with a date you will check it, written down before you start. Movement is a documented kill, an attributable time saving, a survived reorganization, a retirement with reallocated capacity, or a fixed drill failure. Anything else is spend.
- Keep the ratings private. Share the diagnosis with the function being diagnosed and the people who will act on it. Do not build the grid, publish the league table, or set a stage target. The instrument survives only as long as the answers are safe to give.
With the map in hand and a pin in it, one thing remains: the leader's own first ninety days, where this program ends.
Key Takeaways
- Reject decorative maturity models: a stage number describes a destination, while the stall names the structural blockage and therefore the intervention.
- Diagnose stage one by the pilot graveyard, where nothing has a baseline so nothing produces a verdict (MIT's 95 percent, S&P Global's 42 percent); unlock it with one process measured before intervention, and prove movement with a documented, evidence-based kill.
- Diagnose stage two by the capacity ceiling, where method lives in heads and each use case restarts near zero; unlock it with a versioned method library under a named owner plus the apprenticeship staffing rule, and prove it when use cases two and three are faster through reuse.
- Diagnose stage three by the permanence problem, where the four standing capabilities are on loan and funded from project money; unlock it with base-budget conversion and permanent owners, and prove it when the program survives a leadership change intact.
- Diagnose stage four by the estate problem, where forward-facing gates never re-examine running systems until overhead eats throughput; unlock it with the sunset review and reallocation reflex, and prove it with rising throughput and one system retired on evidence.
- Treat stage five as a condition held rather than an endpoint: success removes the pressure that created the discipline, so the leader's job becomes proving continuously through drills that sometimes fail and get fixed.
- Measure by function and report the distribution rather than an average, since cross-functional work runs at the weakest participating function's stage, and accept that stages cannot be skipped because BCG's 70 percent cannot be purchased.
- Run the Self-Location Test privately and from the median function, never as a published scorecard, because ratings built on self-reported evidence become a league table of optimism within two cycles.
Skill.re