Auditing Process Maturity Across Functions
Two meetings, same Tuesday, same 2,400-person company, same badge on every lanyard. At 9 a.m., the accounts payable division reviews its invoice-exception dashboard: cycle time is on the wall, the exception rate is trending down for the eleventh straight month, and the monthly improvement retro has an agenda, an issue log, and pastries. At 2 p.m., three floors up, the readiness auditor asks a sales operations manager for the standard operating procedure behind pipeline qualification and gets a long pause, then: "Honestly? Talk to Doug. He's been doing it nineteen years. He retires in March." The enterprise assessment scored this company's process dimension at 2.1 out of 5, and that number, technically accurate, is also the least useful fact in the building. Because this company does not have a process maturity. It has a terrain: one green, irrigated valley and a lot of rock. This lesson is about mapping that terrain honestly, six functions at a time, before the AI portfolio starts writing checks that the ground cannot cash.
The Terrain Under the Rain
Here is the mental model that makes this entire lesson click, and you should carry it into every enterprise engagement you ever run. AI investment lands on an organization the way rain lands on ground: it pools where the terrain is prepared, and it runs off everywhere else. The same tool, the same license, the same training deck will produce measurable value in the function that runs documented, measured, continuously improved workflows, and will produce a demo, a shrug, and a quiet renewal debate in the function that runs on the retiring veteran's memory. The tool did not change between floors. The terrain did.
Enterprises almost never have "a" process maturity, singular. They have a distribution, and the distribution is usually wide. One function, often the one where a Lean or Six Sigma history lives, where some past continuous-improvement program left behind trained people and standard work, operates at a level the rest of the company would not recognize. The next function over has standard operating procedures (SOPs, the step-by-step documents that let a process survive the departure of the person who invented it) that predate the last two reorganizations. A third function has never written anything down at all, because Doug remembers.
The single averaged score from the enterprise assessment, our storyline's 2.1, flattens all of that into one number, and flat numbers invite flat strategies: one rollout, one standard, one date. You saw where the deep-dive data audit took the data dimension in the previous lesson. This lesson does the same descent for process: from one flat score to a function-by-function map with a verdict attached to every territory.
Why does this matter more than almost any other audit in the sweep? Because of the strongest single finding in the AI-value research. McKinsey's State of AI work found that high performers, the roughly 6 percent of organizations getting real earnings impact from AI, are about three times more likely to have fundamentally redesigned workflows, and that workflow redesign is among the strongest drivers of bottom-line impact from AI. Now run that finding to its enterprise corollary, slowly, because everything in Chapter 4.2 depends on it: you can only redesign what is documented and measured. A workflow nobody has written down cannot be redesigned; it can only be excavated first. A process with no baseline cannot prove that redesign helped; it can only generate anecdotes. Which means process maturity is not one readiness dimension among five. It is the rate limiter on the entire AI portfolio. The portfolio can move only as fast as the terrain allows, function by function, and the sweep's job is to find out, honestly, what the terrain allows where.
You cannot redesign what nobody documented, and you cannot prove value in what nobody measured: process maturity is the rate limiter on the entire AI portfolio.
The stakes are the program's familiar arithmetic. MIT found 95 percent of enterprise GenAI pilots deliver no measurable profit-and-loss return, with missing workflow integration as a leading cause of death: pilots landing on sand. BCG's 10-20-70 rule says 70 percent of AI success is people and process, not algorithms. The process sweep is where that 70 percent gets surveyed before anyone pours a foundation.
The Five Markers, and the Artifact That Proves Each One
The sweep measures the same five markers in every function, scored 1 to 5 against written anchors, and here is the discipline that separates a real sweep from a satisfaction survey: the sweep is roughly 80 percent artifact requests and 20 percent interviews. You learned the asking-for-artifacts move in the framework lesson; this is that move systematized. Interviews tell you how a function feels about its processes. Artifacts tell you what is true. A function head can sincerely believe documentation is "pretty solid"; a returned sample of five SOPs, three of them last touched before the 2019 reorganization, cannot flatter itself. Score the evidence, note the evidence, and let the interview explain the evidence rather than replace it.
Marker 1: Inventory existence
The opening request to every function is the same sentence: "Please send the list of your processes, with a named owner for each." That is the Process Inventory Register you built for one division back in Level 2, now used as a maturity floor for six functions at once. A function that can produce the list within a week, with owners who know they are owners, clears the floor. Most functions cannot. You will receive org charts instead of process lists, tool inventories instead of process inventories, and a surprising number of polite requests for an example of what you mean.
Do not treat the failed request as merely a low score. The request itself is an intervention, and a gentle one: the function that spends two days discovering it cannot list its own processes has learned something no slide deck could teach it, at zero political cost, before any score exists. Several of your future remediation sponsors are created in exactly this moment. Log what came back, score against the anchor, and write one evidence line: "Produced 47-process register with owners in 4 days" or "Produced tool list; no process register exists."
Marker 2: Documentation currency
Existence of documentation is the amateur question. Currency is the professional one. The protocol: sample five processes per function (the function's top three by volume plus two you select at random from whatever inventory exists), request the current SOP for each, and check one date: is this document younger than the last reorganization that touched this function? An SOP describing a team structure that no longer exists is not documentation; it is a historical artifact wearing documentation's clothes. The enterprise assessment vignette's "2019 relic" now becomes a per-function rate: three of five sampled SOPs current in one function, zero of five in another.
Then, one more instrument, borrowed directly from your Level 2 verification pass: the do-it test. Take one sampled SOP per function and have someone who does not perform the process attempt to follow it, step by step, while a real performer watches silently. Count the stops: the undocumented judgment calls, the "oh, that screen changed last year," the step that exists only in muscle memory. A stranger who reaches the end is evidence of real documentation. A stranger stranded at step four is evidence of a document that has been decorative for years. One do-it test per function, same protocol everywhere, twenty minutes each. It is the cheapest truth in the entire sweep.
Marker 3: Measurement presence
Request baselines for the function's top three processes: cycle time, volume, error or rework rate. Not dashboards, not key performance indicators (KPIs) in a strategy deck: operational baselines, with a source system and a date. This is the baseline-or-it-didn't-happen rule from Level 1, promoted to a maturity marker, and it carries the single most important sequencing insight the portfolio will ever get. A function that measures can prove AI value in months, because the before-number already exists. A function that does not measure needs a quarter or more of instrumentation before any pilot claim can be proven, no matter how good the pilot is. Write that consequence into the evidence notes explicitly, because Chapter 4.2 will schedule real money against it.
Marker 4: Stability and exception discipline
Is the process the same thing twice? Mature functions run one process with tracked exceptions. Immature functions run a variant per performer, per region, per mood: a workaround economy where the official process is a polite fiction and the real process lives in personal spreadsheets and forwarded emails, with an exception rate that nobody tracks because nobody agreed what an exception is an exception to.
The measurement instrument is your Level 2 walkthrough craft, repurposed as a probe: the two-performer comparison. Ask two performers of the same process, separately, to walk you through it end to end. The divergence between their walkthroughs is the metric. Same steps in the same order with the same system touches: stable. Two recognizable dialects of one process: workable. Two different jobs that happen to share a name: that is not a process, that is a genre, and an AI workflow pointed at a genre will automate one performer's habits and break everyone else's.
Marker 5: Improvement infrastructure
Does the function have any improvement loop at all? Operational reviews that look at process metrics, retrospectives after failures, a kaizen habit (kaizen: the practice of small, continuous, worker-driven improvements), even a humble issue log that someone actually works through. You are not scoring the sophistication of the ritual. You are scoring whether the muscle exists at all.
And here is the claim this lesson wants on the record, explicitly: of the five markers, improvement infrastructure is the best predictor of post-deployment survival. This is the Level 3 insight applied predictively. Functions with improvement muscle absorb an AI workflow the way they absorb any change: they measure it, review it, tune it, and make it better each quarter. Functions without that muscle turn the same AI workflow into a set-and-forget installation that decays quietly from the day the consultants leave: prompts drifting out of date, exceptions accumulating unhandled, verification eroding, until someone finally asks the avoided question two budget cycles later. Marker 3 tells you whether value can be proven. Marker 5 tells you whether value, once created, will still exist in eighteen months. Weight it accordingly.
One Sweep, Not Six Audits: The Comparative Craft
Running the five markers in one function is an audit. Running them across six functions is a sweep, and the difference is not volume. It is three disciplines of comparison, and they are where enterprise-scale craft actually lives.
Calibration across functions
The grid's entire power is comparability, and comparability is manufactured, not assumed. Same written anchors for every score. Same sample size: five SOPs per function, no exceptions, no "finance is complicated so we sampled nine." Same do-it test protocol, same two-performer probe, and when multiple auditors divide the functions, a calibration session where they score one shared artifact set together and argue until their anchors mean the same thing, exactly as the framework lesson prescribed. This is not methodological fussiness. Functions will compare scores, in the hallway before they compare them in the steering committee, and the only thing that makes an unflattering score survivable is the visible fairness of the method that produced it. "Sales scored 1.7" starts a war. "Every function got the same five requests, the same sample sizes, and the same anchors, and here is the evidence column" starts a remediation plan.
The internal benchmark
One of your six functions is not like the others, and you know exactly why: it is the division where the Level 2 and Level 3 story happened, the one with the mature invoice-exception program, the trained process owners, the baselines going back two years, the monthly improvement ritual that survived a leadership change. It will score above 4 across the markers, and the temptation is to treat it as an outlier to be footnoted. Do the opposite. The benchmark division is the sweep's most valuable finding, for two reasons.
First, it is proof, local proof, that the terrain is climbable. Not a vendor case study, not a conference keynote: this is us. Eighteen months ago that division's process maturity looked like everyone else's, and the delta is method, not magic. That is the replication playbook's credibility argument operating at enterprise scale, and it defuses the most corrosive response to a low score, which is "our work just can't be documented like that." It can. Your colleagues did it. Here is the badge photo.
Second, the benchmark is the sweep's methodology donor. Its inventory template, its SOP format, its baseline definitions, its retro agenda, its two or three people who have actually run this transformation: all of it is transferable, and transferring it is dramatically cheaper than inventing it six times. When the remediation program starts, the benchmark division's artifacts become the starter kit, and its practitioners become the internal faculty. An enterprise that uses its outlier this way turns one division's two-year climb into everyone else's six-month head start.
The politics of scoring functions
Understand what you are actually distributing when the grid goes out: function heads will experience maturity scores as report cards, published, with their name at the top, in front of the people they compete with for budget. This is the pre-wire discipline from Level 2's stakeholder chapter at its most necessary. Every function head sees their own row privately, before the grid convenes, with time to correct factual errors and, crucially, with the remediation offer attached. A score presented as bare judgment creates an enemy who will spend the next year disputing your sample. The same score presented as "here is your row, here is the evidence, here is the funded 90-day plan to move marker 3, and here is what becomes possible for your function when it moves" creates a customer.
And frame the lowest verdict honestly, which happens to mean framing it generously. "Greenfield" is not a euphemism for hopeless; it is, in plain arithmetic, the highest-headroom territory on the map. You learned the goldmine arithmetic in Level 2: improvement potential is largest exactly where the baseline is worst. The greenfield function often contains the single largest win in the entire portfolio, once foundations exist. Say so, with the emphasis on both halves of the sentence.
The Grid Rendered: Six Functions, Five Markers
Here is the sweep's named artifact, the deliverable this lesson exists to teach: the Process Maturity Grid. Functions down the side, the five markers across, every cell scored 1 to 5 with an evidence note behind it, and a verdict column that converts the row into a portfolio instruction. Three verdicts, defined before any function is scored: BUILD-READY (the terrain can absorb AI pilots now), REMEDIATE-FIRST (foundation work precedes pilots, with pilot dates contingent on markers moving), and GREENFIELD (start with discovery, not tools: the function needs to find out what its processes are before anyone automates them).
What follows is our 2,400-person storyline's grid, fully rendered. Every number is hypothetical and illustrative: this is what the artifact looks like filled in, not a benchmark to import.
| Function | Inventory | Currency | Measurement | Stability | Improvement | Avg | Verdict |
|---|---|---|---|---|---|---|---|
| A: Finance shared services (benchmark) | 5 | 4 | 4 | 4 | 4 | 4.2 | BUILD-READY |
| B: Logistics and fulfillment | 4 | 3 | 3 | 3 | 2.5 | 3.1 | BUILD-READY |
| C: Corporate finance (FP&A) | 3 | 3 | 1 | 3 | 3 | 2.6 | REMEDIATE-FIRST |
| F: Human resources | 3 | 2 | 2 | 3 | 2 | 2.4 | REMEDIATE-FIRST |
| D: Customer service | 2 | 2 | 2 | 3 | 2 | 2.2 | REMEDIATE-FIRST |
| E: Sales | 2 | 2 | 1 | 1.5 | 2 | 1.7 | GREENFIELD |
The grid means nothing without its evidence notes, so walk the rows the way the steering committee will.
Function A, the benchmark, 4.2. Produced a 47-process inventory with named owners in four days. SOP currency 78 percent across the sample, and the do-it test stranger reached the end of the invoice-exception SOP with two minor stops. Baselines exist for the top five processes, not just three. Improvement ritual runs monthly with a worked issue log. This row is the proof of climb and the source of the starter kit.
Function B, logistics, 3.1: BUILD-READY. Solid inventory, decent currency, real baselines on warehouse cycle times because the warehouse management system produces them whether anyone asks or not. The soft spot is improvement infrastructure at 2.5: reviews happen but feed no loop. Verdict logic: pilots can land here now, and the pilot itself, run with Level 3 discipline, can install the missing loop.
Function C, corporate finance, 2.6: REMEDIATE-FIRST. Here is the row that teaches the sequencing lesson. Respectable documentation, workable stability, and a measurement score of 1: no operational baselines anywhere. Not weak baselines: none. The forecast process has never had a tracked cycle time; rework on the monthly close is folklore, not a number. The consequence gets written in the verdict column in plain language: any AI pilot in this function is unprovable for approximately six months, because the instrumentation to demonstrate impact does not exist and must be built first. The function's first AI investment is not a tool. It is a baseline pack.
Function D, customer service, 2.2: REMEDIATE-FIRST, with one exception lane. Low scores nearly across the board, but the audit's single happiest finding lives here: the ticketing system has been silently recording timestamps, categories, reopen flags, and resolution codes on every one of roughly 210,000 annual tickets for six years. The baselines exist. Nobody has ever queried them. This is the cheap-instrumentation finding, and every sweep has one: sometimes measurement maturity is not absent, it is unclaimed, sitting in a system nobody asked. Two weeks of analyst time converts marker 3 from a 2 to a 4 and moves the function's first pilot date forward by a quarter.
Function E, sales, 1.7: GREENFIELD. The two-performer probe was decisive: two senior reps walked "the" opportunity qualification process and diverged at nearly every step: different entry criteria, different stages, different systems of record (one of which was a personal notebook). "Every rep has their own process" is the honest finding, and the honest verdict is that there is nothing here for an AI workflow to attach to yet. But the row carries the headroom frame, stated because it is true: sales runs the highest-value decisions in the grid on the least process, which makes it the largest single opportunity in the portfolio once discovery has run. First investment: a Level 2-style discovery engagement. Not tools. Discovery.
Function F, human resources, 2.4: REMEDIATE-FIRST, flagged. Mid-pack scores, plus one flag no other row carries, imported from the bias lesson in Level 3: any future AI pilot touching hiring, promotion, or performance decisions inherits the people-decision caution protocol and, under the EU AI Act's calendar, potential high-risk obligations. The grid is where that continuity gets recorded, so the portfolio prices it in from day one instead of discovering it in legal review.
Now the summary arithmetic, the single line that justifies the sweep's three weeks: two of six functions can absorb AI pilots this quarter. Two. Not six. That number reshapes the CEO's "AI everywhere by June" expectation, and it needs to, because the alternative is four functions' worth of the MIT 95 percent, purchased knowingly. Note also what the grid does to the original flat score: the assessment's 2.1 (weighted by headcount, and sales is the biggest function on the grid) has become a terrain running from 1.7 to 4.2, and every point of that range now has an address, an evidence file, and a next step. The sweep is expectation surgery, and like all good surgery it is delivered with a recovery plan: the same meeting that says "two of six" also hands over the remediation calendar showing when functions C, D, and F earn their pilot dates. That pairing, the hard number with the funded path, is what makes the surgery survivable for the messenger.
Reading the Grid for the Portfolio
The grid's verdict column is not commentary. It is an instruction set for Chapter 4.2, and the arithmetic runs like this.
- BUILD-READY functions supply the portfolio's quick wins. Functions A and B are where the first pilots land, because pilots there can integrate with documented workflows, prove impact against existing baselines, and survive on existing improvement muscle. These are the sequencing fuel of the next chapter: the early, provable wins that buy the political oxygen the slower rows will need.
- REMEDIATE-FIRST functions get foundation investment with contingent pilot dates. C, D, and F receive funded remediation (a baseline pack for C, the ticketing-data excavation for D, documentation and currency work for F), and their pilot dates are written as conditions, not promises: "pilot procurement begins when marker 3 reaches 3." The markers moving is the gate. This converts remediation from virtuous overhead into the visible price of admission, which is the only framing under which anyone funds it.
- GREENFIELD functions get discovery as their first investment. Sales gets a process discovery engagement, walkthroughs, inventory, the Level 2 method entire, before any tool conversation is allowed to start. The discipline here is refusing the seductive shortcut: the vendor who promises their tool "works without process change" is describing a tool that will automate the chaos.
Step back and see what the grid has quietly done to the political economy of the AI program. Without it, "where should we do AI first?" is a lobbying contest, won by whichever function head is most senior, most enthusiastic, or loudest, which is how enterprises end up piloting AI in their least-ready function because its leader sat closest to the CEO. With the grid, the question is terrain analysis: scored evidence, published anchors, verdicts defined before scoring began. That is the Level 2 selection-scorecard move, deciding the criteria before the contestants show up, now operating at function scale, and it is the difference between a portfolio and a popularity contest.
The Uniform Rollout: A Failure Story
Here is what skipping the sweep costs, in a composite story assembled from patterns you will recognize; the figures are illustrative. A retail group's board, impatient with pilots, mandates a uniform AI rollout: one procurement-efficiency workflow, deployed across all twelve regions simultaneously. One contract at $1.8 million, one training program, one go-live date. The logic is presented as fairness and speed: no region left behind, no sequencing debates, maximum negotiating leverage with the vendor. Nobody maps the terrain first, because mapping it would have taken three weeks and the board wanted a date.
Four regions, it turns out, had the maturity to absorb it: documented procurement processes, cycle-time baselines, functioning review habits. In those four, the workflow integrates cleanly and produces a real, measurable result: purchase-order cycle time down 31 percent, roughly $2.1 million in annualized working-capital and labor value across the four. In the other eight regions, the identical tool lands on undocumented regional variants, no baselines, and no improvement habit. Buyers work around it within six weeks. Nothing can be proven either way, because there was never a before-number.
Now the cruelest part, and the part your Level 3 training should make you flinch at in advance: the program is judged on the blended average. Four strong wins and eight nulls average out to "modest, below business case," and the program review runs the aggregate: adoption 46 percent, measured savings well under projection, regional complaints high. The four real successes are statistically invisible inside the blend, exactly the segment-blindness failure you learned to catch in Level 3 measurement, now operating at portfolio scale. The program is labeled a disappointment, the renewal is cut to a maintenance license, and two years of organizational patience for AI is spent. The tool performed identically everywhere. The terrain did not, and nobody had looked.
The sweep would have cost three weeks. It would have sequenced the rollout: four regions now, five after a two-quarter remediation wave, three after discovery. Same tool, same total spend, and the first wave's 31 percent becomes the program's headline instead of its buried secret. The absence of the sweep did not just cost value; it cost the program its narrative, and in enterprise AI the narrative is what funds wave two. Terrain ignored is not terrain escaped. It is terrain encountered later, at full price, with an audience.
What to Do Monday Morning
You can begin the sweep this week with two functions and zero budget. The instruments are requests and twenty-minute sessions.
- Send the five-marker work-sample request to two functions. One page, five asks: your process inventory with owners, the current SOPs for your top three processes plus two we name, your baselines (cycle time, volume, error rate) for the top three, the names of two performers of one process we select, and a description of any improvement ritual you run. Give them one week. What comes back, and what does not, is your first evidence column.
- Run the do-it test on one returned SOP. A volunteer who does not do the job attempts to follow the document while a performer watches silently. Count the stops. Write the count into the evidence notes.
- Run the two-performer divergence probe on one process. Separate walkthroughs, same process, note where the paths split. Steps, order, systems touched. Divergence is the stability metric, and it takes an hour total.
- Draft your Process Maturity Grid with the evidence-note column built in from the start. Functions down the side, five markers across, and behind every cell a one-line citation of the artifact that justifies the score. A grid without evidence notes is an opinion in a table costume.
- Write the three verdict definitions, with their portfolio consequences, before any function sees a score. BUILD-READY, REMEDIATE-FIRST, GREENFIELD, each with one sentence on what it means for investment sequencing. Criteria before contestants: it is the move that makes every hard conversation afterward survivable.
Key Takeaways
- Map process maturity as a terrain, not a score: enterprises run a wide distribution across functions, and AI investment pools where the ground is prepared and runs off everywhere else.
- Treat process maturity as the rate limiter on the AI portfolio: McKinsey's high performers are about three times more likely to redesign workflows, and you can only redesign what is documented and measured.
- Score every function on the five markers: inventory existence, documentation currency, measurement presence, stability and exception discipline, and improvement infrastructure, using written anchors and identical sample sizes.
- Run the sweep as 80 percent artifact requests and 20 percent interviews, because work samples cannot flatter themselves, and use the do-it test and the two-performer probe as your cheapest truth instruments.
- Weight improvement infrastructure heavily: it is the best predictor of whether an AI workflow survives after deployment or decays into set-and-forget entropy.
- Use the internal benchmark function as proof and as methodology donor: "this is us, eighteen months ago, plus method" beats any vendor case study, and its templates and people seed the remediation program.
- Pre-wire every function head with their own row, the evidence, and the funded remediation offer before the grid convenes, and frame GREENFIELD honestly as the highest-headroom territory, because the goldmine arithmetic favors bad baselines.
- Read the verdict column as portfolio instructions: BUILD-READY rows supply the quick wins, REMEDIATE-FIRST rows get foundation work with marker-contingent pilot dates, and GREENFIELD rows get discovery before any tool.
Skill.re