←
AI Readiness & Process Transformation
Strategic · M22 · lesson 22 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

The Enterprise AI-Readiness Assessment: Framework and Scoring

15 min

The proposal sits in the middle of the boardroom table in a bound cover the color of confidence, and it costs $1.4 million. Twelve weeks, a maturity assessment across the whole enterprise, a partner from a firm whose name everyone's spouse would recognize. The CEO has read one page of it. What she is actually looking at is the memory of last year: three flagship AI initiatives scrapped, roughly $2.8 million of spend written down, and a board meeting where she had to explain why her company had joined the 42 percent of organizations that S&P Global found scrapping most of their AI initiatives in 2025. Then she looks up at you, the person who spent the last two years turning a division's invoice-exception process from a 12 percent error rate into a documented, verified, governed workflow, and asks the question this level of the program exists to answer: "Could you build this assessment instead?" You can. The big-firm maturity assessment is not magic. It is dimensions, evidence standards, and scoring discipline, three things you have already run at process scale, plus a layer of enterprise craft that this chapter teaches. This lesson builds the whole instrument.

The Assessment a Burned C-Suite Commissions

At Level 2 you learned to score a process: one workflow, one scorecard, one team's evidence, one decision memo at the end. The instinct is to assume an enterprise assessment is the same exercise multiplied by fifty. It is not, and the difference is not size. The difference is three things that never showed up when you were scoring a single invoice queue: politics, evidence logistics, and the audience.

Politics first. A process score threatens nobody senior. An enterprise readiness score is a ranking of executives, whether you print it that way or not. When the operations function scores 2.1 on process documentation and the finance function scores 3.4, the chief operating officer hears an accusation, and the chief financial officer quietly forwards the chart. Every scoring choice you make will be read by someone whose bonus, headcount, or reputation it touches, and several of them will challenge the number. An assessment that cannot survive a motivated executive's challenge is not an assessment; it is a slide.

Evidence logistics second. At process scale you could walk the floor, pull the log yourself, and interview all eight people who touch the work. At enterprise scale the evidence lives in six functions, forty systems, and a few thousand heads. You cannot collect it all personally, which means the assessment becomes a program you run: reports you commission, surveys you field, interviews you delegate and synthesize, work samples you request. The strategist's job shifts from gathering evidence to designing and quality-controlling an evidence supply chain.

Audience third. The commissioning audience is a C-suite that has usually just been burned. That is not a rhetorical flourish; it is the base rate. MIT found 95 percent of enterprise generative AI pilots deliver no measurable profit-and-loss return. S&P Global's 42 percent scrap rate nearly tripled from 17 percent the year before. Gartner found 63 percent of organizations lack AI-ready data practices, and McKinsey's State of AI survey shows the gap your assessment exists to explain: 88 percent of organizations use AI regularly, yet only about 39 percent can attribute any earnings impact to it. The executives in your kickoff meeting have lived inside those numbers. They are not commissioning optimism. They are commissioning an explanation of why the money burned and a map of what to fix, and they will compare your work, openly or silently, to what McKinsey or BCG would have charged seven figures to produce.

Here is the professional secret that makes this level of the program possible: the seven-figure version is built from parts you already own. BCG's 10-20-70 rule is public knowledge by now: 10 percent of AI success is algorithms, 20 percent is technology and data, 70 percent is people and process. Anyone can quote it. What the consultancies actually sell is the ability to run it: to decompose an organization along those lines, attach evidence to every claim, and produce scores that survive the room. Dimensions, evidence standards, scoring discipline. You have run all three on a single process. This lesson scales them, and the result is the named artifact of this chapter: the Enterprise Readiness Framework, consisting of four parts. The dimension grid. The scoring anchors. The evidence standard. The assessment operating plan. Build all four and you are holding the instrument the bound proposal was selling.

The Framework: Four Dimensions, Fifteen Cells

The skeleton of the framework is the readiness lens you have carried since Level 1: data, process, people, governance. At process scale, each of those was a handful of scorecard questions. At enterprise scale, each dimension splits into three or four sub-dimensions, because "how good is your data?" is not a question a 2,400-person organization can answer in one number without the number becoming meaningless. The splits below are the default grid. You may adapt them to your context, and we will discuss the one iron rule about adapting them in a moment.

The 15-cell grid

DimensionSub-dimensionWhat it measuresPrimary evidence type
DataCritical-dataset qualityAccuracy, completeness, and currency of the datasets AI initiatives would actually consumeSystems (profiling reports)
Ownership and stewardshipWhether critical datasets have named, resourced owners with quality accountabilityDocuments + interviews
Access architectureHow long it takes a legitimate project to get usable access to the data it needsSystems (access-grant logs)
AI-usabilityWhether data is structured, labeled, and licensed in ways AI tooling can consumeSystems + work samples
ProcessDocumentation coverageWhat share of core processes have current, accurate, findable documentationDocuments + work samples
Stability and measurementWhether processes run consistently and have baselines anyone could verifySystems + work samples
Redesign capabilityWhether the function has ever redesigned a workflow, not just automated oneInterviews + documents
Improvement infrastructureStanding mechanisms that capture, prioritize, and execute process changeDocuments + interviews
PeopleSkills distributionWhere real AI capability sits, versus where the org chart claims it sitsSurveys + work samples
Adoption evidence and champion coverageWho genuinely uses AI in real work, and whether every function has credible championsSurveys + systems
Change capacityHow much transformation load the function can absorb given everything else on itInterviews + documents
GovernancePolicy and control maturityWhether AI use is governed by written, enforced, current policyDocuments
Decision structuresWho can approve, fund, pause, or kill an AI initiative, and whether that is written downDocuments + interviews
Risk and compliance postureExposure mapped against the regulatory calendar, including the EU AI Act deadlinesDocuments + systems
Evidence and audit infrastructureWhether AI-touched decisions leave a trail an auditor could followSystems

Fifteen cells. Each cell gets a score from 1 to 5, per function, against written anchors. In the enterprise storyline this chapter follows, a 2,400-person organization with six functions, that is 15 cells times 6 functions: 90 scored cells. Ninety defensible, evidence-backed numbers is what a real enterprise assessment produces, and it is why the operating plan later in this lesson is a six-to-eight-week program and not a workshop.

A note on the governance row, because it carries dates. GPAI means general-purpose AI, the foundation models your teams are already using with or without permission. The EU AI Act's obligations for GPAI providers have applied since August 2, 2025. Transparency obligations for AI-generated content arrive December 2, 2026. High-risk obligations for Annex III use cases (hiring, credit, essential services and their kin) land December 2, 2027, and requirements for AI embedded in regulated products follow August 2, 2028. The risk-and-compliance cell is not scored against vibes; it is scored against that calendar. An organization touching European customers or employees that cannot say which of its AI uses will be high-risk in 2027 has an evidence-backed governance problem with a countdown attached.

Freeze the grid before fieldwork

At Level 2 you learned the constitutional rule of scorecards: fix the dimensions and weights before you collect a single score, because criteria negotiated after the numbers exist are not criteria, they are lobbying. At enterprise scale this rule stops being hygiene and becomes survival. The moment a function scores 1.8 on governance, its leadership will discover a passionate methodological interest in whether "evidence and audit infrastructure" really deserves to be its own cell. If the grid is still soft, that conversation ends with the grid bending and the assessment dying, because every other function watched it bend. So: adapt the sub-dimensions to your industry in week zero, socialize them with the executive sponsor, get sign-off in writing, and then freeze the grid. Dimensions negotiated after scores exist are politics, not measurement. The freeze is what makes the eventual scores worth defending.

Anchor discipline: the 1-3-5 that ends arguments

Every cell gets written scoring anchors at 1, 3, and 5, and the anchors must be observable: descriptions of what an assessor would actually find, not adjectives. "Weak, adequate, strong" is not an anchor set; it is an invitation to negotiate. Here are two anchor sets in full, and they are worth reading slowly because the anchors are where the deliverable's credibility lives.

Data: ownership and stewardship.

  • 1: Critical datasets have no named owners; data questions route by folklore ("ask Priya, she's been here longest"). Quality issues are discovered by downstream failures.
  • 3: Owners are named for most critical datasets, but stewardship time is unbudgeted; ownership is a line in a job description that loses to every deadline. Quality issues are logged but age in a backlog.
  • 5: Named, budgeted stewardship exists for every critical dataset, with quality service-level agreements, monitored metrics, and a quarterly review that has demonstrably changed something.

Process: documentation coverage.

  • 1: Core processes live in heads and inboxes. Asked for the standard operating procedure (SOP, the written instruction set for how work is done), the function produces either nothing or a document last touched before the current team joined.
  • 3: Most core processes have documentation, but currency is unmanaged; some maps describe the process as it ran two reorganizations ago, and nobody can say which without checking against the floor.
  • 5: Core processes have current, versioned documentation with named owners and a review cadence; a new hire or an AI project team could work from the documents without oral tradition.

Now watch what the anchors do in the room. An executive can argue with a 2. "That feels harsh" is a complete sentence, and in a soft assessment it works. But an executive cannot productively argue with anchor text that their own team's evidence matches. When the work-sample request came back with one current SOP and two relics, and the level-1 anchor says "a document last touched before the current team joined," the score is not your opinion anymore. It is a description. The executive's energy redirects from fighting the number to fixing the finding, which is the entire purpose of the exercise. Write anchors as descriptions of observable states, and you convert every future scoring argument into a reading exercise.

The Evidence Standard: What Backs Every Cell

This is the center of the lesson, because it is the exact point where the consultancy-grade assessment separates from the workshop-grade one. Enterprise scores are contestable by design; powerful people have incentives to contest them. Therefore every one of the 90 cells carries, alongside its score, a statement of what evidence backs it and where that evidence came from. Five evidence types, each with its own strengths and its own failure mode.

1. Documents. The SOP census, the policy register, the risk log, the org's own steering-committee minutes. Documents prove what the organization has written down, which is not the same as what it does, but absence is powerfully diagnostic: an organization with zero written AI policy has a governance score before you interview anyone.

2. Systems. Log samples, access reports, data-profiling runs, audit-trail extracts. Commissioned, not self-reported: you do not ask the data team how long access grants take, you ask for the ticketing system's export and compute the median yourself. System evidence is the hardest to argue with and the cheapest per unit of credibility, which is why the operating plan commissions it in week one.

3. Surveys. The shadow-AI census (who is using unsanctioned tools for real work), skills self-assessments, adoption questionnaires. Surveys reach everyone and cost little, and they carry the caveat you learned at Level 2: self-report measures perception, not capability. Every survey-backed cell states that caveat in the deliverable, in writing, so nobody can later claim the assessment confused the two.

4. Interviews. Structured, synthesized, and traceable, per the Level 2 interview discipline: every synthesized claim indexes back to who said it, so when a challenged finding meets the question "who said that?", the answer takes thirty seconds, not a weekend of panic. At enterprise scale you will run dozens of these, which is why the synthesis machinery matters as much as the interview guide.

5. Work samples. The quiet star of the evidence standard, and the move to internalize. Instead of asking a function whether its processes are documented, you ask it to produce its best three artifacts: "send us your strongest current process map, and the baseline data behind it, for one process you would nominate for AI." What comes back is the evidence. A function that returns a crisp, current map with real measurements has just demonstrated documentation coverage, measurement discipline, and improvement capability in one artifact. A function that returns apologies and a diagram from 2019 has also answered the question, more honestly than any interview would have. Asking for artifacts measures more honestly than asking questions, because artifacts cannot be aspirational.

The triangulation rule

One rule governs how evidence becomes a score: no cell is scored on a single evidence type. Every cell needs at least two independent types pointing the same direction before the score is defensible. And when the types disagree, the disagreement is not an inconvenience to be averaged away. It is a finding, often the most valuable one in the assessment. The classic case: the skills survey comes back healthy, 70 percent of a function self-rating as confident AI users, and then the work samples come back and none of the three submitted artifacts shows any evidence of competent AI-assisted work. You score the work samples, not the survey, and you report the gap itself: this function's self-perception runs well ahead of its demonstrated capability. That single sentence, backed by both evidence types, is worth more to the C-suite than either number alone, because it predicts exactly how that function's pilots will fail: confidently.

Self-assessment measures confidence, not readiness, and confidence was never the scarce resource.

The Assessment Operating Plan: Six Weeks, Phased

The fourth part of the Enterprise Readiness Framework is the plan that turns the grid, the anchors, and the evidence standard into a program with dates. For a mid-size enterprise, six functions and a few thousand people, the honest envelope is six to eight weeks. Faster than that and you are running a workshop with extra steps; slower and the organization's attention expires before the findings land. Here is the six-week shape.

Weeks 1-2: document census and commissioned system reports. Request the artifact inventory from every function (SOPs, policies, baselines, risk registers) and, in parallel, commission the system reports: access-grant times from the ticketing system, data-profiling runs on the critical datasets, audit-trail samples, tool-usage logs. Commission these first because they have the longest lead times and the highest credibility per dollar. By the end of week two you already hold objective evidence for perhaps a third of the grid.

Weeks 2-4: interviews and surveys in parallel. Forty-plus structured interviews across six functions, plus the shadow-AI census and skills survey fielded to the whole population. This is the Level 2 synthesis machinery at scale, and at this volume it does not run without AI assistance: transcription, structured extraction against the interview guide, cross-interview synthesis. The traceability index is mandatory, not optional. Forty interviews produce hundreds of synthesized claims, and the first executive challenge will arrive within days of the draft circulating. "Who said that?" must be answerable in under a minute, with the claim traced to the specific interviews behind it, or the whole interview evidence base becomes deniable.

Weeks 4-5: work samples and scoring sessions, per function, with the function's leaders in the room. The work-sample requests go out early in week four with a one-week window (the window length is itself diagnostic; a function that needs three weeks to find its best process map has told you something). Then the scoring sessions, one per function, and the craft here is the Level 2 scoring-session discipline unchanged: every score requires cited evidence, disagreements are logged rather than smoothed over, and conditional scores are allowed ("2, rising to 3 when the stewardship budget approved last month actually staffs"). The strategic choice is in the invitation list. You score each function with that function's leaders present, walking the evidence together, before anything is consolidated. This is the pre-wire principle from your stakeholder training, made structural: no executive should first meet their function's scores in a room full of peers. Scoring with functions costs you a week and buys you the difference between an assessment that lands as a diagnosis and one that arrives as an indictment.

Week 6: cross-function calibration. All six functions' scores go side by side for the first time, and the panel re-reads the anchors before looking at a single number. Outliers get challenged: why is this function's 4 on documentation backed by thinner evidence than that function's 3? For high-stakes cells (anything feeding a funding decision, anything scoring a powerful function low), two assessors score independently from the same evidence file and reconcile their differences on the record. That is inter-rater discipline in operator terms: if two competent people reading the same evidence against the same anchors land two points apart, the problem is the anchor or the evidence, and you fix it before publication rather than after the challenge.

AI's role throughout. Everything from your Level 2 and Level 3 toolkit, at scale: extraction from document mountains, interview synthesis, consistency checking of anchor language, and the adversarial score review, where you hand the model a cell's evidence file and the proposed score and instruct it to argue against you. Plus one genuinely new instrument at enterprise scale: the cross-function consistency sweep. Feed the model all 90 cells with their evidence summaries and one instruction: find every pair of cells where the evidence looks similar but the scores differ, and every pair where the scores match but the evidence quality differs, then flag each for human review. It will catch the calibration drift that six weeks of sequential scoring sessions inevitably produces. And per the verification discipline that has run through this entire program: every AI-produced flag, summary, and synthesis gets verified by a human before it touches the deliverable, because a hallucinated finding in a C-suite assessment is a credential-ending event.

Scope honesty. The last component of the operating plan is a paragraph that goes into the deliverable itself, and drafting it early keeps you honest. A six-week assessment samples; it does not census. You profiled twelve critical datasets, not four hundred. You interviewed forty people, not two thousand four hundred. You reviewed each function's three best work samples, which means you measured their ceiling, not their average. Write the sampling caveat into the deliverable in plain language, exactly as Level 3 taught you to write measurement caveats into pilot results. It costs you nothing, because a C-suite that just watched consultants get quietly walked back from overclaimed findings will trust the assessor who states limits unprompted. Precision about what you do not know is the cheapest credibility you will ever buy.

Ninety Cells at Norvik: The Worked Example

Meet the organization this level's storyline follows. Norvik Group (illustrative, like every number in this section) is a 2,400-person business-to-business services and distribution company: six functions covering sales, operations, finance, supply chain, customer service, and corporate services. Last year Norvik had its scrap year: four AI initiatives launched, three dead, roughly $2.8 million written off, one CEO with a board mandate that begins "before we spend another euro on AI..." You have been commissioned to run the readiness assessment. The alternative on the table was the bound proposal: $1.4 million and twelve weeks from a strategy house.

You run the six-week plan. Ninety cells, scored against frozen anchors, every score carrying its evidence line. The illustrative topline, on the 1-to-5 scale, weighted across functions: data 2.4, process 2.1, people 2.9, governance 1.8.

The governance 1.8 is the headline, and it is the Level 1 shadow-AI pattern at enterprise scale: the document census found zero written AI policy anywhere in the organization, while the shadow-AI census found 71 percent of surveyed staff using unsanctioned AI tools for real work at least weekly. Ungoverned use at scale, with the EU AI Act transparency deadline of December 2, 2026 only months away when the assessment ran, and the Annex III high-risk deadline of December 2, 2027 roughly seventeen months out, with Norvik's customer-service function running AI-assisted decisions that will need classification review before then. That cell scored 1, on document evidence (absence) triangulated with survey evidence (prevalence), and nobody in the room argued, because there was nothing to argue with.

Two cell-level vignettes show the instrument working.

The work-sample request that scored itself. For the process-documentation cell, six functions each received the same request: your best current process map with its baseline data, one week. Finance returned an excellent artifact: current, versioned, with real cycle-time measurements attached. Two functions returned maps dated 2019, one still showing a department that no longer exists. Three functions returned apologies. Against the anchor set you read earlier in this lesson, the scores wrote themselves: the anchor text described what arrived. When the operations lead pushed back in the scoring session ("we document more than this suggests"), the response was not a defense of the number; it was the request email, the deadline, and what came back. He withdrew the objection and asked a better question: what would it take to look like finance? That question is the assessment succeeding.

The 14-day access wall. For the access-architecture cell, the commissioned system report pulled every data-access request ticket from the previous two quarters: 214 tickets, median time from request to usable access, 14 days. Not an opinion, not an anecdote: the organization's own ticketing system describing itself. The cell scored 2, and the finding underneath it explained a piece of the scrap year that nobody had put into words: last year's pilots each burned their first three to four weeks waiting for data access, spending their credibility and calendar before writing a line of configuration. The chief information officer did not enjoy the number, but the evidence type made the conversation short. You cannot pre-wire your way out of a median; you can only fix it.

The economics. Total assessment cost, illustratively: about 50 strategist-days of your time, two commissioned system reports the internal teams produced inside their existing roles, survey tooling the company already licensed, and AI assistance at rounding-error cost. Call it $85,000 to $120,000 fully loaded, against the $1.4 million quote, for a deliverable whose evidence discipline the reader of this program can defend cell by cell. That arithmetic, run once, is the economic foundation of the readiness-strategist credential: the assessment is not cheaper because it is worse; it is cheaper because the method was never the expensive part. The expensive part was believing only four firms in the world could do it.

The failure story: the perception assessment

Now the same engagement, run the way it usually gets run. A different firm, same size, same scrap-year scar tissue, decides an external assessment is too slow and too expensive. The transformation office books a two-day executive workshop instead. Sticky notes, breakout groups, a facilitator with excellent energy. Each function self-scores the four dimensions: "we're probably a 4 on data; our warehouse project just finished." By the second afternoon there is a maturity radar chart, and it is beautiful: consensual, balanced, hovering encouragingly between 3 and 4 on every axis. It goes into the strategy deck. Funding decisions cite it.

Eight months later the first two funded pilots hit the real data foundation. The first spends five weeks waiting for access grants the radar chart said were a solved problem, then discovers the vendor master file has no owner, three competing versions, and a 40 percent duplicate rate that nobody was accountable for finding. The second pilot, built on the assumption of "4 on data," dies quietly in month four when its outputs cannot be reconciled against a source of truth, because there isn't one. Neither pilot's post-mortem blames the radar chart. Post-mortems never do. The chart just stops appearing in decks, deleted from the template without a decision anyone can date, the way fictional evidence always exits: silently, after the money is spent. The workshop measured the confidence in the room, and the confidence was real. But the S&P Global 42 percent is substantially a census of confident organizations. Nobody scraps an initiative they expected to fail; the scrap year is what expected success looks like when it was never underwritten by evidence. The two-day workshop did not skip the assessment. It skipped the evidence, and delivered the assessment's cost as a delayed, compounding invoice.

One frame to carry out of the contrast: the perception assessment and the Enterprise Readiness Framework produce the same deliverable format. Dimensions, scores, a chart. The entire difference, the whole seven-figure difference, is what stands behind each number: in one case a sticky note, in the other an evidence line any executive can pull and test. When you compete with the big firms, and at this level you are, evidence discipline is the product. Everything else is formatting.

The bridge. The topline told Norvik's C-suite where the organization stands: data 2.4, process 2.1, people 2.9, governance 1.8. It did not yet tell them what to do about any of it, and a topline without a remediation path is a diagnosis without a treatment plan. The next four lessons go dimension by dimension into the enterprise-scale audits behind those four numbers, in the order the evidence usually demands: data first. The 14-day access wall and the unowned vendor master are not anecdotes; they are the next lesson's subject.

What to Do Monday Morning

You do not need a commissioned engagement to start building the instrument. Every step below is executable this week, inside your current role.

  1. Draft your 15-cell grid. Start from this lesson's table, adapt sub-dimensions to your industry, and write the adaptation down with a date. This is the freeze candidate; treat it as versioned from day one.
  2. Write full 1-3-5 anchors for the three cells you know best. Observable states only: what an assessor would find, never adjectives. Test each anchor by asking whether two colleagues reading it against the same evidence would land on the same score.
  3. Write the evidence standard for every cell. One line each: which of the five types (documents, systems, surveys, interviews, work samples) backs this cell, and which second type triangulates it. Any cell you can only back one way is a design flaw to fix now.
  4. Commission one system report: access-grant times. Ask whoever runs the ticketing system for two quarters of data-access requests and compute the median yourself. It is the cheapest objective evidence in the building, and whatever the number is, you have started assessing.
  5. Design one work-sample request. Draft the exact email asking a function for its best current process map and the baseline behind it, with a one-week window. You will learn more from what a pilot version of this returns than from a month of interviews.
  6. Calendar the six weeks. Even as a draft for a future engagement: document census and commissioned reports in weeks 1-2, interviews and surveys weeks 2-4, work samples and per-function scoring sessions weeks 4-5, and the cross-function calibration session booked, with names, in week 6. A plan with a calibration session on the calendar is a program; everything else is an intention.

Key Takeaways

  • Recognize what changes between scoring a process and scoring an organization: not size but politics, evidence logistics, and a burned C-suite audience that will compare your work to a seven-figure consultancy deliverable.
  • Build the Enterprise Readiness Framework as four parts: the 15-cell dimension grid (data, process, people, governance, each split into enterprise sub-dimensions), the 1-3-5 scoring anchors, the per-cell evidence standard, and the assessment operating plan.
  • Freeze the grid before fieldwork begins; dimensions negotiated after scores exist are politics, not measurement, and one visible bend destroys every other cell's authority.
  • Write anchors as observable states, because an executive can argue with a 2 but cannot productively argue with anchor text that their own team's evidence matches.
  • Back every cell with at least two of the five evidence types (documents, systems, surveys, interviews, work samples), commission system evidence rather than accepting self-report, and treat disagreement between evidence types as a finding: self-perception versus demonstrated capability is its own insight.
  • Run the assessment as a six-to-eight-week phased program: commissioned reports first, interviews and surveys in parallel with mandatory traceability, per-function scoring sessions with leaders in the room, and a cross-function calibration week with independent double-scoring where stakes are high.
  • Use AI for extraction, synthesis, adversarial score review, and the cross-function consistency sweep, and verify every AI-produced claim before it reaches the deliverable, because one hallucinated finding ends the credential.
  • State scope honestly in the deliverable itself: a six-week assessment samples rather than censuses, and the sampling caveat, written unprompted, buys more C-suite trust than any confident overclaim, because the 42 percent scrap statistic is substantially a census of confident organizations.