←
AI Readiness & Process Transformation
Capable · M21 · lesson 21 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

The People Readiness Scorecard

15 min

It is the Sunday night before the steering committee, and the assessor's desk holds everything this chapter built: an interview synthesis pack with forty coded themes, a stakeholder heat map with three red cells, a skills matrix with a priced training plan, and a shadow-AI roster naming eleven hidden adopters. Four artifacts, six weeks of careful work, maybe ninety pages of evidence about the people who will make or break the invoice-exception pilot. The committee has given people readiness twelve minutes on Thursday's agenda. And here is the fact the assessor has to make peace with before opening the laptop: the committee will read none of it. Not the synthesis pack, not the heat map, not the matrix, not the roster. Committees do not act on artifacts. They act on scores with evidence behind them. The data readiness assessment arrived last month as a verdict and a number. The process baseline arrived as a cost per transaction. If people readiness arrives as a stack of documents and a feeling, it loses every argument in that room to the dimensions that arrived as numbers. This lesson is about folding four rich artifacts into one defensible score, without losing the chain of evidence that makes the score worth anything.

The Dimension That Refuses to Be a Number

Recall the arithmetic this whole program runs on. BCG's 10-20-70 rule says that in AI transformations, roughly 10 percent of the effort is algorithms, 20 percent is technology and data, and 70 percent is people and process. The 70 percent is where pilots live or die, and MIT's autopsy of the 95 percent of enterprise GenAI pilots that produced no measurable return found the causes squarely inside it: no workflow integration, no learning loop, adoption without transformation. Everyone in your steering committee has heard some version of this by now. Nobody disputes that people matter.

And yet watch what happens in the room. The data readiness report says NOT READY, and it says it with a number: 38 percent of historical invoice-exception records lack the resolution code the model needs. The process baseline pack says the current process costs $19.60 per exception across 26,000 exceptions a year. These dimensions speak committee. They compress into figures, the figures compress into a decision, and the decision gets made in the twelve minutes available. Then people readiness stands up with a synthesis pack and a heat map and says, in effect, "it is complicated." The committee nods respectfully, and the pilot plan absorbs exactly zero of the six weeks of people evidence, because there was nothing shaped like a decision input to absorb.

This is the quiet irony of the 70 percent: the most decisive readiness dimension is the least quantifiable one, and in a committee setting, the dimension that does not arrive as a number loses to the dimensions that do. Not because the committee is lazy or shallow, but because a committee is a machine for comparing things, and you cannot compare a heat map to a cost per transaction. The fix is not to make the committee read ninety pages. The fix is the same discipline you learned at the end of Chapter 2.3, when four data artifacts became the one-page data readiness report: compression with a traceable chain. State the conclusion as a number, and make every number recomputable from an exhibit behind it. What worked for the most quantifiable dimension works for the least quantifiable one. It just takes more craft, and the craft is this lesson.

The named artifact you will build is the People Readiness Scorecard: four subscores, each rated 1 to 5 against written anchor definitions, an evidence line per subscore tracing back to a specific artifact, a weighted roll-up into one number, and a two-sentence narrative that travels with the number wherever it goes. It sits beside process readiness and data readiness in the full assessment, and in Chapter 2.5 all three feed the process-selection scorecard and the one-page readiness report. This is the chapter's closing move: everything the last four lessons produced, folded into the one deliverable a committee can actually use.

Anchors, or You Have Built a Mood Ring

Before the four subscores, one piece of measurement craft that decides whether your scorecard is an instrument or a decoration. A score of 3 means nothing by itself. If you rate change appetite a 3 and a colleague assessing a different process rates theirs a 4, is their team really more ready than yours, or does your colleague just grade warmly? If the sponsor challenges your 2, what exactly are you defending? Without written definitions of what each number looks like in observable terms, a scorecard is a mood ring: it reports the temperature of the person holding it, and everyone in the room quietly knows it.

The fix is the anchor definition: for each subscore, a short written description of what a 1 looks like, what a 3 looks like, and what a 5 looks like, phrased in terms of things you could point to. Not "low morale" but "two or more interviewees independently described the last change initiative as something done to them, and no interviewee volunteered an improvement idea." Anchors do two jobs. They make scores comparable across assessors and across processes, which matters the moment your organization assesses its second candidate process. And they make scores defensible under challenge, because the argument stops being "why did you feel it was a 2?" and becomes "here is the written definition of a 2, and here are the observations that match it." The first argument you lose on politeness. The second you win on evidence.

Write your anchors before you score anything, and keep them in the scorecard document itself, not in your head. The definitions below are a working set for the four subscores. Adapt the wording to your organization, but keep the structure: 1, 3, and 5 defined in observable terms, with 2 and 4 as the in-between judgments they honestly are.

The Four Subscores, Slowly

Each subscore compresses one or two of the chapter's artifacts. That is deliberate: the scorecard is not new research, it is the chapter's research made legible. Every subscore gets a number, an evidence line, and nothing else on the page. The artifacts themselves become the appendix.

Subscore 1: Change Appetite

Source artifacts: the interview synthesis pack and its tension patterns. Change appetite measures whether the people in the affected process have the emotional and practical slack to absorb a new way of working right now. It is read from interview themes: how people talked about the last change, whether anyone volunteered ideas, whether frustration pointed at the process or at each other.

  • A 1 looks like: prior initiative scar tissue dominates the transcripts. Multiple interviewees describe the last rollout in the language of something inflicted on them. Active cynicism ("we know how this goes") appears as a coded theme. No slack: the team is at or over capacity, and any change reads as more work on top of full plates.
  • A 3 looks like: mixed signals with real engagement. Skepticism is present and vocal, but it is skepticism about specifics, not about the idea of improvement. At least one theme shows people frustrated with the current state in a way that wants fixing. Capacity is tight but not crushed.
  • A 5 looks like: teams are asking to pilot. A recent change succeeded and people reference it as evidence that change here can work. Slack exists: someone can absorb three weeks of ragged transition without the queue collapsing.

Now the counterintuitive scoring wisdom, and it matters enough to slow down for: a loudly skeptical team that engages is healthier than a quietly compliant one that does not. The team that argues with you in interviews, that pokes holes in the pilot design, that says "this will fail unless you fix the intake step first," is a team that has already started doing the work of adoption. Argument is engagement wearing armor. The team that nods, agrees with everything, and offers nothing has told you precisely nothing except that they have learned it is safer not to be on the record. When you score change appetite, silence scores lower than argument. Assessors who miss this rate the polite team a 4 and the argumentative team a 2, and get the readiness picture exactly backwards.

Subscore 2: Skills Coverage

Source artifact: the skills matrix, read straight. Skills coverage measures whether the capabilities the pilot needs will exist in the affected teams by the time the pilot needs them. The judgment is three factors multiplied together: gap size (how many critical-path skills are missing), trainability (whether the gaps close with training or require hiring), and calendar (whether they close before the pilot start date, not eventually).

  • A 1 looks like: critical-path skills are missing and no training plan is priced. The matrix has red cells on skills the pilot cannot run without, and the closure plan is a hope, not a line item.
  • A 3 looks like: gaps are identified and a training plan exists on paper, but it is unfunded, unowned, or its timeline lands after the pilot start. The distance to readiness is known; the path is not yet paid for.
  • A 5 looks like: gaps are priced, funded, and closable before pilot start. The matrix's red cells each carry a named intervention, a dollar figure, an owner, and a completion date that beats the pilot calendar.

Notice what this subscore is not: it is not a judgment about whether the team is smart. It is a logistics question about whether a specific set of capabilities arrives before a specific date, and the skills matrix already did the hard work. If your matrix from earlier in this chapter carries a priced plan, this subscore is twenty minutes of translation.

Subscore 3: Champion Strength

Source artifacts: the shadow-AI survey roster and the stakeholder heat map, read together. Champion strength measures whether the pilot has credible internal advocates where it needs them. And here is the trap to avoid: this is not a headcount, it is a coverage measure. Eleven enthusiastic champions who all sit in a team the pilot never touches are worth less than two champions inside the team that carries the new workflow. What you are scoring is champions with influence in the right teams, which is why the roster alone is not enough: you cross-reference it against the heat map to see where the enthusiasm actually sits relative to where the pilot lands.

  • A 1 looks like: no champions at all, or champions only in teams the pilot does not touch. The affected teams face the change with no internal voice that has already made the tools work.
  • A 3 looks like: at least one credible champion inside the main affected team, but coverage gaps remain: an affected team with nobody, or champions who have enthusiasm but no standing with their peers.
  • A 5 looks like: opt-in champions inside every affected team, including at least one convert with high blocking power: someone who could have stopped this and has instead chosen to carry it. A converted potential blocker is the single most persuasive artifact of people readiness that exists, because everyone in the building knows what their opposition would have meant.

Subscore 4: Resistance Exposure

Source artifact: the stakeholder heat map, inverted. Where champion strength scores what pulls the pilot forward, resistance exposure scores what can stop it, discounted by what you can do about it. The judgment is again three factors: how many red cells the heat map shows, how much blocking power sits in them, and whether credible moves exist for each. A red cell with a move planned and an owner assigned is managed exposure. A red cell with a shrug next to it is a live threat.

  • A 1 looks like: high-blocking-power resistance with no credible moves. Someone who can stall or kill the pilot is opposed, and the assessment has no realistic plan that changes that.
  • A 3 looks like: red cells exist and most carry moves, but at least one serious cell is unresolved: the move is drafted but not yet made, or it depends on someone else acting.
  • A 5 looks like: no red cell lacks a move, and the top blocker is already engaged: the conversation has happened or is scheduled, with the right person carrying it. Exposure remains, but none of it is unaddressed.

Score this one with the heat map physically open. Every point of this subscore should be traceable to specific cells, and when the sponsor asks "why only a 3?", your answer is a cell reference, not an impression.

The Roll-Up, the Weights, and the Rule That Saves Careers

Four subscores become one number through a weighted average, and the craft here is not the arithmetic, it is the sequencing: weights are a judgment declared before scoring, never after. Default to equal weights, 25 percent each. Adjust only with a stated reason written into the scorecard: a pilot that demands heavy behavior change from frontline staff might weight change appetite up to 40 percent; a pilot needing rare technical skills on a tight calendar might weight skills coverage up. What you may never do is score first and tune the weights until the roll-up says what someone wants it to say. Weights chosen after scoring are not weights, they are a thumb on the scale, and any analyst on the committee can smell it.

The arithmetic itself is deliberately boring. With equal weights and subscores of 3, 2, 4, and 3: (3 x 0.25) + (2 x 0.25) + (4 x 0.25) + (3 x 0.25) = 0.75 + 0.50 + 1.00 + 0.75 = 3.0. One decimal place, no more. A scorecard reporting 3.04 is claiming a precision the instrument does not have, and false precision invites exactly the credibility attack you built the anchors to survive.

Now the rule that saves careers, and it deserves its own paragraph because it is the emotional center of this lesson: a low score is a finding, not a failure. The scorecard does not exist to grade the organization or to flatter the sponsor. It exists to route effort. A people score of 2.1 with a funded plan to reach 3.5 by pilot start is a healthier pilot than an unexamined assumption of readiness, because the 2.1 tells everyone exactly where the next dollar and the next conversation should go. The assessor who reports a low people score with moves attached is doing the job, not sabotaging the pilot. The assessor who launders a low score into a comfortable one is loading a delayed-action failure into the pilot and signing it. Chapter 2.5 will teach the full craft of presenting findings a sponsor does not want to hear; for now, hold the principle, because the scorecard is where you will need it first.

A low people score with a priced path attached is the job done well. A high score without a chain behind it is a guess wearing a suit.

AI Drafts the Scores. You Own Them.

Scoring four dimensions against written anchors from ninety pages of evidence is exactly the kind of compression work AI accelerates, and exactly the kind where unverified AI output is most dangerous, because a plausible score with plausible-sounding evidence is the easiest thing in the world for a model to produce. Three AI moves earn their place in this workflow, each with the human firmly downstream.

Move one: subscore proposals with forced contrary evidence. Give the model your anchor definitions and the relevant artifact, and ask for a draft: "Propose a change-appetite score for this process using these anchors. Cite the three strongest supporting quotes from the synthesis pack and the two strongest contrary ones." The contrary-evidence clause is the whole trick. A model asked only to justify a score will smooth the record into a tidy story; forced to surface the evidence that cuts against its own proposal, it hands you the tension instead of hiding it. This is the anti-smoothing move, and it belongs in every scoring prompt you write.

Move two: anchor consistency checking. Once your organization scores more than one process, drift creeps in: a 3 on process A quietly means something different from a 3 on process B. Feed the model both scorecards and the anchor definitions and ask it to flag inconsistencies: "These two processes both scored 3 on change appetite. Compare the evidence lines against the anchor definition and tell me whether the same standard was applied." The model is genuinely good at this kind of side-by-side pattern check, and it is a check almost no human ever performs unprompted.

Move three: the argue-both-ways stress test. For any score you feel uncertain about, run: "Argue this score down one point. Now argue it up one point." Read both arguments side by side. What this exposes is which evidence is load-bearing: if the argue-down case leans entirely on one interview quote, your score rests on one person's bad Tuesday, and you should go weigh that quote deliberately. If the argue-up case is thin and stretchy, your score is probably not too low. The two arguments cost three minutes and routinely save a scorecard from its weakest number.

Then the boundary, stated plainly: the human owns every final number and the narrative. AI proposes; you dispose. And verification is mechanical, not aspirational. Every subscore's evidence line must trace to a specific, checkable location in a specific artifact: a transcript theme count ("blame-fear theme, 9 of 14 interviews"), a matrix cell, a roster entry, a heat-map cell. If an evidence line cannot be traced, the subscore is not scored yet. And any score that moved after AI drafting gets its reason logged: "AI proposed 4, human set 3, because two contrary quotes came from the two most senior clerks and outweigh volume." The Verification Log you have kept since Chapter 2.1 absorbs scorecards too. Six months from now, when someone asks why appetite was a 3, the log answers in one line, and that one line is the difference between an instrument and an alibi.

The Worked Example: The Invoice-Exception Scorecard

Time to pay off the whole chapter. All numbers below are hypothetical and illustrative, but the shape is exactly what your own scorecard should look like. Our assessor is scoring people readiness for the invoice-exception pilot at a mid-sized distributor: accounts payable (AP) clerks handling the exceptions that fall out of automated invoice matching. The four artifacts are on the desk. Here is the scorecard they become.

SubscoreScoreEvidence lineWeight
Change appetite3Blame-fear theme in 9 of 14 transcripts (synthesis pack, theme 4); cut against by pride-in-catching-errors theme (6 of 14) and open frustration with queue backlog (11 of 14)25%
Skills coverage2, rising to 4Two critical-path gaps in matrix cells B3 and B5; $28k training plan priced and funded, completion three weeks before pilot start25%
Champion strength4Two opt-in AP clerk champions (roster entries 3 and 7) inside the affected team; AP team-lead convert (roster entry 1, heat map cell C2, formerly amber)25%
Resistance exposure3One unmovable red cell (heat map D1, senior approver, high blocking power); escalation move drafted and on the sponsor's desk awaiting the conversation25%

Walk through the judgment inside each number. Appetite lands at 3, not lower, despite the blame-fear theme being the loudest signal in the pack, because the contrary evidence is real: the pride theme and the queue frustration are the voices of a team that wants the work to be better, and remember, this team argued in interviews rather than going quiet. Skills is the matrix talking: a 2 today because two critical-path gaps are open, a 4 conditional because the $28k plan from the skills-matrix lesson is priced, funded, and beats the calendar. Notice the compounding: a number built two lessons ago is reused here without new work. That is what an assessment that compounds looks like. Champions is a 4 on coverage grounds: champions inside the affected team, plus the team-lead convert, but not a 5 because the downstream approvals team has nobody. Resistance is a 3: every red cell has a move, but the top one is not yet resolved, only escalated.

The roll-up, equal weights declared before scoring: (3 + 2 + 4 + 3) / 4 = 3.0 today. And conditional on two named events, the training plan completing and the sponsor conversation landing: (3 + 4 + 4 + 4) / 4 = 3.75 by pilot start. That conditional structure is the deliverable's real power. A single static number tells the committee where you are; the conditional pair tells them where you are, where you could be, what it costs to get there, and who has to act. It prices the path, not just the position. Committees fund paths.

The two-sentence narrative that travels with the number, verbatim, ready to copy into the deck and the report:

"People readiness for the invoice-exception pilot is 3.0 today, held down by two open skills gaps and one unresolved high-power resistance cell, and held up by strong champion coverage inside AP and a team that engages loudly rather than complying quietly. Conditional on the funded $28k training plan (completes three weeks pre-pilot) and the sponsor's conversation with the senior approver (on his desk now), the score reaches 3.75 by pilot start: both conditions are priced, owned, and dated."

Now the failure story, because there is always a version of this that goes wrong, and it goes wrong in one specific, predictable way. A different assessor, same company in a parallel universe, faces the same Thursday committee. The sponsor has made it gently clear that the pilot "needs to happen this quarter," and the assessor, reading the room, reports people readiness as a reassuring 4. No anchors, no evidence lines, no conditional. The justification on the slide: "the team seems positive." The committee, comparing a clean 4 against the data dimension's messy conditional verdict, waves people readiness through without a single question. The pilot launches directly into the unexamined blame-fear theme: clerks who believe exceptions get people blamed do not feed honest corrections into a tool that records everything, and adoption stalls at 30 percent by week six. In the postmortem, someone pulls the readiness assessment and asks the assessor to walk through how the 4 was derived. It cannot be walked through, because it was never derived. It was chosen. A number without a chain is a guess wearing a suit, and in the postmortem the suit comes off in front of everyone. The pilot's failure was expensive; the assessor's credibility going down with it is the part that follows them to the next role. This is MIT's adoption-without-transformation arriving exactly on schedule, except this time it arrived with a named person's signature on the door it walked through.

Same evidence existed in both universes. The only difference was whether the number had a chain. The scorecard is the 70 percent made legible: BCG's people-and-process share of the transformation, arriving at the committee in the only form a committee can act on, with every link back to the evidence intact. And in Chapter 2.5, this score meets its siblings: people, process, and data scores assemble into the process-selection scorecard and the one-page readiness report, the goldmine hands-on where the whole assessment becomes one decision document. Everything you fold carefully now unfolds cleanly there.

What to Do Monday Morning

You have the four artifacts from this chapter, or you are close. Here is the fold, in order.

  1. Write anchor definitions for all four subscores in your own context. Take the 1-3-5 anchors from this lesson and rewrite them in your organization's observable terms: what does "prior initiative scar tissue" actually sound like in your transcripts? Do this before scoring anything, and save the anchors into the scorecard document itself.
  2. Draft each subscore with its evidence line. Use the AI proposal prompt with forced contrary evidence ("three strongest supporting, two strongest contrary"), then set the final number yourself. Every evidence line must cite a specific artifact location: a theme count, a matrix cell, a roster entry, a heat-map cell. No traceable citation, no score.
  3. Run the argue-both-ways stress test on your shakiest score. "Argue it down one point, then up one point." Identify which evidence turned out to be load-bearing, and go weigh that evidence deliberately before you lock the number.
  4. Compute the roll-up with declared weights. Equal weights unless you write a reason to deviate, and write the reason before you look at what it does to the total. Show the arithmetic in the document. Log any score that moved after AI drafting, with its reason, in your Verification Log.
  5. Write the two-sentence narrative, including the conditional path. Sentence one: the score today and what holds it there, both directions. Sentence two: the conditional score, the named events that unlock it, and confirmation that each is priced, owned, and dated. If you cannot write sentence two, that gap is itself a finding: your assessment has no funded path, and the committee should hear that too.

Key Takeaways

  • Compress the chapter's four people artifacts (synthesis pack, heat map, skills matrix, shadow-AI roster) into the People Readiness Scorecard, because committees act on scores with evidence behind them, not on artifacts, and BCG's 70 percent must arrive as a number or lose every argument to the dimensions that do.
  • Write anchor definitions (what a 1, a 3, and a 5 look like in observable terms) before scoring anything; anchors are what make scores comparable across assessors and defensible under challenge, and a scorecard without them is a mood ring.
  • Score change appetite remembering that silence scores lower than argument: a loudly skeptical team that engages is healthier than a quietly compliant one that does not.
  • Read skills coverage straight from the skills matrix as gap size times trainability times calendar, and champion strength from the roster and heat map as coverage in the right teams, never as a headcount.
  • Invert the heat map for resistance exposure: red cells times blocking power, discounted by credible moves, with a 5 meaning no red cell lacks a move and the top blocker is already engaged.
  • Declare weights before scoring, default to equal, deviate only with a written reason, and show the roll-up arithmetic to one decimal place.
  • Report low scores as findings, not failures: a 2.1 with a funded plan to 3.5 is a healthier pilot than an unexamined assumption of readiness, and the conditional score structure prices the path, not just the position.
  • Trace every evidence line to a specific artifact location, force contrary evidence into every AI-drafted score proposal, log every post-AI score change with its reason in the Verification Log, and own every final number and the narrative yourself.