←
AI Readiness & Process Transformation
Capable · M19 · lesson 19 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Scoring a Real Process End to End

15 min

There are nine sheets of paper taped to the conference room wall. Eight of them are anchor definitions, one per criterion, printed large enough to read from any chair: what a 1 looks like, what a 3 looks like, what a 5 looks like, in observable terms. The ninth sheet says three words: EVIDENCE OR NOTHING. Around the table sit the accounts payable manager who owns the invoice-exception process, the AP team lead who lives inside it, the data steward who knows where its records are buried, and, in the corner with a notebook and strict instructions to observe, the finance director sponsoring the whole assessment. You are at the whiteboard, facilitating. On the table are four artifacts, each with a version number on its cover page: a verified process map, a standard operating procedure, a baseline pack, and a people readiness scorecard. In the next two hours, this room will produce eight numbers, and those eight numbers will decide whether a six-figure pilot gets nominated, remediated, or refused. Every instrument you have built in this level was built for this session. This lesson is that session, cell by cell, including the fight.

The Room Where the Number Gets Made

Last lesson you built the Process-Selection Scorecard: eight criteria, four on the value side, four on the feasibility side, each with written 1-to-5 anchors, each with a weight, and a decision threshold at 3.2. An instrument on paper is a promise. Its first public scoring is where the promise is kept or broken, because everyone who will ever be asked to trust a scorecard number is really trusting the process that produced it. If the first scoring is rigorous, the scorecard becomes an institution. If the first scoring is a guess dressed in decimals, the scorecard becomes a prop, and props get ignored the first time they say something inconvenient.

So the first rule of the session is about who is in the room. Scoring is done with the process's residents, not to them. The process owner is there because she is accountable for the thing being judged. The team lead is there because he touches the exceptions daily and will catch any claim that flatters the process. The data steward is there because data readiness is where pilots die, and Gartner's numbers say so twice: 63 percent of organizations lack AI-ready data practices, and through 2026, 60 percent of AI projects without AI-ready data will be abandoned. The sponsor is there to watch, not to score, because a sponsor's raised eyebrow moves numbers, and numbers moved by eyebrows are not evidence. And you, the assessor, hold the pen and the anchors, and you score nothing alone.

The ground rules go up before the first criterion, and you read them out loud even though they are on the wall:

  • Anchors decide, not adjectives. Every score must match the printed anchor text. "Feels like a 4" is not a sentence anyone gets to finish.
  • Every cell gets a citation. A score enters the record only with a named piece of evidence: an artifact, a version number, a page. No citation, no score, no exceptions, including for the facilitator.
  • Disagreements get logged, not averaged. If two people land on different scores, both positions and both citations go into the record before anything is settled. Surfaced disagreement is signal. Suppressed disagreement is a landmine with a delay fuse.
  • The AI drafts and challenges; it never proposes a score. More on this later, because the reason is subtle and the rule is absolute.

Why this much ceremony for eight numbers? Because of where those numbers go. MIT's GenAI Divide research found that 95 percent of enterprise GenAI pilots deliver no measurable profit-and-loss return, and the autopsy traced the failures to organizational causes: no workflow integration, no learning loop, adoption without transformation. The scoring room is the last cheap moment to catch those causes. Every score you inflate here becomes a budget line later; every gap you hand-wave here becomes a month-five surprise. The session is two hours. The pilot it green-lights is nine months. The exchange rate on honesty in this room is roughly one uncomfortable minute now per week of zombie pilot avoided later.

The other way to do this, and how it ends

Hold the counterexample in view the whole session, because it is the default you are displacing. At a different organization, one we will keep anonymous and illustrative, the sponsor of an AI program scored his organization's top candidate process alone, at his desk, in twenty minutes. His reasoning was efficiency itself: "I know these processes. I don't need a committee to tell me what I already know." Value criteria: all 5s. Feasibility criteria: all 4s. Citations: none, because he was the citation. The sheet looked decisive, leadership funded the pilot on it, and for four months the project ran on the fumes of that confidence. In month five it hit a data gap, a core field that existed in two systems with two incompatible meanings, and the integration work to fix it was quoted at more than the pilot's remaining budget. The pilot died. The postmortem contained one sentence that should be framed in every scoring room: the organization had the evidence, it just was not in the room. Their own data steward knew about the field conflict. Nobody asked her, because asking her would have taken a meeting. Scoring alone is not efficiency. It is evidence suppression with a calendar benefit.

The Value Side: Four Cells, Four Citations

The candidate on the table is the process this entire level has been assessing: accounts payable invoice-exception handling. You have walked its floor, mapped it, verified the map, drafted and verified its SOP (standard operating procedure, the written instruction set for how the work is done), baselined its costs, profiled its data, and scored its people. Now the artifacts come out of their folders, and the walk through the criteria begins on the value side.

Criterion 1: run-rate value. The anchor for a 4 reads: "annual fully loaded cost between $250,000 and $750,000, from a documented baseline." The process owner opens the baseline pack, version 1.0, to its summary page: 1,150 exceptions per month at roughly $31 each, an annual run rate of about $428,000, with a 12 percent rework rate inside it. That is a citation, not a recollection: the pack documents its sources and its counting rules, which is exactly why you built it. Score: 4. Citation: Baseline Pack v1.0, summary page. Thirty seconds, zero debate, and notice why: the debate already happened weeks ago, during the baselining, where it was cheap. Evidence gathered early is disagreement prevented later.

Criterion 2: pain intensity. Here comes the session's first correction, and it comes from the process owner, who says, honestly and warmly, "this one feels like a 5. My team is drowning." You point at the wall. The anchor for a 5 requires documented pain: attrition data, escalation logs, customer complaints tied to the process. The anchor for a 4 requires "consistent pain themes across a majority of interviewed staff, plus at least one measurable symptom." Feelings are not evidence, and the room holds the line, gently: what can we cite? The team lead pulls up the interview synthesis from the assessment: 11 of 14 interviewees independently described fear of blame as the reason exceptions sit untouched in queues, exact quotes attached. The baseline pack supplies the measurable symptom: a p90 resolution time of 9 days, meaning the worst 10 percent of exceptions take longer than 9 days to clear while the median case moves in a fraction of that. Two citations, one score. Score: 4, not 5. The process owner nods, and something important just happened: the instrument corrected its most senior friend in the room, politely, using paper instead of rank. That is the moment the scorecard earned its first unit of credibility.

Criterion 3: volume and growth. The anchor for a 3: "volume sufficient for measurable pilot results within one quarter, but flat or slow-growing." The baseline pack again: 1,150 exceptions per month, stable across the trailing twelve months, no seasonal spike, no growth trend. Plenty of volume to measure a pilot against, no compounding urgency. Score: 3. Citation: Baseline Pack v1.0, volume table. A lesser room would round this up "because AP matters." This room does not, because a 3 that is really a 3 makes the 4s believable.

Criterion 4: strategic visibility. The anchor for a 4: "named by a senior executive in a formal forum within the last two quarters." The sponsor, from the corner, offers the citation herself: the chief financial officer flagged invoice-exception cost by name in the Q2 budget review, and the minutes say so. Score: 4. Citation: Q2 budget review minutes. Note what the sponsor did and did not do: she supplied evidence when she had it. She did not supply opinion. That is the observer role working.

Value side complete: 4, 4, 3, 4. Four cells, four citations, one rejected adjective. Twenty-five minutes on the clock.

The Feasibility Side, and the Fight in Cell Six

Criterion 5: process readiness. The anchor for a 4: "process mapped and verified against reality; SOP exists and matches the map; baseline measured." You lay the artifacts on the table like exhibits, because that is what they are. Process map v1.1, carrying its verification block: the version that survived the invented-step hunt, checked against real invoices, with the corrections dated and initialed. SOP v1.0, with its verification note recording who walked the floor and what changed between draft and verified. Baseline Pack v1.0. Score: 4, and the citation is three document titles with version numbers. Feel what the version numbers are doing here. In a scoring session, "we have a process map" is a claim; "Process Map v1.1, verified on these invoices, superseding the retired draft" is a fact with a paper trail. Every hour of document discipline this level demanded of you was a deposit, and this cell is the withdrawal.

Cell six: the fight

Criterion 6: data readiness. And now the session earns its existence, because the data steward and the team lead disagree, out loud, by two full points.

The data steward argues for a 2, and she has a citation: the data-gap findings from the assessment show that the real reason for each exception lives in a free-text comment field for 61 percent of cases, in eleven spelling variants of "PO mismatch," while the structured reason-code field is decorative. Any AI triage pilot needs reliable reason codes; today they do not exist. The anchor for a 2 reads: "a critical field is missing or unreliable, with no committed fix." She reads it aloud and rests her case.

The team lead argues for a 4, and he has a citation too: the data readiness report that went to leadership got two remediations approved and funded. The dropdown enforcement fix, configuration only, zero dollars, owned by him, due in two weeks. And retention of rejected invoices, switched on already, accumulating the history a pilot will need. The gaps are real, he says, but they are dying gaps, and scoring the corpse is unfair to the patient.

Here is where a weak facilitator averages: two plus four, call it three, moving on. Do not do that, because silent averaging destroys exactly the information the committee needs most. Instead you run the disagreement protocol, and it has four steps.

  1. Log both positions with their citations. The scoring record gets both lines verbatim: "Steward: 2, per Gap Findings, free-text reason field, 61 percent of cases. Team lead: 4, per Data Readiness Report, two funded remediations, dated and owned." The disagreement is now an asset that travels with the score, instead of a tension that evaporates into a number.
  2. Re-read the anchor text aloud. Not paraphrased. Read. The anchor for a 3 says: "critical fields have known gaps, and every blocker gap has a funded, dated, owned fix that lands before pilot start." The room goes quiet, because the anchor was written for precisely this situation, back when nobody was defending a position.
  3. Settle on the anchor, not between the people. The gaps are real (so not a 4: a 4 requires fields reliable today). The fixes are funded, dated, and owned (so not a 2: a 2 requires no committed fix). The anchor text says this exact configuration is a 3.
  4. Attach the condition in writing. The score enters the record as: "Data readiness: 3, conditional on the two funded remediations landing by pilot start. If either slips, this score reverts to 2 and the nomination returns to committee." Both the steward and the team lead sign off, because each of them can see their evidence alive inside the settlement.

This is conditional scoring, and it is the honest middle path between optimism and paralysis. The condition is not a footnote; it travels with the score forever after, into the readiness report, into the committee deck, into the pilot's stage gates. Remember the Gartner figure: 60 percent of AI projects without AI-ready data will be abandoned through 2026. The conditional score is the mechanism that keeps your organization out of that statistic, because it makes the data fix a load-bearing part of the decision instead of a hopeful aside. The sponsor's pilot in our failure story died in month five on a data gap. This process's equivalent gap just got a tripwire attached to it, in writing, with two witnesses.

Cells seven and eight: the system feeds itself

Criterion 7: people readiness. No debate needed, because this dimension arrives pre-scored. The People Readiness Scorecard from earlier in this level rolled change appetite, skills coverage, champion strength, and resistance exposure into a 3.0, held down by two open skills gaps and one unresolved resistance cell, with a documented conditional path to 3.75 by pilot start, both conditions priced, owned, and dated. Score: 3. Citation: People Readiness Scorecard v1.0, including its conditions. Pause on what just happened structurally: a whole scorecard became one cell of a bigger scorecard. Dimensions feeding dimensions. The assessment is not a stack of documents; it is a system, and this is the cell where you can see the plumbing.

Criterion 8: solution maturity. The anchor for a 4: "proven solution category with multiple established vendors and documented deployments in comparable organizations." Invoice-exception triage qualifies: it is exactly the kind of back-office, document-heavy, high-volume work where MIT found the clearest returns, and the same research found that externally partnered solutions succeed roughly twice as often as internal builds, which matters because this would be a buy, not a build. Score: 4. But the record gets a countervailing note, and you insist on it: category maturity is claimed by vendors, and vendor claims are marketing until verified. Gartner's agent-washing finding, thousands of self-described agentic vendors of which only about 130 were judged real, is the standing reason every shortlisted vendor still gets the full due-diligence treatment from Level 1, benchmark verification and reference calls included. A 4 on solution maturity is permission to shortlist, not permission to trust.

The Arithmetic in the Open

Eight scores exist. Now the math happens on the whiteboard, in the open, in front of everyone, because arithmetic done in private is just another eyebrow. The weights come from the scorecard you built last lesson: 40 percent of the total on the value side, 60 percent on feasibility, because feasibility is where the 95 percent went to die. All figures here are, as always, illustrative numbers for a hypothetical organization, but the method is exact.

CriterionScoreWeightWeightedCitation
Run-rate value415%0.60Baseline Pack v1.0 ($428k run rate)
Pain intensity410%0.40Interview synthesis (11 of 14); p90 9 days
Volume and growth310%0.30Baseline Pack v1.0 (1,150/month, flat)
Strategic visibility45%0.20Q2 budget review minutes
Process readiness415%0.60Map v1.1; SOP v1.0; Baseline Pack v1.0
Data readiness3*20%0.60Gap findings; funded remediations (conditional)
People readiness3*15%0.45People Readiness Scorecard v1.0 (conditional 3.75 path)
Solution maturity410%0.40Vendor landscape; MIT partnered-solution finding; due diligence pending
Weighted total100%3.55Threshold: 3.2

0.60 plus 0.40 plus 0.30 plus 0.20 plus 0.60 plus 0.60 plus 0.45 plus 0.40 equals 3.55, against a nomination threshold of 3.2. The verdict, spoken aloud and written down: NOMINATE, with two conditions attached. Condition one: the two funded data remediations land by pilot start, or the data score reverts and the nomination returns. Condition two: the people readiness path to 3.75 executes as priced and dated. The asterisks in the table are not decoration; they are the two places where this nomination is a promise rather than a fact, and everyone downstream deserves to see them.

Everything on that whiteboard now becomes the session's named artifact: the Scoring Record. One document, versioned like everything else you build, containing five things: the eight scores with their weights and the arithmetic; the citation for every cell, artifact names and versions included; the disagreement log, both positions verbatim with their evidence; the conditions, worded as tripwires with owners and dates; and the verdict with the threshold it was measured against. The Scoring Record is what makes the number auditable. Six months from now, when someone asks "why did we pick AP exceptions?", the answer is not a memory. It is a document that shows its work, and next lesson it becomes the one-page readiness report that goes in front of the committee.

A score without a citation is an opinion wearing a number. Surfaced disagreement is signal; suppressed disagreement is a landmine.

The two candidates the scorecard refused

The same week, the same room, the same wall of anchors scored two other candidate processes, and what happened to them matters as much as the nomination, because a scorecard's credibility comes from what it refuses. The contract-renewal process arrived as everyone's favorite: high value side, roughly 4.3, real money, real executive attention. Its feasibility side scored 1.9: no verified map, a critical data source living in one retiring analyst's spreadsheet, no funded fixes. The threshold rule from last lesson is unambiguous: a feasibility side below 2.5 is remediate-first, no matter how loud the value siren sings. The room held, the siren was resisted, and the process owner left not with a rejection but with a prioritized remediation list: the exact, costed steps that would make his process nominable next cycle. Done right, a refusal is a gift with a roadmap attached. The third candidate, a monthly report-formatting task, scored a gaudy 4.5 on feasibility, clean data, simple process, eager team, and 1.8 on value: barely $30,000 a year at stake, illustratively, and no strategic visibility. Verdict: not now, politely, in writing. Automating it would produce a demo, a press release for the intranet, and membership in MIT's 95 percent, where adoption without transformation lives. One nomination, one remediate-first, one not-now. That distribution is what a working instrument looks like. A scorecard that only ever says yes is not an instrument; it is a rubber stamp with arithmetic.

What the AI Did in the Room, and What It Was Never Allowed to Do

AI was working throughout this session, and the shape of its involvement is a lesson in itself, because this level's discipline applies here with full force: the tool drafts and challenges, people decide, and everything AI-touched gets verified.

Job one: live consistency checking. During the session, the assessor's laptop ran a standing prompt: here are the eight anchor definitions verbatim, and here is each score with its cited evidence as we log it; flag any score that sits more than one point from what the anchor text implies for that evidence. It flagged twice. Once on pain intensity, where the initial "5" had no citation yet, the same catch the room made on its own, and the agreement between machine check and human check is itself useful signal. Once on solution maturity, where it noted the evidence cited was vendor-published and suggested the record say so, which is how the due-diligence caveat ended up in writing. Both flags were suggestions to humans, verified by humans against the actual anchor text before anything changed.

Job two: drafting the Scoring Record. The AI turned the session's running notes into the structured record, tables, disagreement log, and conditions, in minutes instead of an evening. Then the record went through the verification habit like any AI draft: every citation checked against the actual artifact, every quote in the disagreement log confirmed by the person quoted. The data steward found one error: the draft rendered her position as "the data is not ready," a summary, where the record requires her actual cited claim, the 61 percent free-text figure. Summaries drift; citations do not. The correction took one minute and is exactly why humans sign the record.

Job three: the post-session adversarial pass. After the room emptied, you ran the challenge prompt: here is the full Scoring Record with all cited evidence; argue each score down one point using only the evidence cited, no outside assumptions. This is the cheapest red team you will ever hire. Any score that survives the pass is defensible in front of a committee, because the committee's skeptic cannot construct an attack the pass did not already try. Any score that folds gets re-examined before the report goes out, not after the CFO finds the soft spot in a live meeting. In this session's pass, seven scores held. The eighth, strategic visibility, wobbled: the pass pointed out that one mention in one budget review is a thin citation for a 4 if the anchor says "named in a formal forum," singular being technically satisfied but weakly. The score stood after review, but the readiness report will carry the citation explicitly so the committee can weigh it themselves. That is the pass working: not changing numbers, but hardening the record around them.

And the rule that governed all three jobs: the AI never proposes a score during the session. Not once, not as a starting point, not "just to calibrate." The reason is anchoring, the same bias that makes the first number spoken in a negotiation gravitationally warp every number after it. Put an AI-suggested "3.8" on the shared screen and the facilitation dies on the spot: the residents stop reasoning from evidence to anchor text and start reasoning toward or against the machine's number, and the sponsor in the corner starts wondering why the humans disagree with the software. The BCG 10-20-70 arithmetic applies even inside a scoring session: the algorithms are the small part, and the 70 percent, people confronting evidence together and owning the resulting number, is the part the pilot's survival actually depends on. The tool drafts records and manufactures challenges. People decide. Accountability stays human, in this room and every room after it.

What to Do Monday Morning

You have the instrument from last lesson and, from this level, the artifacts to feed it. Here is how to run your own first scoring session.

  1. Schedule the session with the process's residents, not around them. Two hours, one candidate process (two more later in the week if you have them). Invitees: process owner, a hands-on team lead, the data steward, you facilitating, sponsor as silent observer. If the data steward "can't make it," reschedule; an empty steward chair is the failure story waiting to repeat.
  2. Print the anchors and stack the artifacts. All eight anchor definitions on the wall, large type. Every artifact you will cite on the table with version numbers visible: map, SOP, baseline pack, gap findings, people scorecard. If an artifact is missing, that is a finding, not an inconvenience; score the cell with what exists and log the gap.
  3. Read the ground rules aloud, then walk all eight criteria in order. Every score gets a citation into the record as you go. When someone offers a feeling ("this is definitely a 5"), point at the anchor and ask for the artifact. Expect to reject at least one uncited score; that rejection is the session's credibility being minted.
  4. Log every disagreement with both positions and both citations, then settle on the anchor text. Re-read the anchor aloud, verbatim. If the honest settlement needs a condition, write the condition as a tripwire: what must land, by when, owned by whom, and what the score reverts to if it slips.
  5. Do the arithmetic on the whiteboard, in the open. Weights, multiplication, total, threshold, verdict. Then have the AI draft the Scoring Record from your notes, and verify every citation and quote against the source artifacts before anyone signs.
  6. Run the adversarial pass the same day. "Argue each score down one point using only the cited evidence." Re-examine anything that folds before the record leaves your hands. Then file the Scoring Record as v1.0, conditions and disagreement log intact, ready to become next lesson's one-page readiness report.

Key Takeaways

  • Treat a scoring session as an evidence proceeding, not form-filling: every cell gets a score, a citation to a named artifact with a version number, and a logged disagreement wherever humans differed.
  • Score with the process's residents in the room (owner, team lead, data steward, silent sponsor); the failure story's org had the evidence and lost the pilot because the evidence was not invited to the meeting.
  • Reject uncited scores out loud, even from allies: "pain feels like a 5" became a cited 4 (11 of 14 interviews, p90 of 9 days), and that public correction is where the instrument's credibility gets minted.
  • Log disagreements verbatim instead of averaging them silently, then settle on the anchor text: the data readiness fight (steward's 2 versus team lead's 4) resolved to a conditional 3 that both could sign.
  • Attach conditions as written tripwires that travel with the score forever: "3 conditional on the two funded remediations landing by pilot start" is what keeps Gartner's 60 percent data-abandonment statistic at bay.
  • Show the arithmetic in the open: eight weighted scores summing to an illustrative 3.55 against a 3.2 threshold produced a nomination with two conditions, and the whiteboard math is part of the audit trail.
  • Value the refusals as much as the nomination: the high-value contract-renewal process (feasibility 1.9) went to remediate-first with a roadmap, and the easy report-formatting task (value 1.8) got a polite not-now.
  • Use AI for consistency flags, record drafting, and the adversarial "argue each score down" pass, but never let it propose a score during the session; a number on a screen anchors the room, and people, not tools, own the decision.