Building the Process-Selection Scorecard
The pilot-selection meeting is scheduled for ninety minutes and the outcome was decided before anyone sat down. The VP of Sales has been in the CEO's ear for three weeks about an AI proposal-writing tool she saw at a competitor. The CEO himself came back from a conference with a keynote still ringing in his head and a vendor's business card in his jacket. And somewhere on the third floor, the manager of the invoice-exception team, the process with a verified $428,000 annual run rate, a 12 percent rework rate, and a nine-day tail that a baseline pack has already measured to the dollar, is not in the meeting. He was not invited. His process, the one that most needs help and is most ready to receive it, will not be nominated, because nomination in most organizations is not an analysis. It is a lobbying contest. The loudest voice wins, the best-connected use case jumps the queue, and the selection that determines whether the organization joins MIT's 95 percent or its 5 percent gets made on volume, charisma, and proximity to power. This lesson builds the instrument that ends the contest: the Process-Selection Scorecard, the signature artifact of this entire level, where everything you have built since the process inventory finally assembles into a single decision machine.
The Lobbying Contest You Are About to Replace
Start by being honest about how pilot selection actually happens, because the scorecard only makes sense as a replacement for something, and the something is ugly. In most organizations there is no selection process at all. There is a nomination ritual dressed as one. A use case arrives with a sponsor attached, and the sponsor's seniority, persistence, and calendar access do the work that analysis should be doing. The proposal-writing tool gets picked not because proposal writing is expensive, painful, measurable, or ready, but because its champion attends the right meetings. The invoice-exception process gets skipped not because it scored poorly, but because it was never scored. Its owner does not present to the executive committee. Misery without a microphone is invisible.
Now hold that ritual against the failure record this program is built on. MIT's GenAI Divide study found that 95 percent of enterprise generative AI pilots deliver no measurable profit-and-loss return. S&P Global found that 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before. Gartner predicted that 30 percent of generative AI projects would be abandoned after proof of concept, and that through 2026, 60 percent of AI projects without AI-ready data will be abandoned outright. Those numbers are not primarily verdicts on technology. They are verdicts on selection. A pilot that was chosen by lobbying inherits every weakness of the lobbying: nobody checked whether the process was stable, whether the data existed, whether the people would adopt, whether a proven tool category even existed for the problem. The pilot was doomed at the moment of nomination, and everyone spent nine months and six figures finding that out slowly.
The Process-Selection Scorecard replaces lobbying with calculation. Two axes: value (would winning here matter?) and feasibility (could we actually win here?). Every candidate process scored on the same declared criteria, with declared weights and pre-committed thresholds, published before any candidate is scored. That last clause is the entire trick, and it deserves a name: the constitutional moment. Rules written before you know who they help are fairness. Rules written after you know who they help are politics wearing math. When the criteria, weights, and thresholds are signed and dated before the first candidate touches the instrument, the VP of Sales can still argue, but she has to argue about the rules in general, not about her use case in particular, and arguments about rules in general are cheap, early, and honest. That is where you want the politics: at the rules stage, where a fight costs a meeting, not at the results stage, where a fight costs the instrument's life.
Rules written before you know who they help are fairness. Rules written after you know who they help are politics wearing math.
The Value Axis: Would Winning Here Matter?
Take the two axes slowly, because the craft is in the criteria, and the criteria only work if you understand what each one is actually measuring. The value axis asks one question: if the pilot succeeded completely, would anyone be able to find the success on a financial statement or a strategic agenda? Four criteria cover it.
Annual run-rate cost
This is where Chapter 2.2's baseline packs pay their dividend. The invoice-exception process is not "a process that seems expensive." It is a process with a verified annual run rate of about $428,000: roughly 1,150 exceptions a month at about $31 each, measured, documented, and signed. A run rate is the denominator of every future improvement claim. A 20 percent improvement on $428,000 is $85,600 a year, a sentence a CFO (chief financial officer) can act on. A 20 percent improvement on "a lot of manual work" is a shrug. Processes without baseline packs can still be scored, but they score from estimates, and the anchor definitions should explicitly discount estimates relative to verified figures, because a scorecard that treats a guess and a measurement as equal teaches the organization to stop measuring.
Pain intensity
Run rate measures size; pain measures suffering per unit. Error rates, rework loops, cycle-time tails (the p90 that runs nine days when the median runs two), escalations, and the interview evidence from Chapter 2.1's synthesis work: the themes where operators described the process with words like dread and firefight. Misery is a value signal. A process that people hate is a process where improvement gets felt, adopted, and defended, and the interview corpus you built earlier in this level is admissible evidence here. A criterion that lets coded interview themes move a score is a criterion that lets the people who live in the process testify without attending the meeting.
Volume and growth
A painful process handled forty times a year is a candidate for a checklist, not a pilot. Volume is what turns per-unit improvement into annual dollars, and growth is what turns this year's annual dollars into next year's bigger ones. A process growing 25 percent a year is quietly moving up the value axis while the committee deliberates. Score today's volume and let a documented growth trend earn a point of upgrade, with the trend cited, not asserted.
Strategic visibility
This is the criterion most scorecard builders are too squeamish to write down, and their squeamishness ruins their instruments. The truth: a modest process that the CFO personally watches will generate more organizational support, faster decisions, and better executive patience than a bigger process nobody above director level has heard of. That is a real component of value, because pilots die of executive indifference as often as they die of technical failure. You have two choices: state visibility honestly as a weighted criterion, or watch it get smuggled into the scoring anyway through mysteriously generous scores on other criteria. An honest instrument names its politics and prices them. A dishonest instrument claims to be apolitical and then bends. Write the criterion, anchor it (a 5 means an executive sponsor already owns the KPI, a key performance indicator, that this process moves; a 1 means no one above the process owner would attend a readout), and cap its weight so it flavors the ranking without ruling it.
The Feasibility Axis: Could We Actually Win Here?
The value axis is where executives live. The feasibility axis is where pilots die, and it is where every artifact this level built becomes a scored input. BCG's 10-20-70 rule says AI transformation effort is roughly 10 percent algorithms, 20 percent technology and data, and 70 percent people and process. Look at the four feasibility criteria and notice what they are: the feasibility axis is the 70 percent, plus the 20, turned into numbers. This is the axis the lobbying contest never checks, which is why the lobbying contest feeds the 95 percent.
Process readiness
Is the process stable, documented, and measurable? The evidence is Chapter 2.2's verified artifacts: a process map that survived a walkthrough, an SOP (standard operating procedure) that operators confirmed against reality, a baseline pack with signed numbers. A process that changes shape every quarter, lives in three people's heads, and has never been measured is not a pilot candidate; it is a documentation project wearing a pilot costume. The anchor language should reference the artifacts by name: a 5 has a verified map, a current SOP, and a signed baseline pack; a 3 has a draft map and partial measurement; a 1 has tribal knowledge and an org chart.
Data readiness
Straight from Chapter 2.3: the data readiness report and its verdict. READY, READY WITH CONDITIONS, NOT READY, each with named gaps and remediation prices. Gartner's finding that 63 percent of organizations lack AI-ready data practices, and that 60 percent of AI projects without AI-ready data will be abandoned through 2026, is the reason this criterion exists and the reason its anchor definitions should be merciless. A NOT READY verdict with structural gaps is a 1 no matter how exciting the use case is. The scorecard is where the data audit stops being a report and starts being a gate.
People readiness
Straight from Chapter 2.4: the People Readiness Scorecard's roll-up, with its anchored subscores for change appetite, skills coverage, champion strength, and resistance exposure. You already did the hard compression work in the last chapter; here the number simply takes its seat. A process whose team includes documented shadow-AI adopters and a strong champion scores high. A process whose owner tops the resistance heat map scores low, and no amount of value on the other axis changes what happens when a pilot lands on hostile ground.
Solution maturity
The one feasibility criterion that looks outward: does a proven category of tool exist for this problem, or would we be pioneering? Invoice-document extraction is a mature category with dozens of vendors and years of production evidence. A bespoke agent that renegotiates freight contracts autonomously is a research project. MIT's finding gives this criterion its teeth: externally partnered and purchased solutions succeeded roughly twice as often as internal builds. Pioneering is sometimes worth it, but the pioneer tax must be priced into the score, not discovered in month seven. Anchor it plainly: a 5 means multiple established vendors with referenceable deployments in your industry; a 3 means the category exists but is young; a 1 means you would be the case study.
The Instrument Craft: Anchors, Weights, Thresholds
Eight criteria are a list. What turns a list into an instrument is three pieces of craft, and each one exists to close a specific loophole that politics will otherwise crawl through.
Anchors: every criterion gets an observable 1, 3, and 5
You learned anchor discipline building the People Readiness Scorecard; now it generalizes to all eight criteria. For each one, write what a 1, a 3, and a 5 look like in observable, evidence-citable terms, before scoring begins. The reason is the same as it was in Chapter 2.4: unanchored scores drift toward whoever speaks last. In a room without anchors, "I'd call that a 4" is a social move, and the scores end up measuring the meeting's power gradient instead of the organization's processes. With anchors, a score is a claim about evidence against a written standard, and the argument becomes checkable. Anchors are also what let two different assessors score the same candidate a month apart and land within half a point of each other, which is the property that makes the scorecard an institution instead of a performance.
Weights: strategy, declared in advance and signed
Equal weights are the honest default, and the right choice when you have no strategic reason to deviate. But weights are also where strategy legitimately enters the instrument. A cash-strapped organization weights annual run rate up, because it needs the pilot to pay rent. An organization coming off its own version of the 42 percent scrap year weights the feasibility axis up, because it cannot afford another corpse and needs a win more than it needs a big win. Both are defensible. What makes them legitimate is sequence: the weight sheet is debated, justified with one line of reasoning per weight, and signed by the sponsor with a date, before any candidate is scored. Get the politics done at the rules stage, where it is cheap. A weight argued about in the abstract is strategy. The same weight adjusted after the results are visible is corruption, and everyone in the room will know it even as they nod.
Thresholds: pre-committed, including the one that saves you
Two thresholds, written down before scoring. First, a nomination floor: below a stated total, no pilot this cycle, no matter what else is true, because a weak field should produce zero pilots rather than the least-bad pilot. Second, and more important, a feasibility gate: if any feasibility dimension falls below a stated floor, the candidate goes to remediation first, regardless of its total. This is the readiness-gated logic that Gartner's abandonment predictions point toward, turned into a standing rule, and it exists to defeat one specific monster: the high-value, low-feasibility candidate. Call it the siren-song quadrant, because committees keep steering into it, and it is worth a paragraph to understand why.
Value is legible to executives. A $2 million run rate, a strategic account, a board-level pain point: these translate instantly in an executive room, because executives are trained readers of value. Feasibility is legible to operators. An undocumented process, a data set with no resolution codes, a team whose champion just resigned: these translate instantly on the third floor and hardly at all in the boardroom, because feasibility failures are invisible until you are standing in them. And the selection meeting is full of executives. So the room systematically overweights the axis it can read and discounts the axis it cannot, and the beautiful, doomed, high-value candidate sails through nomination on the strength of everything the room can see, toward the rocks the room cannot. The feasibility gate is Odysseus tied to the mast: a rule adopted in a calm moment, precisely because you know the song is coming and you know it will sound wonderful.
The 2x2: the communication device
The scorecard computes; the 2x2 communicates. Plot every scored candidate on value versus feasibility and four quadrants fall out, each with a standing disposition:
| High feasibility | Low feasibility | |
|---|---|---|
| High value | Quick wins: nominate now. This is where pilots come from. | Remediate first: the siren-song quadrant. Fund the fix (documentation, data, people), then rescore next cycle. |
| Low value | Capability builds: cheap places to learn. Sometimes worth one small pilot purely to build organizational muscle. | Not now: decline politely, in writing, with the scores attached. |
The quadrant chart is what goes on the steering-committee slide, because a committee can absorb a 2x2 in ten seconds and the full scoring workbook never leaves the appendix. But the chart is only the projection. The instrument, with its anchors, weights, and thresholds, is the artifact.
AI's Three Jobs, and the Test the Instrument Must Pass
This is an AI-assisted lesson, so be precise about where the assistant helps and where it must not be allowed near the controls.
AI drafts anchor language. Writing observable 1-3-5 descriptions for eight criteria is exactly the kind of structured drafting a model does well and a human verifies quickly. Feed it the criterion, your artifacts (baseline pack figures, data report verdicts, the people scorecard's anchor style), and ask for anchors that cite evidence types, not adjectives. Then verify every draft against the verification habit you built in Chapter 2.2: no anchor ships until a human confirms it describes evidence your organization actually produces.
AI stress-tests the criteria. This is the highest-leverage use, and it is adversarial by design. Prompt: "Here are my eight criteria and anchors. Describe a candidate process that would score 4 or higher on this instrument and still be a terrible pilot. Then describe a process that would score below 3 and be an excellent pilot." Every exploit the model finds is a hole in the instrument. A common first-round find: a high-volume, well-documented process that scores beautifully but is about to be eliminated by a system migration, which tells you to add a stability horizon to the process-readiness anchors. Iterate until the model's exploits stop being plausible. You are red-teaming a measurement instrument the way a security team red-teams a network, before the adversary (which in this case is your own politics) gets its turn.
AI consistency-checks scores later. Once real scoring begins, next lesson and beyond, the model can compare score-and-evidence pairs across assessors and flag the pattern where one assessor's 4 matches another's 2, or where a written justification does not actually satisfy the anchor it claims. Same discipline as the people scorecard's consistency pass, wider surface.
The human owns weights and thresholds absolutely. Weights are strategy: they encode what this organization needs most this year, and that judgment belongs to the sponsor whose name is on the sheet. Thresholds are risk appetite: how weak a field still deserves a pilot, how low a readiness dimension can go before it gates. Delegating those to a model is not efficiency; it is abdication, and it ends with you in front of a steering committee explaining your roadmap with "the AI said so," which is a sentence careers do not survive. The model advises on everything and decides on nothing. That line is the whole governance model of this program in miniature.
Verification: the instrument must explain the past before it may predict the future
Here is the step almost everyone skips, and the one that separates an instrument from a slide. Before the scorecard is used on any live candidate, test it the way you would test any measurement device: against known cases. Pick two from your organization's own history: one obvious success (a project that worked, quietly and provably) and one famous failure (the pilot everyone remembers going down). Score both, honestly, using the anchors and the evidence that existed at the time. Then check: does the instrument agree with hindsight? Does it pass the success and flunk the failure, for the right stated reasons? Call this the retrodiction test. An instrument that cannot retrodict the organization's own past has no business claiming authority over its future, and running the test in front of the sponsor is the single fastest way to earn the instrument institutional trust. When the scorecard demonstrably would have caught the disaster everyone still winces about, the room stops treating it as paperwork and starts treating it as a guardrail.
The Worked Example: One Division's Instrument, Signed and Tested
Here is a complete illustration, with every figure hypothetical, from a distribution division building its first Process-Selection Scorecard. Eight criteria: four value (annual run-rate cost, pain intensity, volume and growth, strategic visibility), four feasibility (process readiness, data readiness, people readiness, solution maturity). Each scored 1 to 5 against written anchors. Two anchors, excerpted in full so you can copy the pattern:
Anchor: annual run-rate cost.
- 1: Estimated annual cost under $75,000, or no credible estimate exists. No baseline pack.
- 3: Annual cost between $150,000 and $400,000, supported by at least a structured estimate (volume x loaded cost per unit) reviewed by the process owner.
- 5: Annual cost above $400,000, drawn from a verified baseline pack with signed figures. Verified packs score one point higher than estimates of the same magnitude.
Anchor: data readiness.
- 1: No data readiness report exists, or the report's verdict is NOT READY with structural gaps (missing fields, no history, unresolved access or privacy flags).
- 3: Verdict is READY WITH CONDITIONS, all conditions priced, remediation under 90 days and funded.
- 5: Verdict is READY: required fields present and populated above the report's stated thresholds, history sufficient, access and privacy flags cleared in writing.
The weight sheet, signed by the division sponsor on March 4, with one line of reasoning per weight: run-rate cost 30 percent ("coming off last year's budget cuts, the CFO funds paybacks, not adventures"), pain intensity 10, volume and growth 5, strategic visibility 5 ("real, but capped: visibility flavors, it must not rule"), process readiness 15, data readiness 15 ("Gartner's 60 percent abandonment figure is about us if we ignore this"), people readiness 10, solution maturity 10 ("MIT's 2x buy-versus-build finding, priced in"). Total 100. Thresholds, pre-committed on the same signed page: total of 3.2 or higher to nominate this cycle; any feasibility criterion below 2.0 sends the candidate to remediate-first regardless of total.
Then the retrodiction test, run in front of the sponsor. Known failure: the division's 2024 customer-service chatbot, dead in seven months, still a sore subject. Scored with the evidence that existed at launch: value lands at 4.1 (big volume, visible pain, the CEO loved it), feasibility lands at 1.8, with process readiness at 1 (the escalation process it sat on was undocumented and had no baseline; nobody could ever prove what the bot changed) and data readiness at 2. The feasibility gate trips twice over. The instrument flunks the chatbot, for the exact reasons the post-mortem later found. Known success: the 2023 invoice-matching automation, the division's one quiet win. Value 3.4, feasibility 4.2 (documented process, clean structured data, a willing team, a mature vendor category), total 3.8: nominate. The instrument passes the success. Two for two against hindsight. The sponsor, who lived through the chatbot, signs the weight sheet a second time, unprompted, and that signature is the moment the scorecard becomes the division's constitution rather than the assessor's spreadsheet.
The failure story: the night the math became negotiable
Now the cautionary tale, because this instrument has one unforgivable failure mode and you need to see it kill. A committee at another company, hypothetical but assembled from patterns you will recognize, builds a genuinely good scorecard: eight criteria, written anchors, the works. They score eleven candidates. The CIO's favorite, a contract-analysis tool for the legal team, finishes third, behind two unglamorous back-office processes. There is a silence in the room, and then the phrase that ends everything: "we just felt volume deserved more emphasis." The weights are adjusted, after scoring, in the meeting, until the favorite finishes first. Everyone nods. The minutes record a unanimous, data-driven decision.
Six months later the contract pilot is dead, at a hypothetical cost of $310,000, and it died exactly where the original weights said it would: people readiness 2.1, a legal team that never wanted it and never fed it. But the pilot is not the real casualty. The scorecard is. Next cycle, nobody argues about evidence against anchors, because everyone now knows the endgame: score first, then bend the weights until the powerful win. The organization has learned that the math is negotiable, and a negotiable scorecard is worse than no scorecard at all, because it launders lobbying into the language of analysis and stamps politics with a decimal point. Weights changed after scoring is the one unforgivable move. The antidote costs one signature and one date, applied before the first candidate is scored, and the discipline to treat that signed page the way you treat a contract: amendable next cycle, in the open, with reasons; untouchable this cycle, forever.
This is also what the playbook means when it calls this chapter the goldmine. The scorecard's value is not that it picks flattering winners. It is that it kills the doomed pilot in week one, at the cost of a meeting, instead of month nine, at the cost of $310,000 and the organization's remaining patience. And it makes the surviving nomination defensible: when anyone asks "why this process?", the answer is a signed instrument, a scored field, and a quadrant chart, not "the VP wanted it." In a world where 42 percent of companies are scrapping initiatives and 95 percent of pilots cannot prove a return, a defensible nomination is a competitive asset. Next lesson, you will run this instrument end to end on a real candidate and watch the scoring discipline work at full depth.
What to Do Monday Morning
- Draft your eight criteria with anchors. Four value (run-rate cost, pain intensity, volume and growth, strategic visibility), four feasibility (process, data, people, solution maturity). Write an observable 1, 3, and 5 for each, citing the evidence types your organization actually produces: baseline packs, data report verdicts, people scorecard roll-ups. Use AI for the first draft; verify every line.
- Run the adversarial stress test. Ask the model for a process that would score well and be a terrible pilot, and one that would score badly and be a great one. Patch the anchors until the exploits stop working.
- Propose weights with one line of reasoning each. Start from equal, deviate only with a stated strategic reason, and cap strategic visibility so it flavors rather than rules.
- Get the sponsor signature on the weight sheet, with a date. This is the constitutional moment. Do not score a single candidate before the ink is on the page.
- Pre-commit both thresholds in writing. The nomination floor (no pilot below the line, even in a weak field) and the feasibility gate (any dimension below the floor means remediate first, regardless of total).
- Run the retrodiction test. Score one past success and one past failure with the evidence that existed at the time, in front of the sponsor. If the instrument disagrees with hindsight, fix the instrument before it touches a live candidate.
Key Takeaways
- Replace the lobbying contest with calculation: the Process-Selection Scorecard scores every candidate on declared criteria, declared weights, and pre-committed thresholds, published before any candidate is scored.
- Score value with four criteria: verified annual run-rate cost (the baseline pack's dividend), pain intensity (misery is a value signal), volume and growth, and strategic visibility stated honestly as a capped criterion instead of smuggled in later.
- Score feasibility with the level's own artifacts: process readiness from Chapter 2.2's verified maps and baselines, data readiness from Chapter 2.3's report verdicts, people readiness from Chapter 2.4's scorecard, and solution maturity with MIT's 2x buy-versus-build finding priced in. The feasibility axis is BCG's 70 percent, quantified.
- Anchor every criterion with observable 1, 3, and 5 definitions written before scoring, because unanchored scores drift toward whoever speaks last.
- Treat weights as declared strategy and thresholds as pre-committed law: the sponsor signs the weight sheet, with a date, before scoring, and weights changed after scoring is the one unforgivable move that teaches the organization math is negotiable.
- Gate on feasibility to survive the siren-song quadrant: value is legible to executives, feasibility is legible to operators, and a room full of executives will steer into the high-value, low-feasibility monster unless a standing rule ties it to the mast.
- Keep weights and thresholds human: AI drafts anchors, stress-tests criteria adversarially, and consistency-checks scores, but strategy delegated to a chatbot ends with "the AI said so" in front of a steering committee.
- Run the retrodiction test before first use: an instrument that cannot correctly flunk your organization's famous failure and pass its quiet success has not earned authority over its future.
Skill.re