Your Readiness Baseline: Scoring Your Own Org Honestly
At a leadership offsite in March, a facilitator hands eight executives a card each and asks one question: on a scale of one to ten, how ready is this organization for AI? The cards come back: a three, two fives, a six, three sevens, and a nine. Same company, same room, same quarter. The COO laughs, the CIO does not, and the group spends forty minutes debating whose number is right, which is forty minutes debating feelings, because not one of the eight numbers is attached to a single piece of evidence. Six months later this company launches two pilots anyway, aimed by the nine. Both stall. You already know from Chapter 1 which statistical bucket they land in. This lesson exists so that when the card comes to you, you do not write a feeling. You write a number you can defend line by line, because you spent one week earning it. This is your first full hands-on pass at the instrument this whole program is built around: the readiness audit, in its 20-question baseline form, scored by you, on your own organization, honestly.
You Cannot Navigate from an Unknown Position
Every discipline that moves things through space starts the same way. A ship's navigator fixes position before plotting a course. A physical therapist measures range of motion before prescribing exercise. A CFO closes the books before forecasting the year. The AI version of this rule is the entire lesson of Level 1 compressed into one sentence: the organizations that joined MIT's 95 percent did not fail because they chose the wrong destination; they failed because they never established where they were starting from. No process baselines, so no provable value. No data check, so a pilot built on fields that were 30 percent blank. No people signal, so a rollout into a team that had already decided the tool was a headcount threat. The unmeasured organization is not just more likely to fail; it is structurally incapable of proving success, which in a budget cycle is the same thing.
You have spent this level collecting the pieces of the fix. Chapter 1 gave you the failure record and the governance floor. Chapter 3 gave you the four readiness dimensions one at a time: people readiness and the change curve, process readiness and the afternoon triage, data readiness and the five-question pre-check, shadow AI as a discovery dataset. What you have not yet done is what this lesson does: assemble those pieces into one instrument, run it end to end, and produce the single page that starts everything else.
Three properties make this baseline the right first move, and each one removes an excuse. It costs no money: every question is answered by looking at things your organization already has, or conspicuously does not have. It needs no permission: you are reading documents, walking processes, and having coffee conversations, all of which sit comfortably inside any operations role. And it fits in one week: five evenings or a few carved-out afternoons, as the worked example below will show with a stopwatch running. When someone tells you their organization cannot afford a readiness assessment, they are usually picturing a consulting engagement with a six-figure invoice. You are about to build the version that costs a week of attention.
One more framing before the instrument itself. This program keeps returning to a claim: the readiness audit is the goldmine, the repeatable, sellable, career-defining service at the center of the AI-ready operations career. Level 2 will teach you to run it with AI assistance at scale, and later levels will teach you to price it. But the first audit you ever run should be manual, and it should be on the organization you know best, because you cannot detect dishonest answers in a client's self-assessment until you have caught yourself producing them.
The Honesty Problem, and the Evidence Rule That Fixes It
Self-assessments inflate. Not because people lie, but because the question "do we have documented processes?" gets silently rewritten in the mind of the answerer into "do I remember documentation existing?", which becomes "yes, there is a folder." The eight cards at the offsite ranged from three to nine because each executive was scoring a different private question: the nine was scoring enthusiasm, the three was scoring the state of the data warehouse, and nobody was scoring evidence. BCG's 10-20-70 rule says 70 percent of AI success lives in people and process; the corollary nobody says out loud is that people and process are exactly the dimensions where self-flattery is cheapest, because there is no error log to contradict you.
The fix is mechanical, and it is the design principle behind all 20 questions that follow: every question is anchored to evidence, not opinion. You never ask "is our documentation good?" You ask "when I walked three SOPs against three real executions of each, how many divergences did I find?" (an SOP, a standard operating procedure, is the written step-by-step for how a task is meant to be done). The first question invites a feeling. The second question has an answer that two different people would arrive at within one point of each other, which is the practical test of an honest instrument.
If you cannot point to the evidence, the score is zero. Not "probably one." Zero.
Every question scores 0, 1, or 2, with written anchors so you are matching your organization to a description, not sliding a feeling along a dial. Three rules of engagement before you start. First, score what exists today, not what is planned: "we are about to hire a data engineer" scores the current state, which is no data engineer. Second, when you are torn between two scores, take the lower one; the instrument's value is diagnostic, and an inflated score deprives you of the diagnosis you paid a week to get. Third, write one line of evidence next to every score as you go. The evidence lines are not bureaucracy; they are the raw material of the memo you will write at the end, and they are what makes your number defensible when the nine-card executive challenges it.
The Artifact: The 20-Question Readiness Baseline
Here is the instrument in full: five questions per dimension, each with the question, the evidence that answers it, and the scoring anchors. Everything here is a compressed, field-ready version of a Level 1 lesson you have already read, so nothing should surprise you; what is new is seeing it as one connected examination.
People: five questions
- P1. Shadow-AI presence. Are people already using AI for real work, officially sanctioned or not? Evidence: five to eight no-blame coffee conversations across departments, asking "what do you already use AI for?" and counting. Anchors: 0 = no idea, or a ban with no visibility behind it; 1 = anecdotes but no count; 2 = a named list of who uses what, for which tasks. Remember the reframe from the shadow-AI lesson: discovered users are a strength, not a scandal; they are your use-case discovery dataset.
- P2. Champion density. Are there named people who have already reshaped their own work with AI and would show a colleague how? Evidence: names on paper, and at least one of them walking you through their workflow. Anchors: 0 = nobody identifiable; 1 = one or two enthusiasts, unnamed and unsupported; 2 = named champions in at least two departments who have demonstrated a real workflow to you.
- P3. Change history. How did the last significant system or process rollout actually go? Evidence: the last rollout's adoption reality 90 days after go-live, from someone who lived it, not from the project's closing slide. Anchors: 0 = it stalled, was worked around, or quietly died; 1 = adopted late, with workarounds still visible; 2 = adopted on schedule, with a feedback loop that fixed early problems.
- P4. Fear handling. Has "will this take my job?" been asked out loud, and did leadership answer it concretely? Evidence: a written statement about job impact and redeployment, or its absence. Anchors: 0 = the topic is unaddressed or suppressed; 1 = verbal reassurance only; 2 = a written commitment on what happens to displaced hours, plus a visible skills path. The people-readiness lesson's point stands: fear is a calculation, and only evidence changes a calculation.
- P5. Skills floor. Do the people whose work AI would touch understand what these tools are and how they fail? Evidence: ask ten people across roles what a hallucination is and how they would catch one; count competent answers. Anchors: 0 = fewer than three of ten; 1 = pockets of literacy, concentrated in one team; 2 = seven or more of ten, or structured literacy training with completion records.
Process: five questions
- PR1. Documentation-reality match. Do written SOPs describe what actually happens? Evidence: pick three SOPs, watch or reconstruct three real executions of each, and count divergences. This is the single most clarifying walk in the whole instrument. Anchors: 0 = no SOPs, or divergence so broad the documents are fiction; 1 = SOPs exist and are recognizable, with real drift; 2 = documents match reality with only minor, known drift.
- PR2. Stability and exceptions. Would three performers of the same process describe the same steps, and what share of cases leave the standard path? Evidence: three short interviews plus an exception count from one week's real cases. Anchors: 0 = every case is handled bespoke; 1 = a stable core exists but exceptions exceed roughly a quarter of volume; 2 = stable path, and exceptions fall into a few named classes.
- PR3. Baseline metrics. For at least one serious candidate process, do measured numbers exist for cycle time, volume, error rate, and cost per unit? Evidence: the four cells, filled from real data, dated. Anchors: 0 = none exist anywhere; 1 = partial or estimated figures; 2 = a current, measured four-cell baseline for at least one process. Chapter 1 taught you why this cell block decides everything: no baseline, no provable value, ever.
- PR4. Process inventory. Does a list of the organization's processes exist, with owners and rough volumes? Evidence: the document itself, and its last-updated date. Anchors: 0 = no list; 1 = a partial or stale list; 2 = a current inventory with named owners. Without this, "which process should AI touch first?" is answered by whoever talks loudest.
- PR5. Redesign muscle. Has this organization ever actually redesigned a workflow, as opposed to installing a tool on top of an unchanged one? Evidence: a before-and-after process map from any past project, plus one measured delta. Anchors: 0 = never, or nobody can produce an example; 1 = once, painfully, and the maps are gone; 2 = practiced, with maps and measured results someone can show you. McKinsey's finding gives this question its weight: workflow redesign is among the strongest drivers of AI impact, and organizations that have never redesigned anything do not start with AI.
Data: five questions
- D1. Use-case-relative pre-check. Has the five-question data pre-check been run on at least one named use case? Data readiness is relative to a use case, never to the enterprise. Evidence: the pre-check's output, including a 50-record sample review. Anchors: 0 = never run, or run and failed everywhere; 1 = run with mixed results and no follow-up; 2 = passed for at least one named use case.
- D2. Access reality. When someone last needed a usable extract of an important dataset, how long did it actually take? Evidence: one historical fact, not a policy statement: the elapsed time of the last real request. Anchors: 0 = weeks, or it never arrived; 1 = days, and only via favors; 2 = a documented, repeatable path measured in hours.
- D3. Field quality. For the fields a real use case needs, what does a sample actually show? Evidence: pull 50 records and count blanks, obvious errors, and inconsistencies in the three to five fields that matter. Anchors: 0 = more than a fifth of records have problems; 1 = between roughly one in twenty and one in five; 2 = under five percent, with the known caveats written down.
- D4. Rights clarity. Do you know what you are allowed to do with the data: customer contracts, privacy obligations, vendor terms? Evidence: a written answer from legal or compliance covering one specific dataset and one specific use. Anchors: 0 = nobody knows and nobody has asked; 1 = a verbal "should be fine"; 2 = a documented answer for at least one dataset.
- D5. Source of truth. Where does a business-critical number actually live? Evidence: trace one KPI from a leadership slide back to its origin and count the hops. Anchors: 0 = competing spreadsheets with competing answers; 1 = a system of record patched by manual exports; 2 = one governed source that the slide provably comes from. Gartner's pairing hangs over this whole dimension: 63 percent of organizations lack AI-ready data practices, and through 2026, 60 percent of AI projects without AI-ready data will be abandoned.
Governance: five questions
These five are the five-artifact governance floor from Chapter 1, restated as evidence checks. They are also the five things an enterprise customer's AI questionnaire will ask for, with the EU AI Act's transparency deadline of December 2, 2026 giving the clock real teeth.
- G1. System inventory. Is there a dated list of AI systems in use, including the unofficial ones your P1 conversations surfaced? Evidence: the document. Anchors: 0 = none; 1 = official tools only, shadow use excluded; 2 = current, dated, and includes discovered shadow tools.
- G2. Acceptable-use policy. Is there a short, readable policy with concrete examples, and do people know it exists? Evidence: the policy, plus five people asked where to find it. Anchors: 0 = none, or a ban nobody follows; 1 = a policy exists but sampled employees cannot locate it; 2 = one page, examples included, findable by the people sampled.
- G3. Data boundaries. Is there a written rule about what may never be pasted into which tools? Evidence: the rule, tool-specific and category-specific. Anchors: 0 = nothing written; 1 = a vague "be careful with sensitive data"; 2 = named data categories mapped to named tools.
- G4. Accountability lines. Does every AI-touched decision have a named human owner? Evidence: names attached to each line of the G1 inventory. Anchors: 0 = no owners anywhere; 1 = owners for some systems, gaps elsewhere; 2 = a name on every line. Accountability stays human; a committee is not a name.
- G5. Incident path. When an AI output goes wrong, does everyone know who to tell, and is the path blame-free enough that they actually would? Evidence: the documented path, plus three people asked what they would do. Anchors: 0 = no path; 1 = a path exists on paper but sampled people do not know it; 2 = documented, known, and used at least once.
Scoring: the Bands, and Why the Shape Beats the Total
Twenty questions, two points each, 40 points maximum, ten per dimension. Add it up, then read the band, and read it as a speed limit rather than a grade.
| Band | Score | Honest reading | The move |
|---|---|---|---|
| Foundation | 0-13 | Readiness work is the roadmap. A pilot launched here is a donation to the 95 percent. | No pilots yet. Sequence the repairs, starting with the cheapest evidence gaps, and re-score in a quarter. |
| Selective | 14-27 | Ready in patches. Most mid-size organizations that have never run this instrument land here. | Pilot narrowly inside the strongest quadrant while repairing the weakest one in parallel. |
| Ready | 28-40 | The floor exists, the evidence habits exist. | Run the Level 2 and Level 3 method at pace; your constraint is now selection and sequencing, not readiness. |
The total earns the headline, but the profile writes the plan, and the profile matters more. A flat 16 (four everywhere) and a spiky 16 (eight in process, five in people, two in data, one in governance) are entirely different organizations. The spiky profile names your sequence for you. Strong process plus weak governance means fix the floor first, because the enterprise customer questionnaire, or the regulator, will arrive before your pilot matures, and a strong pilot inside a governance vacuum is a liability with good metrics. Strong people plus weak process means your shadow users are sprinting ahead of your documentation: channel that energy into the afternoon triage before someone automates a process nobody has mapped. Strong data plus weak people means the pilot will demo beautifully and die of quiet non-adoption, the slowest and most expensive of the deaths you studied in Chapter 1. In every case the rule is the same: the weakest dimension sets the risk, the strongest dimension sets the venue, and the pilot, if the band allows one at all, goes where the strength is while the repair crew works where the weakness is.
The Deliverable: the One-Page Readiness Baseline Memo
The score is not the deliverable. The memo is. A number in your notebook changes nothing; one page on the right desk starts everything, and this memo is the first entry in the portfolio you will finish building in the Level 1 capstone, where this instrument is the core. The template has five parts, and it fits on one page or it has failed:
- Header: organization or unit assessed, date, assessor, and the method in one line ("20-question evidence-anchored baseline; one line of evidence per score; full scoring sheet attached").
- Scores: the four dimension scores out of ten, the total out of 40, and the band, stated without cushioning language.
- Top three gaps, with evidence: each gap in one sentence, each followed by the specific evidence line that produced it. The evidence is what makes this memo unanswerable; "our data readiness is weak" invites debate, while "the claims dataset our warehouse claim rests on is four spreadsheet exports reconciled by hand monthly" invites action.
- One recommended next step per gap: small, dated, owned, and free or nearly free. Not "improve data governance"; instead "run the five-question pre-check on the claims-letter dataset by the 15th; I will do it."
- One nomination: the single process this baseline suggests should enter the afternoon triage first, with one line on why (usually: it is the only one with a measured baseline, or the strongest quadrant touches it).
Address it to the most senior person who has ever said the words "we should be doing something with AI." That person has a problem: they are exposed to the question and have no evidence underneath them. You are handing them a position fixed to the chart, three named repairs, and a nomination. In most organizations this single page is the moment its author stops being someone who executes AI decisions and starts being someone consulted on them, which is the career pivot the next lesson is about.
Five Evenings at Halden: a Worked Example
Here is the instrument run end to end, in a hypothetical composite built from the patterns this program teaches. Halden Insurance Services is a fictional 340-person commercial insurance intermediary. An operations manager there, eight years in the role, decides to run the baseline after a Monday meeting in which three directors give three confident, incompatible answers about the company's AI readiness. She gives it five evenings, roughly eleven hours total, and spends nothing.
Evening one: process. She pulls the three most-used SOPs (new policy setup, mid-term adjustment, claims first notification) and walks each against three recent real cases from the ticketing system. The walk takes three hours and produces the week's first honest number: eleven divergences across the three documents, including an entire approval step in the adjustment SOP that was retired two years ago but never deleted, and a claims triage step performed daily that appears in no document at all. PR1 scores 1: the SOPs are recognizable, but the drift is real and nobody had counted it. Her interviews the next morning score PR2 at 1 (stable core, but exceptions near 30 percent of adjustment volume). PR3 scores 2, and this is Halden's quiet asset: the claims team measured its letter-drafting cycle last year for a staffing case, so a real four-cell baseline exists for exactly one process. PR4 scores 0 (no inventory exists), PR5 scores 0 (nobody can produce a before-and-after map of anything). Process: 4/10.
Evenings two and three: data. The company's standing claim, repeated in two board decks, is "we have a data warehouse." She traces the claims-performance KPI from the latest deck backward and watches the claim dissolve: the number is assembled monthly from four spreadsheet exports out of two systems, reconciled by hand by one analyst, with version-named files as the audit trail. D5 scores 0, and the evidence line writes itself. Her 50-record sample of the claims extract finds 9 records with blank or misformatted loss-date fields, 18 percent, so D3 scores 1, barely. D2 scores 1 (the extract took the analyst four days and two favors last quarter). D1 scores 1 (a pre-check has never been run, but her sample work this week is most of one, so she scores the current state and notes the fix is nearly done). D4 scores 0: nobody has asked legal what client-contract clauses say about processing claims data with AI tools. Data: 3/10.
Evening four: people. Seven coffee conversations, framed as curiosity rather than audit, surface what prohibition-minded leadership would call a problem and this program calls a signal: 19 people across four departments already use AI weekly for real work, drafting client emails, summarizing policy wordings, reformatting schedules. P1 scores 2, because she now has the list. Two of the nineteen have genuinely redesigned their own workflow and light up when asked to demonstrate: P2 scores 2. P3 scores 1 (the last system rollout limped in four months late but did eventually stick). P4 scores 0 (the job question has been asked in two town halls and answered with a joke both times). P5 scores 1 (four of ten sampled people could describe a hallucination and how they would catch it). People: 6/10, Halden's strongest dimension, carried almost entirely by the shadow users nobody in leadership knows about.
Evening five: governance. The fastest and bleakest evening. No AI system inventory (G1: 0). An IT bulletin from last year saying "do not use ChatGPT for client data," which five of five sampled employees could not locate and two of nineteen shadow users had ever heard of (G2: 0, and she notes the ban is fiction in the P1 evidence). No written data boundaries beyond that bulletin (G3: 0). One genuine bright spot: the compliance officer's professional-indemnity register already names an accountable owner for every regulated decision process, which covers part of what G4 wants (G4: 1). A well-worn complaints-escalation path that people actually know, which is half of an incident path (G5: 1). Governance: 2/10.
Total: 15/40. People 6, Process 4, Data 3, Governance 2. Selective band, bottom third, with a spiky profile whose message is legible at a glance: real human energy, one measured process, unreliable data plumbing, and almost no floor under any of it.
The memo, as written
She writes the page on Sunday and sends it to the COO on Monday. Lightly condensed:
Readiness Baseline: Halden Insurance Services. Method: 20-question evidence-anchored baseline, scored [date], one evidence line per score, scoring sheet attached. Scores: People 6/10, Process 4/10, Data 3/10, Governance 2/10. Total 15/40: ready in patches, not ready at pace. Gap 1, governance floor (2/10). Evidence: no AI system inventory exists; the 2025 usage bulletin could not be located by any of five employees asked; 19 colleagues already use AI weekly with no boundaries or incident path. Next step: build the five-artifact governance floor (inventory, one-page acceptable-use policy, data boundaries, named owners, incident path) in three weeks; I will draft artifacts one and two by the 12th; artifact four extends the existing PI register rather than duplicating it. Gap 2, data reliability (3/10). Evidence: the claims-performance KPI is assembled monthly from four manual spreadsheet exports; an 18 percent defect rate in loss-date fields in a 50-record sample; no legal opinion on client-data rights. Next step: run the full five-question data pre-check on the claims-letter dataset only, and obtain one written rights answer from compliance, by month end. Not a warehouse project; one dataset, one use case. Gap 3, process documentation (PR1/PR4). Evidence: 11 divergences found walking three SOPs against nine real cases; no process inventory exists. Next step: one-afternoon process triage with the two team leads on the 20th; I will bring the template. Nomination: the claims-letter process enters triage first: it is the only process at Halden with a measured baseline (cycle time, volume, error rate, cost per letter, dated last year), it sits inside our strongest people quadrant (two of the 19 shadow users are on that team), and it should be our first and only pilot candidate once the governance floor exists. Recommendation: no pilot before the floor; floor in three weeks; pre-check in four; triage on the 20th; re-score the baseline in one quarter.
Notice what the memo does not do. It does not ask for budget, does not name a vendor, does not promise transformation, and does not use the word "strategy." It fixes a position, names three repairs with dates and an owner (her), and nominates one process for one afternoon. It is the least dramatic document in Halden's AI conversation, and it is the only one attached to evidence, which is why, in organizations like this one, it is the document that ends up steering.
What to Do Monday Morning
- Block the week. Five evenings or three afternoons in your calendar, this week or next, labeled honestly: readiness baseline. The instrument dies as a someday intention; it lives as a calendar entry.
- Print the 20 questions and start with the SOP walk. Three SOPs, three real executions each, count the divergences. It is the highest-yield three hours in the instrument and it teaches your hands what evidence-anchored means.
- Run the coffee conversations. Five to eight people, across departments, no-blame framing: "what do you already use AI for?" Write down names, tools, and tasks. You are simultaneously answering P1, P2, and half of G1.
- Trace one number and sample one dataset. Follow a KPI from slide to source and count the hops; pull 50 records of the dataset your most-discussed use case would need and count the defects. Two hours, and D2, D3, and D5 are scored.
- Score all 20 with the anchors, taking the lower score wherever you hesitate, and write the one-line evidence next to each. Total it, band it, and read the profile before the number.
- Write the memo and send it. One page, five parts, addressed to the most senior person exposed to the AI question. File the scoring sheet in your portfolio; the Level 1 capstone builds directly on it, and Level 2 will show you how to run this same instrument with AI doing the heavy lifting.
Key Takeaways
- Fix your position before plotting any course: the organizations that joined the 95 percent were the unmeasured ones, structurally unable to prove even their successes.
- Anchor every readiness question to evidence, not opinion; self-assessments inflate silently, and the walk, the sample, and the trace are what keep the score honest.
- Run the full 20-question baseline across people, process, data, and governance in one week, with no budget and no permission: it is Level 1 assembled into a single instrument.
- Score 0/1/2 against written anchors, take the lower score when torn, and record one evidence line per question; the evidence lines become the memo.
- Read the band as a speed limit: 0-13 means readiness work is the roadmap and no pilots yet; 14-27 means pilot narrowly in the strongest quadrant while repairing the weakest; 28-40 means run the method at pace.
- Trust the profile over the total: the weakest dimension sets your risk, the strongest sets your pilot's venue, and a spiky profile writes your sequence for you.
- Ship the one-page Readiness Baseline Memo (scores, top three gaps with evidence, one owned next step each, one process nominated for triage) to the most senior person exposed to the AI question.
- Keep the scoring sheet: this baseline is the core instrument of the Level 1 capstone, the manual first pass at the goldmine audit that Level 2 teaches you to run AI-assisted and at scale.
Skill.re