Auditing People, Skills, and Culture
The chief people officer slides a single page across the table. It is the people dimension from the enterprise readiness assessment, and it says 2.9 out of 5, with a footnote: 71 percent of surveyed staff report using unsanctioned AI tools for real work. The CEO reads it twice and asks the question this whole chapter exists to answer: "Is that number about 2,400 people, or about the twelve people who filled in the survey honestly?" Nobody in the room knows. And underneath the silence sits something heavier than a data-quality problem. Two years ago this company scrapped most of its AI initiatives in a single budget cycle, the internal shorthand is "the scrap year," and every one of those 2,400 people watched it happen. Whatever the audit finds about skills and appetite, it will be read by an organization that remembers. This lesson is about auditing people, skills, and culture at enterprise scale: honestly, safely, and without destroying the trust you are trying to measure.
Three Things No Single Team Can Show You
In Level 2 you built the process-level people assessment: interviews synthesized into themes, a stakeholder heat map, a skills matrix for one redesigned workflow, a shadow-AI survey with an amnesty framing, and finally the People Readiness Scorecard that gave one team's readiness a defensible number. That toolkit still works, every instrument in it reappears in this lesson, but at process scale it was answering a small question: does this one team have the appetite, the skills, and the champions to carry one pilot? At enterprise scale, three phenomena appear that no single team can show you, because they only exist in the aggregate.
The first is culture. A team has a mood; an organization has a culture, which is something closer to accumulated case law. Culture is the organization's settled beliefs about change, formed by every initiative that ever over-promised, every transformation that quietly became a logo on a coffee mug, every "this time is different" that turned out to be the same. You cannot interview one team and find it, because each team holds only a fragment of the record. And in this company the most recent precedent is brutal: the scrap year. S&P Global found that 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before, which means the scrap year is not an unusual scar. It is one of the most common formative experiences in the modern enterprise, and it is in the room for every conversation your audit will have.
The second is distribution. At process scale you asked "how ready is this team?" At enterprise scale the question "how ready are we?" is almost meaningless, because skills and appetite are never evenly spread. An average hides everything that matters. The real enterprise question is a terrain question, the same logic you applied to processes in the last lesson, now applied to humans: where is readiness pooled, and where is it absent? A company where AI capability averages 30 percent might have one function at 52 percent and another at 11, and those two functions need entirely different investments, sequenced differently, priced differently. The audit's job is to draw that terrain, not to compute its average.
The third is institutional memory. Organizations remember. They remember the 2023 chatbot that answered customer questions confidently and wrongly. They remember the offshoring wave and who was told their roles were "evolving" the month before the announcement. They remember the last transformation program, the one with the roadshow and the branded lanyards. When your AI program launches, every employee will price it against that memory, not against the press release. MIT's autopsy of the 95 percent of failed GenAI pilots found adoption without transformation everywhere: tools used, nothing changed. One reason is exactly this: people extend effort in proportion to what history tells them effort earns. Audit the memory, or be ambushed by it.
So the enterprise people audit measures three things the process audit could not: the culture the organization has accumulated, the distribution of capability and appetite across its terrain, and the institutional memory that will price the next program before it launches. And it does so under one cardinal rule that doubles down on everything Level 2 taught about anonymity and amnesty.
At enterprise scale, the people audit measures trust while spending it. Aggregate, anonymize, and act on systems, because 2,400 people's trust is the asset being measured, and measuring it as surveillance destroys it mid-measurement.
Hold that rule through everything that follows. The output of this lesson is a named artifact, the People Readiness Atlas: a function-by-function map of skills distribution, adoption evidence, champion coverage, change capacity, and the culture read, with every claim traceable to its evidence type. Four audit streams feed it. We take them in order.
Stream One: Skills Distribution at Scale
The Level 2 skills matrix listed roles down the side and required skills across the top for one workflow. You cannot interview 2,400 people, so at enterprise scale the matrix goes statistical, and statistics done casually about people is how audits get laughed out of steering committees. Done carefully, it is honest arithmetic a non-statistician can run and defend.
The task-based survey: ask about behavior, not confidence
The core instrument is a self-assessment survey, but not the kind that asks people to rate their "AI proficiency" from 1 to 5. Confidence ratings measure personality, not capability; the cautious expert scores herself 2 and the enthusiastic novice scores himself 4, and the data is noise wearing a chart. Instead, the survey is task-based, the have-you-done-this design: "In the past month, have you used an AI tool to summarize a document longer than five pages? To draft a message you then edited and sent? To check its output against a source before using it?" Concrete, recent, binary. People are far more accurate about what they did last month than about what they are. Attach the caveat in writing when you report it: this is still self-report, and self-report drifts optimistic, which is exactly why the survey never travels alone.
The work-sample calibration: thirty volunteers, twenty minutes each
Here is the move that turns a survey into evidence: calibrate it with sampled work-sample tests. Recruit roughly 30 volunteers spread across the functions, paid in goodwill and a genuine thank-you, and give each one a 20-minute practical exercise matched to the survey's task claims: here is a messy meeting transcript, produce a usable summary with an AI tool of your choice; here is a flawed AI draft, find the two errors we planted. Score against a simple rubric. Now compare: if a function's survey says 60 percent of its people are capable of a task and the work samples from that function say 25 percent performed it competently, you have a calibration ratio, and you apply the deflator to that function's survey scores before they enter the atlas. You are not accusing anyone of lying. You are doing what every good pollster does: measuring the gap between report and behavior, and correcting for it. State the method in one paragraph in the atlas appendix and it survives any challenge a steering committee can throw, because the challenge "how do you know the survey is right?" has the answer "we tested it."
Distributions, not averages
The output of stream one is never a company-wide number. It is a per-function distribution: "Function B has a deep prompt-capable core, roughly half its people demonstrably task-capable, and a long novice tail; Function E is uniformly novice, with no measurable capable core at all." Those two sentences dictate two different training programs at two different price points, a distinction Chapter 4.4 will build on when it designs the enablement plan. A single average of the two would have dictated one mediocre program that fits neither, which is how training budgets die. BCG's 10-20-70 rule holds that 70 percent of AI transformation effort belongs in people and process; stream one is where that 70 percent stops being a slogan and becomes a map of exactly which people, where.
Stream Two: Adoption Evidence, Because Behavior Beats Sentiment
Stream one measures what people can do. Stream two measures what they actually do, and it outranks every sentiment question on the survey, because behavior is the only readiness signal that cannot be performed for the auditor. McKinsey's State of AI research found 88 percent of organizations already using AI regularly somewhere; the enterprise question is never whether your workforce uses AI. It is where, with what, and whether officially.
The enterprise shadow census and the CEO's confession
This is the Level 2 shadow-AI survey scaled up, and the amnesty mechanics scale with it. At process level, a team lead confessing their own unsanctioned ChatGPT use unlocked honest answers from twelve people. At enterprise level, the leader-confession mechanic reaches its most powerful form: the CEO, on video or at the all-hands, describing the unsanctioned tool they personally used last week and declaring the amnesty in their own name. It works because it inverts the org chart's usual information flow: the person with the most to lose from candor goes first. The census itself asks the same behavioral questions as Level 2: which tools, for which tasks, how often, what would you need to do this officially. Anonymous by design, reported only in aggregates, with an opt-in line at the end for anyone willing to be contacted as a potential champion. That opt-in line quietly feeds stream three.
Sanctioned-tool telemetry: aggregate only, and say why out loud
If the organization has sanctioned tools, their usage logs are legitimate evidence, under one absolute rule: read at aggregate level only, usage by function, never by person. State the rule publicly before you pull a single log, and state the reason: the moment one individual's usage appears in one deck, every future survey in this company returns fiction, and the audit function itself becomes untrusted infrastructure. The telemetry's value is the cross-check it enables, not the individuals it could expose.
The demand-and-distrust signal
Now cross-reference three layers: the shadow census, the sanctioned telemetry, and the workaround inventory your process sweep produced in the previous lessons. Where shadow use is high and official adoption is low, you have found the audit's most valuable pattern: a demand-and-distrust signal. Demand is proven, people are already doing the work with AI; distrust is proven, they are refusing the official channel to do it. That means the tooling is wrong or the policy is, and the audit says which by asking: the census's "what would you need to do this officially" question sorts the answers into "the sanctioned tool cannot do what my phone can" (a tooling finding) and "I do not want my name on AI output here" (a policy and culture finding, and a bridge to stream four). MIT's research described this shadow economy precisely: employees privately crossing the GenAI divide while their employers' official programs stall. Your audit turns that anecdote into a per-function map.
Streams Three and Four: Champions, Capacity, and the Culture Read
Champion coverage: gaps matter more than counts
The census opt-ins give you a champion roster. The rookie move is to report its size; the enterprise move is to map it against the organization chart, because coverage gaps matter more than counts. Nineteen champions sounds healthy until the map shows all nineteen sitting in two functions, at which point you are looking at a broadcast tower with no receivers in four buildings: enthusiasm will circulate where it already lives and never cross the walls. The atlas therefore shows champion density per function, and every empty zone becomes a named recruitment target for Chapter 4.4's enablement plan. Champion quality matters too, and Level 2's distinction returns: an enthusiast is excited about tools, an influence-holder is the person colleagues actually consult before changing how they work. You verify quality by sampling the colleagues, not the champions: "who would you ask before trusting a new tool with real work?" Two or three names recur per team, and if none of them are on your roster, your roster is a fan club.
Change capacity: load arithmetic before psychology
The fourth stream is the hardest, so start with its easiest half: change capacity as arithmetic. Change capacity is a budget, and before you speculate about willingness, list what is already spending it. What else is this organization digesting? The enterprise resource planning (ERP) migration, the system that runs finance and inventory and is currently consuming every operations lead's evenings. The reorganization that moved 300 people in March and whose dust has not settled. The layoff rumor nobody official will confirm and everybody unofficial has priced in. Build the initiative-collision calendar: every major program touching each function over the next three or four quarters, on one page. It is the least glamorous artifact in the atlas and often the one leadership acts on first, because it converts "the people are not ready" into "these two functions have no absorption capacity until Q4, and here is what is consuming it." That is a sequencing input, and Chapter 4.2's roadmap lesson will consume it directly.
The culture read: three evidence types, zero fake precision
Then the culture read itself, done with humility, from three evidence types that triangulate each other.
- Interview themes at scale. Run 40-plus interviews across functions and levels, coded exactly per the Level 2 synthesis method, but aimed at institutional memory: "Tell me about the last big change here. What happened to the people who went early?" Do not ask people to describe the culture; the answer to the early-adopter question IS the culture, told as a story the interviewee believes is just history.
- The survey's trust items. One or two behavioral trust questions ride along on the census: "If an AI tool you used produced an error that reached a customer, would you report it?" This is the psychological-safety proxy, and it predicts whether your future incident playbooks will ever be triggered by anything other than luck. Level 5's culture lesson builds a full instrument here; the audit needs only the early-warning light.
- Behavioral residue. Artifacts again, the auditor's favorite witness because artifacts cannot perform enthusiasm. Does the improvement-suggestion channel get responses, or does the suggestion box have cobwebs? When did a frontline suggestion last change a standard operating procedure (SOP)? Residue tells you what the organization does with initiative, which is what your champions are about to find out the hard way if you do not check first.
And then the discipline that separates a credible culture read from a consulting cartoon: report culture as themes with strength counts, never as a single score pretending to precision. "Early adopters got extra work, not recognition: 22 of 43 interviews" is evidence. "Culture: 2.7" is astrology with a decimal point. Some things are honestly qualitative, and saying so, in a document otherwise dense with calibrated numbers, is the credibility move that makes the numbers believable too.
| Stream | Primary instruments | Atlas layer produced | Discipline that keeps it honest |
|---|---|---|---|
| Skills distribution | Task-based survey plus sampled work-sample tests | Per-function capability distribution | Calibration deflator; distributions, never averages |
| Adoption evidence | Shadow census, aggregate telemetry, workaround inventory | Adoption-evidence map with demand-and-distrust flags | Amnesty sponsored from the top; aggregate-only rule |
| Champion coverage | Census opt-ins mapped to org chart, colleague sampling | Champion density map with named empty zones | Gaps over counts; influence verified by colleagues |
| Change capacity and culture | Collision calendar, coded interviews, trust items, residue | Change-capacity index and the culture read | Load arithmetic first; themes with counts, no fake score |
Assembling the People Readiness Atlas
Four streams become one artifact through three synthesis disciplines, and these disciplines are where the atlas earns the word "audit" instead of "survey."
Triangulation, with divergence reported as a finding. Every atlas claim carries its evidence type, and where evidence types disagree, the disagreement goes in the atlas rather than getting smoothed away. Self-report versus work samples, sentiment versus telemetry, interview claims versus behavioral residue: when a function's self-reported capability runs double its demonstrated capability, that perception-capability gap is itself an atlas layer, because a function that believes it is ready and is not will skip training, trust unverified output, and generate your first incident. The gap is not measurement error. It is one of the most operationally useful things the audit finds.
The anonymity architecture, at scale. Level 2's anonymity promise becomes engineering. Before any data is collected, set minimum cell sizes: no statistic is reported for any group smaller than N, with N typically between 8 and 15 depending on how identifiable roles are, and cells below N roll up to the next level or do not report. Write the rule into the survey's front page so respondents can see the protection before they answer. One paragraph of jurisdictional awareness belongs here, not legal advice, just competence: in many European jurisdictions, works councils have codetermination rights over employee monitoring and surveys, and in any jurisdiction the partnership with human resources (HR) and, where applicable, the works council should be sealed before launch, not negotiated after a grievance. An audit that surprises its own works council has already published its most important cultural finding, about itself.
The asset frame, the audit's political spine. Here is the single reframe that determines whether the atlas launches a program or a crackdown. When the census reports 71 percent unsanctioned use, the compliance reflex reads a violation: 1,700 people breaking policy. The asset frame reads an inventory: roughly 1,700 people who taught themselves a new class of tool, on their own time, with zero training budget, and who are telling you exactly which tasks they think it is good for. This is the Level 1 shadow-AI lesson completing its arc at enterprise scale. The atlas's headline is what the organization already HAS: self-trained users, proven demand, named champions, and a free requirements document for the sanctioned stack. Lead with the asset, and the same executives who would have funded a lockdown will fund an enablement plan instead. Lead with the violation, and read the failure story below.
The worked example: one enterprise atlas in five vignettes
All numbers hypothetical and illustrative; the shapes are what to learn.
Skills. The task-based survey across 2,400 people finds 55 percent claiming comfort using AI for work tasks. The 30 work samples calibrate that to roughly 30 percent demonstrably task-capable, and the terrain is wildly uneven: Function B calibrates at 52 percent, Function E at 11. One deflator paragraph in the appendix; two different training designs already visible; a perception-capability gap flagged for the two functions where self-report ran hottest.
Adoption. The shadow census, launched the week after the CEO's confession at the all-hands, draws a 64 percent response rate, which for an anonymous enterprise survey is itself a cultural data point. Findings: 71 percent have tried AI tools for work, 38 percent use them weekly, the tool census finds 11 distinct tools in circulation of which 2 are sanctioned, and the top demanded tasks are summarization, drafting, and data cleanup. Two functions light up as demand-and-distrust zones: heavy shadow use, near-zero sanctioned use, and census comments pointing at the tooling, not the policy.
Champions. Forty-one opt-ins. The coverage map shows them pooled in four functions and entirely absent from Functions E and F, which are now named recruitment targets with owners, not vague hopes.
Change capacity. The collision calendar finds the ERP go-live sitting in Q3, squarely on top of the two functions the process sweep rated most AI-ready. That single row will move the entire roadmap in Chapter 4.2, and it cost one afternoon of calendar archaeology. People capacity is a sequencing input; this is what that sentence looks like in practice.
Culture. Three themes with counts from 43 coded interviews. "Early adopters got extra work, not recognition": 22 of 43, the incentive finding Chapter 4.4 must answer before recruiting a single champion in the empty zones. "Mistakes here travel upward fast": 17 of 43, the psychological-safety flag that tells you error-reporting will not survive contact with reality unprotected. "The veterans take craft pride in doing it right": 15 of 43, the asset, because Level 2's verification story already proved craft pride converts into the best verification culture money cannot buy.
The failure story: the surveillance own-goal
Now the same audit, run by the compliance reflex. A firm of similar size decides to measure workforce AI readiness properly, which its steering committee interprets as individually. It mandates an "AI Skills Certification Quiz," name-attached, completion tracked by manager, and quietly enables browser telemetry on AI-related domains, announcing both in a compliance memo with the phrase "failure to complete may be reflected in performance reviews." The numbers arrive on schedule and look magnificent: 96 percent quiz completion, pass rates suspiciously clustered just above the threshold, and reported shadow-AI use collapsing from an early informal estimate near 70 percent to 9. Leadership celebrates the 9 percent as policy success. It is nothing of the kind: the usage moved to personal phones, where it is now invisible to telemetry, untouchable by policy, and permanently outside any future amnesty, because the amnesty's credibility was spent in the memo. Two of the firm's genuine champions, the influence-holders whose names colleagues actually mention, quietly stop championing; one tells a colleague, in a phrase that travels further than any memo, "I am not putting my name near this." The works council files a grievance over the telemetry, freezing all people-analytics work for two quarters. Eighteen months later the firm pays a consultancy roughly $400,000, hypothetically but plausibly, to run a "listening exercise" and rebuild the trust its own audit spent, and the consultants' first recommendation is an anonymous survey with a leadership amnesty, the instrument the firm could have run itself for the cost of a survey license. The audit produced numbers and destroyed the thing it measured. Measure people the way you would want to be measured, or the measurement is the incident.
The bridge: four dimensions, two pages
With this lesson, the enterprise audit's fourth dimension is in hand: strategy and leadership, data and technology, process maturity, and now people, skills, and culture. Four rigorous documents that no executive committee will ever read end to end. The next lesson closes the chapter by building the artifact they will read: the readiness heat map, two pages that compress all four dimensions into the picture leadership acts on, without breaking a single evidence chain on the way down.
What to Do Monday Morning
The atlas takes weeks; its foundations take one week, and three of these five steps cost nothing but calendar time.
- Draft the task-based skills survey and its calibration plan together. Fifteen have-you-done-this-in-the-past-month questions mapped to your real workflows, plus a one-page plan for 30 work-sample volunteers at 20 minutes each with a simple rubric. Design them as one instrument; a survey without its calibration plan is a press release.
- Get the CEO's amnesty confession on the calendar. Not approval for a survey: a date, a venue, and a first-person story of their own unsanctioned use. If the sponsor will not confess, you have your first culture finding before the census launches, and the census should wait.
- Draw the champion coverage map from opt-ins you already hold. Level 2's survey, pilot volunteers, anyone who ever raised a hand: plot them against the org chart today and mark the empty zones. The gaps are recruitment targets for Chapter 4.4, and leadership grasps an empty-buildings map faster than any statistic.
- Build the initiative-collision calendar for the next three quarters. Every major program touching each function, one page: the ERP go-live, the reorg aftershocks, the audit season. One afternoon of work, and it becomes the sequencing input Chapter 4.2's roadmap cannot responsibly be built without.
- Set your minimum cell sizes and seal the HR and works-council partnership before any data is collected. Write the no-stats-under-N rule into the survey's front page, and get the partnership agreed in writing. Anonymity architecture designed after collection is an apology, not an architecture.
Key Takeaways
- Audit the three phenomena only enterprise scale reveals: accumulated culture, uneven distribution of skills and appetite, and the institutional memory that prices every new program against the last one, including the scrap year that 42 percent of companies share.
- Replace "how ready are we?" with "where is readiness pooled and where is it absent?", and report per-function distributions, never company-wide averages that fit no one.
- Calibrate self-report with sampled work-sample tests: 30 volunteers and 20 minutes each turn a survey into evidence, and the deflator between claimed and demonstrated capability is a finding in itself.
- Weigh behavior over sentiment: run the shadow census under a CEO-sponsored amnesty, read sanctioned-tool telemetry at aggregate level only, and treat high shadow use with low official adoption as a demand-and-distrust signal the census can diagnose.
- Map champion coverage against the org chart and hunt gaps, not counts; verify champion quality by asking colleagues who they actually consult, and hand the empty zones to Chapter 4.4 as recruitment targets.
- Compute change capacity as load arithmetic first: the initiative-collision calendar lists what is already spending the change budget, and it feeds the roadmap's sequencing directly.
- Report the culture read as themes with strength counts from interviews, trust items, and behavioral residue, never as a single score pretending to precision; honest qualitative reporting is what makes the calibrated numbers credible.
- Frame the atlas around assets, protect it with minimum cell sizes and an HR and works-council partnership sealed before launch, and remember the cardinal rule: measure people the way you would want to be measured, or the measurement is the incident.
Skill.re