Culture: From Fear to Fluency
The wall panel outside the fourth-floor meeting rooms says "Fail fast, learn faster" in a confident sans-serif, and it cost about eleven thousand pounds to install. Eleven days before it went up, in the room behind it, an analyst named Ravi told a project review that the pricing assistant had been rounding a discount field the wrong way for six weeks. The functional director asked him, in front of nine people, why he had not caught it earlier, and the meeting moved on in under two minutes. Nothing formal happened to Ravi: no warning, no note on a file, no change to his objectives. Something informal happened to the nine people watching, and it was permanent: they learned the price of being the person who says the number is wrong. Two of them find a similar problem next quarter. Neither reports it. The error class runs another five months and costs a little over 140,000 pounds in credited discounts before a customer, not an employee, points it out. The panel is still on the wall, read by four hundred people, and it has changed none of their behaviour, because culture is not what an organization writes down. It is what people watch happen to the first person who tests it.
The Most Over-Invoked, Under-Specified Word in Transformation
Every serious analysis of why AI programs fail arrives at people. BCG's 10-20-70 rule puts 10 percent of the difficulty in algorithms, 20 percent in technology and data, and 70 percent in people and process. MIT's autopsy of the 95 percent of enterprise generative AI pilots that produced no measurable profit-and-loss return named adoption without transformation as a central pattern: the tool was used, the organization around it did not change. McKinsey found 88 percent of organizations using AI regularly while only about 39 percent could attribute any earnings impact to it. All of it points at the same 70 percent, and every executive who reads it arrives at the same sentence: we need a culture of experimentation.
Then they are handed nothing. "Culture" is where transformation advice goes to become an adjective. The leader asks what to do on Tuesday and receives a values framework, an internal brand, a roadshow, and a psychological-safety module. These are not stupid things. They are simply not mechanisms, and a leader with a real problem and a non-mechanism produces the only thing a non-mechanism can: a poster. The stakes are not decorative: S&P Global found 42 percent of companies scrapped most AI initiatives in 2025, up from 17 percent.
Culture is a prediction, not a mood
Here is the definition this lesson uses, chosen because it is testable rather than inspiring. Culture is what people in your organization predict will happen to them if they do a particular thing. Not what they feel, and not what they believe about the company's values: what they expect, concretely, for themselves, next Tuesday, if they report the error, push back on the model, or admit the number is bad.
The second half is the operationally important half. Those predictions are built almost entirely from observed events involving other people, not from statements addressed to the observer. Someone judging what happens if they raise a problem does not weight the chief executive's video and the incident involving their colleague equally. They weight the incident at close to 100 percent, and here is the uncomfortable part: they are right to. The statement is a claim about intent: cheap, often sincere, and routinely overridden by the pressures on a middle manager at 4pm on a Thursday. The incident is evidence about behaviour under real pressure, and any competent risk professional prefers evidence to claims. Your workforce is doing competent risk assessment about you.
An analogy from a discipline that solved this decades before AI existed. Nobody assesses a factory's safety culture by reading the signage. They ask what happened the last time a line worker stopped the line. If the worker was thanked, the stoppage was investigated as a system finding, and the worker's shift was not quietly rescheduled, the plant has a safety culture regardless of what is on the wall. If the worker was asked why they were so quick to escalate, the plant has a signage budget. Nobody is forming their view of your AI program from the values statement. They are forming it from what happened to Ravi.
The definition earns its place by converting an unactionable noun into a finite list of events. If culture is formed by observed outcomes, culture change is not a campaign but a sequence of moments in which the observable outcome either matches the stated intent or contradicts it. You cannot change culture directly. You can determine what happens in the handful of moments where the organization is watching, and that is the whole of the job.
The Artifact: The Six Culture Test Moments
The moments that decide an organization's real culture are not evenly distributed. A few situations carry nearly all the signal, because they are the ones where the stated intent is expensive to honour and cheap to abandon. This lesson's deliverable is the Culture Test Moments: six situations, each with the audience genuinely watching, the healthy response, the damaging one, and the lesson each teaches. Use it as a diagnostic by working backwards (which of the six have you faced, and what did people observe?) and as a checklist by working forwards, deciding each remaining response while nothing is at stake and your judgement is uncontaminated by a specific case.
Moment 1: The first reported AI error
Who is watching: not the whole company, but everyone who might report the next one, the people closest to the work and furthest from the program's leadership.
The healthy response: the reporter is thanked publicly, by name, and specifically, before any analysis of the error begins, and the fix is credited to them when it lands. The incident is then discussed as a system finding: what in the design let this reach production, what control should have caught it, what changes. The damaging response is subtler than shouting, and it is what nearly every competent organization does: a professional, well-run investigation that opens with who approved this and when.
The teaching point surprises leaders most: the sequence matters more than the content. Most organizations do thank the reporter, at the end, after the inquiry has cleared them. Gratitude in that order reads as gratitude conditional on innocence, an accurate description of what it was, and every observer decodes it within a day. If the thanks come first and unconditionally, the reporter is safe even if the investigation later finds they contributed to the problem. That is the distinction between a blameless reporting culture and a merely polite one.
Moment 2: The first killed project, and what happens to its people
Who is watching: everyone deciding whether to put their name on something uncertain, which is everyone whose work matters to the portfolio.
The healthy response was set out in this level's kill-discipline lesson and belongs here for a different reason. The team is visibly reassigned to good work within days, ideally in the same communication as the kill, and the sunset is described as the system working, not as anyone's failure. The damaging response takes three forms: limbo, where nobody says anything for six weeks; quiet reassignment to unglamorous work, read instantly and correctly as a demotion; and the sponsor's visible disappointment, which teaches that the stage gates are a formality wrapped around an expectation of success.
The lesson is the price of honesty. If ending a project is safe, your intake fills with candid people bringing uncertain ideas. If ending one costs its team their standing, every review becomes a defence, every owner an advocate, and kill discipline decays into theatre within two cycles. The reallocation is not kindness. It is what makes the gates credible.
Moment 3: The first time a metric and the new process conflict
Who is watching: every team that suspects the new way of working will be punished by the old scorecard, which in most organizations is every team.
The situation arrives predictably: a team is asked to run a human verification gate on AI-assisted output, the gate takes eleven minutes per case, and the team's key performance indicator (KPI, the number their manager is judged on) is throughput per day, which nobody adjusted. The healthy response: the metric is changed and the change is announced with its reason attached, as in "we asked you to add a control, so the throughput target moves from 46 to 38 while the control is in place, and here is when we review it." The damaging response costs nothing in the moment and everything afterwards: the team is told to do both.
This is where "we support the transformation" becomes either true or visibly false, and it is the moment most often lost, because losing it requires no decision at all: doing nothing is the damaging response. The observers are not judging whether leadership cares. They are judging whether leadership will spend something, and a metric change costs political capital in a way a message does not, which is exactly why it is believed.
Moment 4: The first time a junior person overrides an AI output and turns out to be right
Who is watching: everyone weighing whether to trust their own judgement against a system management bought and praised.
The healthy response: the override is highlighted, by name, as exactly the behaviour the human gate exists to produce. The person's reasoning, not just their conclusion, becomes training material for everyone operating that gate, and the case enters the override review as a signal about the model rather than an anomaly about the person. The damaging response requires no malice at all: silence. The override is logged, the outcome is fine, nobody mentions it. Silence teaches that the system's output is the default truth and human review is a formality that occasionally, embarrassingly, produces friction.
This is the best opportunity your program will get to establish that verification is real work rather than a rubber stamp, and it is nearly always wasted, because nothing bad happened and busy people do not celebrate the absence of a problem. Treat it as you would a near-miss report in a safety programme: the highest-value event of your week, and the one most likely to go unremarked.
Moment 5: The first time a senior leader's own idea fails the criteria
Who is watching: everyone who suspects the rules are for other people, the default assumption in any organization that has run governance before.
The healthy response: the leader accepts the finding publicly, in the room, without renegotiating it, thanks the analyst who produced it by name, and ideally asks that analyst to help reshape the idea so it can pass properly. The damaging response has three respectable-looking variants: an exception granted "given the strategic importance," a quiet re-scoring after a word with the assessor, or an override recorded so blandly that nobody outside the room hears about it.
Here is the arithmetic leaders underestimate. One exception here costs more than a year of governance investment, because it does not merely bend one decision. It reclassifies the entire apparatus from binding to advisory, permanently and retroactively, in the mind of everyone who hears about it, and they will all hear about it. Governance that applies to everyone except the person who commissioned it is a review service whose opt-out is seniority. The reverse is cheaper and just as powerful: a senior leader accepting an unwelcome finding in public is the highest-yield ninety seconds in the transformation.
Moment 6: The first honest bad quarter
Who is watching: everyone who will report a number next quarter, including the people who prepare the numbers you present upwards.
The healthy response: the bad result is reported early rather than at the last moment, diagnosed openly, presented with the response already attached, and delivered without a scapegoat. The damaging response comes in two flavours, and the softer one does more damage: a reframed narrative that finds good news inside the bad quarter, or a visible search for the person responsible.
The connection to your measurement disciplines is direct. An organization that punishes bad numbers gets good numbers and bad outcomes. The reporting layer meets the demand it is given: ask for accuracy and you get accuracy with unpleasant quarters in it, early enough to act; ask for reassurance and you get reassurance, and you learn what was happening when it grows too large to reframe. Every measurement system built in this level, the benefits ledger, the override telemetry, the escape rate, rests on how this moment goes.
| Moment | Who is actually watching | What the healthy response teaches |
|---|---|---|
| 1. First reported AI error | Everyone who might report the next one | Reporting is safe before the facts are known |
| 2. First killed project | Everyone choosing whether to own something uncertain | Ending well is career-neutral, so gates can be honest |
| 3. First metric conflict | Every team on an unchanged scorecard | Support for the new way is something leadership pays for |
| 4. First correct junior override | Everyone deciding whether to trust their own judgement | Verification is real work with authority behind it |
| 5. First senior idea that fails the criteria | Everyone who suspects rules are for other people | The system is binding, including upwards |
| 6. First honest bad quarter | Everyone who will report a number next quarter | Accuracy is worth more than comfort |
Your culture is not what you announce. It is what people watched happen to the first person who tested it.
Mechanics That Make the Healthy Response Likely
Knowing the six moments leaves the outcome dependent on a leader being wise, unhurried, and present on the day, and none of those can be relied on. The following mechanics make the healthy response the path of least resistance, so it survives a bad week, a travel schedule, and eventually the leader's departure.
Rituals that carry the message without you
A ritual is a scheduled event with a fixed shape that produces the right behaviour whether or not anyone feels like it. It is the cheapest cultural technology available and the most durable, because it is written into an agenda rather than into a personality. Four are worth installing.
- The catch of the quarter. A standing operating-review item where the most useful thing anyone caught, an error, a bad output, a control that did not fire, is named with the person who found it. Four minutes, and it converts moment 1 into routine, so the tenth reporter needs no courage.
- Kill announcements from the sponsor. The person who championed a project announces its sunset, states the criterion it failed, and names where the team has gone. Not the programme office, not governance: the sponsor, because the signal lives in who is willing to say it.
- The override-cluster review. A regular session examining patterns in human overrides, where the model's weaknesses are described plainly and without embarrassment. This normalises the idea that the system has known limits, the precondition for disagreeing with it in a live case.
- Opening the quarterly with a problem. The first slide is something not working, presented by the person accountable for it. If the quarterly always opens with a win, the organization learns the running order and prepares accordingly, and you lose the early warning you need most.
Rituals matter disproportionately at enterprise scale for a plain reason: a large organization changes its transformation leadership at least once during a multi-year program. Values that lived in one leader's conduct leave with that leader. Values encoded in a standing agenda item survive the handover, because the successor inherits a meeting with a shape and finds it easier to run than to redesign.
Language discipline, because vocabulary travels faster than policy
A senior leader's phrasing is copied downward within weeks, without anyone deciding to: into team meetings, then written updates, then how a supervisor explains a change to a new joiner. That makes vocabulary a governance instrument, and three habitual phrases do compounding damage.
| Phrase | What it actually teaches | Replacement |
|---|---|---|
| "The AI made a mistake" | Locates agency in the tool, making the system an actor with its own conduct rather than something you operate and answer for, and makes the error nobody's to fix. | "Our claims triage produced a wrong output and our check did not catch it. Here is the control change." |
| "We're all excited about this" | Speaks for people who are not, and marks the speaker as someone in front of whom a reservation cannot be raised. | "Some of this will be better and some will be disruptive. I want to hear both, and here is where." |
| "This won't affect anyone's job" (when unknown) | Spends a promise you may not keep. Broken once, every later statement from leadership is discounted, including the true ones. | "I do not know yet. Here is what I do know, the date I will know more, and the commitment I can make now." |
The third replacement is the hardest and most valuable, and it deserves care rather than a formula, which is why a later lesson in this chapter is devoted entirely to the conversation about jobs. Note the mechanism: a specific commitment you can keep, plus a dated promise to say more, beats a comforting generality every time, because your workforce is evaluating your reliability, not your warmth.
Fluency as a destination, defined in observable terms
"AI fluency" risks becoming as empty as "culture", so here it is in behaviours a manager could observe on a Tuesday. An AI-fluent organization has four visible properties.
- People describe what a system is good and bad at, without hype or dismissal. The junior operator can name the two categories where the tool is unreliable and why. Neither "it's amazing" nor "it never works" appears in an assessment.
- Disagreement with an AI output is normal and evidence-based rather than deferential. Someone says "the model has this wrong because the contract term is non-standard," and the conversation is about the contract term, not about whether disagreeing is acceptable.
- Ideas arrive with mechanisms attached. Not "we should use AI in procurement" but "supplier onboarding takes nine days, six of them document checking, and the mechanism is extraction plus a verification gate." Fluency shows up as a change in the shape of suggestions.
- Nobody needs permission to escalate a concern. The route exists, is known, and does not pass through the person whose deadline the concern threatens.
Defined this way, fluency is measurable through signals your program already collects, so no culture survey is required. Intake quality gives you mechanisms: the share of candidates arriving with a stated problem and a stated mechanism. Override rates and reasoning quality tell you whether verification is real. Incident reporting volume tells you whether reporting is safe, and the healthy pattern is counter-intuitive: reports rising while escapes to customers fall. The shadow census, your anonymous count of unsanctioned tool use, tells you whether the sanctioned path is trusted, and healthy programs show shadow use falling as sanctioned use rises. MIT's finding of a thriving shadow AI economy inside organizations whose official pilots were stalling is that same measurement at national scale.
The Honest Limit of Culture Work
This section is short because it needs believing rather than admiring. Culture work cannot compensate for a broken incentive system or a dishonest message about jobs. If a team's bonus depends on a throughput number the new process makes unreachable, no ritual, no vocabulary change, and no amount of leadership humility will produce adoption, and none should. Those employees are responding correctly to the system they can observe. Fear, in the framing this program established at the previous level, is a rational inference rather than an attitude problem, and inferences are corrected by changing the evidence, not by improving the reassurance.
There is a sharper version of the warning. A leader who runs a culture programme while the incentives contradict it is not merely wasting money. They will be read as manipulative, and fairly, because changing how people feel about a system while leaving the system intact is a working definition of manipulation. That perception is expensive, slow to reverse, and worse than doing nothing.
So the sequencing rule from the incentive-repair lesson holds, and it is the operating instruction for everything above: repair the system first, then let the observable events do the cultural work. Run the four audit questions (where do the saved hours go, who bears the transition cost, which existing metrics conflict with the new process, and is it safe to report a mistake), fix what they surface, and only then invest in rituals and language. Where the honest answer about jobs is difficult, and in most real transformations it is, do not paper over it here: that conversation gets its own lesson, including what to say when the truthful answer is that some roles will change substantially and you do not yet know which.
Get both right, the incentives and the six moments, and the harder question follows: how to make readiness itself continuous rather than a programme with an end date. That is the flywheel, and it is next.
Worked Example: Two Years of Norvik's Moments
Norvik Group is the 2,400-person business-to-business industrial distributor this level has followed: a scrapped year, then a governed rebuild, now two years into a funded programme across six functions. All figures below are illustrative, shown for shape rather than benchmarking.
Moment 1 arrived in week five of the first wave, far earlier than leadership expected, which is itself the lesson: these moments arrive during implementation, not during the culture phase planned for later. A claims-and-credits clerk noticed the categorisation assistant was mis-filing a class of returns from one product family. She told her team lead, who did something small and correct: he named her in the Friday note, said she had found a pattern affecting an estimated 310 items, and promised her credit for the fix. The note said nothing about how the error got through; that ran the following week as a design review. Reported issues across the wave the following month ran at roughly four times the prior rate. Ninety words in a Friday note did that.
Moment 5 arrived in year two and was much harder. A function head proposed a customer-facing recommendation use case that failed the intake screen on a missing mechanism: it named a benefit but could not say which process step would change, or what evidence would show it. The analyst who wrote the assessment was seven years his junior. At the leadership meeting he accepted the finding in one sentence and asked the analyst to help reshape it. It cost him perhaps ninety seconds of visible discomfort. Over the next two months, three people in unrelated functions described that exchange, unprompted, as the moment they concluded the gates were real. What travelled was not the decision, which few cared about, but the conduct.
Moment 3 was handled badly the first time, and that account matters more than the successes. A pricing team was asked to run a verification gate on AI-drafted quotes while holding its existing throughput target. Nobody decided to squeeze them; the metric lived in a different system, owned by a different function, and nobody connected the two. The result was two months of rushed reviews: the gate was performed, ticked, and effectively empty, which is worse than no gate because it produces documented assurance that nothing was checked. It surfaced only when that team's override rate showed as an outlier at 3 percent against a 17 percent programme average. The target was cut from 46 to 38 quotes per day, announced with the reason attached, and the transformation lead told the operating review that the delay had been the programme's error rather than the team's. On the evidence of the next quarter's escalations, that last sentence was the most useful part of the fix.
The fluency indicators over the two years moved as the previous section predicts. Intake candidates arriving with a mechanism rather than an aspiration rose from 38 to 71 percent. Override reasoning improved from a field that was blank or said "wrong" in half of cases to structured reasons usable in the cluster review. Incident reports rose while escapes to customers fell, which reads as alarming on a one-line dashboard and is the healthy signature: finding more, shipping less. Shadow use in the anonymous census fell from 31 percent weekly to 12 percent as the sanctioned path became faster than the workaround. Not one of these numbers came from a culture survey.
The failure story: the culture campaign that was tested in week six
A retail bank launches an AI culture initiative with real money and sincerity behind it: a values framework with four pillars, an internal brand, a leadership roadshow through eleven sites, and a psychological-safety module every people manager completes. The budget is in the high six figures. Nobody in this story is a villain.
Six weeks after launch, an analyst in a lending division escalates a concern about a customer-facing AI output: affordability summaries look wrong for a category of self-employed applicant. She raises it with her line manager, whose objectives that quarter are dominated by a delivery date. He handles it as a performance conversation, not aggressively: he suggests she is being perfectionist, notes the deadline is immovable, and asks her to raise such things earlier next time. The concern is not passed on, and she stops pursuing it.
The story reaches the division within a fortnight, the way these stories always travel, and the rest of the bank within a quarter. The campaign runs on as scheduled, landing on an audience that now knows precisely what this organization does when tested. Attendance is good, the post-programme survey is warm, and reporting volumes do not move.
The postmortem finding is worth memorising, because it is the general case: the campaign was sincere, well-funded, well-executed, and irrelevant, because it never touched the manager's incentives or the escalation path. He responded correctly to the pressures on him: a hard date, and no recognition for surfacing a problem that would move it. And the escalation route ran through the one person with a reason to suppress it. Two structural facts, both fixable in an afternoon, outweighed an entire culture programme. An organization's culture is the set of outcomes it produces in the moments people are watching, and a campaign is never one of them.
What to Do Monday Morning
The six moments are not hypothetical for most readers. Some have happened already, and the organization has drawn its conclusions. Start there.
- Audit the moments you have already faced. Write down honestly which of the six have occurred and what people observed. Ask two people close to the work rather than trusting your own recollection: your view is the sponsor's view, and the sponsor was not the one being watched.
- Script the responses for the ones still ahead. For each remaining moment, write the first sentence you will say and the sequence you will follow, particularly the rule that thanks precede analysis in moment 1. Decide now, because on the day you will be under time pressure with a stakeholder's reputation in the room.
- Install one ritual this month. One, not four. The catch of the quarter is the cheapest start: four minutes on a standing agenda, a named person, a specific find. The test of a real ritual is whether it happens when you are on leave.
- Remove one phrase from your vocabulary and your leadership team's. Pick the one you use. "The AI made a mistake" is the most common and most corrosive. Say the replacement out loud until it is automatic.
- Trace your escalation path on paper. Follow a hypothetical concern from a front-line analyst to the person who can act on it, and check whether it routes through the manager whose deadline that concern threatens. If it does, you have the bank's structure regardless of your values framework, and rerouting it is an afternoon's work with a larger return than any programme in your plan.
- Check the honest limit before spending anything. Run the four incentive-repair questions on the affected teams. If incentives contradict the culture you describe, fix them and postpone the culture work rather than running both, because running both is what gets a sincere leader read as manipulative.
Key Takeaways
- Define culture as what people predict will happen to them, formed by observed events involving colleagues rather than statements addressed to them, which turns an unactionable noun into a finite list of moments you control.
- Run the Culture Test Moments as diagnostic and checklist: the first reported error, the first killed project's people, the first metric conflict, the first correct junior override, the first senior idea that fails the criteria, the first honest bad quarter.
- Put gratitude before investigation in moment 1, because thanks delivered after an inquiry clears the reporter reads as gratitude conditional on innocence.
- Treat the correct junior override as your best chance to prove verification is real work, and silence about it as the damaging response, since a near-miss costing nothing is the event most likely to pass unremarked.
- Refuse the single exception in moment 5: one senior override reclassifies the governance apparatus from binding to advisory, permanently, and costs more than a year of investment in it.
- Install rituals rather than relying on leadership conduct (the catch of the quarter, sponsor-delivered kill announcements, the override-cluster review, a quarterly that opens with a problem), because rituals survive turnover and personalities do not.
- Police the vocabulary, replacing "the AI made a mistake," "we're all excited," and unfounded job guarantees, since a leader's phrasing is copied downward faster than any policy.
- Measure fluency through signals you already collect (mechanism-bearing intake, override reasoning quality, rising reports alongside falling escapes, shrinking shadow use), and sequence honestly by repairing incentives first, because a campaign run against contradictory incentives is correctly read as manipulation.
Skill.re