Agentic Operations: The Next Readiness Bar
The quarterly risk review at a mid-sized insurer is forty minutes in and running to time when the chief risk officer asks a question that is not on the agenda. "How many agents do we have running right now, and who can stop them?" The room handles the first half badly: four, someone thinks, unless the claims triage assistant counts, in which case five, and there is something in procurement the vendor calls an agent that nobody here has seen. The second half produces no answer at all: silence, then the operations director saying, honestly, "I would have to find out." Nothing has gone wrong here. Every one of those deployments was approved, scoped, and monitored. And yet the CRO has just found the seam where readiness stops being a property of any project and starts being a property of the enterprise. This lesson is about that seam, and about which of the things you need on the other side of it are slow enough to build that you should start before you need them.
The Three Assumptions That Quietly Stop Holding
Every readiness discipline in this program was designed for a world with a particular shape. The AI system produces an output (a draft, a summary, a classification, a forecast) and a human decides what to do with it. Baselines, verification, and the human gate all assume that artifact exists and that someone has time to look at it. Underneath it all sits one assumption so basic the program has rarely stated it: the machine proposes, the human disposes.
Agents break that shape, because agents act. Level 3, Chapter 3.4 taught you how to govern one properly: the agent-fit test asking whether the workflow needs autonomy or just needs automating, oversight patterns matched to reversibility classes R0 through R3, the Control Spec with its least-privilege permission set and audit log at reconstruction standard, the four test suites, and the survival checklist built against Gartner's three cancellation drivers. That machinery is correct, and this lesson extends rather than repeats it. Assume you have it.
What follows is what happens to that machinery when the count changes. When agent-shaped work stops being a governed exception and becomes a common pattern, three assumptions underneath your readiness stack stop holding, quietly and roughly at once.
First, oversight assumes a reviewable artifact and a human with capacity. At one agent that works: the Level 3 Gate Spec answered a staffing question with arithmetic (actions per week, minutes each, reviewer hours, hire or reallocate accordingly). At forty agents acting continuously the arithmetic does not get harder, it gets impossible, and nobody will fund the attempt.
Second, permissions were designed for humans with jobs. Identity and access management (IAM, the discipline deciding who can reach which system) was built around a person who joins on a Monday, holds a role for two years, is reviewed annually, and leaves through an offboarding checklist. Software identities need narrow access, scoped to a purpose, granted for hours rather than years, at a speed no ticket queue was built for.
Third, the audit trail assumed decisions, not decision chains. The Decision Record standard's six fields describe one decision made at one moment by one accountable party. When one agent's output becomes another agent's input, the question that matters at the incident review ("why did this actually happen?") is answered by no single record. It spans several systems, several vendors, and possibly several functions.
A word on honesty, before the content
This is a forward-looking lesson, and forward-looking lessons are where instructional material usually starts lying. So, the convention: every capability below is marked Established where the practice exists today with a track record you can inspect, and Projected where it is a reasoned extrapolation. A projection is not a guess, but it is not evidence either, and you should be able to tell your steering committee which you are quoting.
The leader's task is not to predict how far agents go, which nobody can do credibly. It is narrower and more useful: identify which readiness investments pay off under any plausible agent scenario, and make those now. The argument for acting early is not enthusiasm, it is an asymmetry in build times. Each capability below takes months to a year to build properly; an agent takes weeks to deploy. Organizations fail this transition not because they lack the capability but because they start building it on the day they first need it.
Governance that exists only per agent is not governance of agents. Every deployment can be controlled while the fleet is uncontrolled, and the fleet is what fails.
The Artifact: The Agentic Readiness Bar
The Agentic Readiness Bar is four capabilities an organization needs before agent-shaped work becomes common, each paired with a current-state test you can run this month. It is not a maturity model and produces no score, but four uncomfortable facts about your own organization and a defensible posture on each.
| Capability | What changes at fleet scale | Current-state test |
|---|---|---|
| Oversight capacity | From a staffing question to an architectural one | Compute review hours per thousand agent actions, then model ten times the volume |
| Machine identity infrastructure | From a per-agent control to an operational service | Time a scoped agent identity provisioning, and find who recertifies the existing ones |
| Chain-aware accountability | From one decision record to a reconstructable chain | Pick two connected automated steps, reconstruct why the second happened, and time it |
| Containment and blast radius | From per-agent rollback to correlated, fleet-wide action | Ask whether you could stop all agent action in one function inside five minutes, and whether anyone tried |
Capability one: oversight capacity as an architectural problem
At workflow scale, oversight was a staffing question and the Gate Spec answered it with arithmetic: two hundred agent actions a week, six minutes each, twenty hours, most of a reviewer. That calculation stays sound until the volume moves an order of magnitude, at which point it delivers an answer nobody will fund. This is the shift: at fleet scale, oversight stops being a staffing question and becomes an architectural one, because human review cannot scale linearly with agent action volume. Three capabilities substitute for linear review, and they are not equally mature.
Exception-based oversight means humans see only what the system flags: the low-confidence case, the out-of-pattern amount, the action the agent itself judged unusual. Established in part, and limited. Exception routing is well-proven in operations (statistical process control has worked this way for decades), but the honest limitation is current and specific: this control requires the system to be good at knowing what it does not know, and today's models are unevenly calibrated at exactly that. A confidently wrong model produces a case that never gets flagged. So design exception routing, and a sampling layer underneath it that reviews a share of the unflagged population precisely because the flag cannot be trusted alone.
Portfolio-level review means a human reviews patterns across many actions rather than individual ones: the distribution of amounts this week against last, the rate of an action type by hour, the shape of the exception population, the drift in how often the agent picks one path over another. Projected as a primary control; established as a skill in adjacent fields. Quality engineers, fraud analysts, and clinical safety reviewers already read distributions for a living. What is new is asking a process owner who has spent fifteen years reviewing individual cases to change what "reviewing" means, and that is not acquired by renaming the job. The reviewer of the future reads distributions, not documents, and organizations that want that person in three years start building them now.
Structural constraint means the permission set does the work review used to do. This is the strongest and most available answer today, and it is the Level 3 rule promoted from supplement to primary strategy: the strongest gate is the missing key. If the agent structurally cannot issue a refund above a threshold, cannot touch a customer record outside its segment, cannot reach the payment rail at all, those actions need no review because they cannot occur. Established: every control function knows this logic as segregation of duties. Permissions are not free. Every one you grant is a future review queue.
The current-state test. Compute your oversight hours per thousand agent actions today, using timesheet-grade numbers rather than the business case estimate. Multiply the volume by ten and read the resulting full-time equivalent (FTE) figure out loud in a governance meeting. The number is not the finding. What people say next is, because that is when most organizations discover their oversight model has an expiry date.
Capability two: machine identity and permission infrastructure
Level 3 required every agent to have its own identity, distinct from any human's, with least-privilege access and attributable actions. For one or two agents that is a control you implement by hand and document in the Control Spec. At fleet scale it becomes an infrastructure question, and the infrastructure most enterprises have was built for people.
The term of art is non-human identity (NHI): a service account, an application programming interface (API) key, a machine credential, an agent identity. Most large organizations already have more of these than employees, and most are undocumented, over-permissioned, and immortal. Agents do not create the problem; they make it acute, adding identities that are numerous, short-lived by design, and able to act with consequence. What you need is the ability to provision, scope, time-bound, review, and revoke them as a routine operation rather than as a project. Four capabilities carry that load.
- Purpose-bound, time-limited credentials that exist for the duration of a task and expire without anyone remembering to revoke them. Established: short-lived credentials are standard in modern cloud identity platforms. The gap is policy and habit, not technology.
- Delegation that carries constraints rather than expanding them. When agent A hands work to agent B, B inherits A's limits, not the union of both permission sets. Partly established, partly projected: the principle is old (a power of attorney works this way), the tooling is young, and this is worth putting to any vendor in writing.
- Automated recertification of machine access at the cadence humans already have. If you review human access quarterly and have never reviewed a service account, that asymmetry is the finding. Established: extending a mature audit practice to non-human identities is a policy change, not engineering.
- A single view of what every non-human identity can reach: one owned list showing every agent and service identity, its permissions, owner, expiry, and last use. Established in tooling, rare in practice. This is what the CRO in the opening scene was asking for.
The current-state test. How long does it take, today, to provision a scoped identity for a new agent, and who reviews the ones that already exist? Most organizations find the answers are "weeks" and "nobody." The second is the cheapest fix in this lesson: adding non-human identities to an access review that already runs for humans is a scope change, not a programme. It is also the finding most likely to matter in your next audit, because the EU AI Act's high-risk obligations under Annex III land on December 2, 2027, and every one assumes you can say who or what did a thing.
Capability three: chain-aware accountability and evidence
This is the shift that most needs naming, because it has no settled answer. The Decision Record standard from Level 3 has six fields describing one link: what was decided, by whom, on what inputs, with what AI involvement, against what policy, at what time. Run it well and any single automated step is defensible. Now put two agents in sequence: the first classifies an incoming case and enriches it, the second acts on the enriched case, and both keep excellent records. Six weeks later a customer disputes the outcome and the question at the review is not "what did the second agent do?" It is "why did this happen at all?" The answer lives in the first agent's classification, in a different system, possibly from a different vendor, joined to the second record by nothing but a timestamp and a hope.
Three requirements make chains reconstructable.
Correlation identifiers that survive across systems: one identifier, generated when the work item enters the chain, carried by every system that touches it, present in every log line. Established: ordinary distributed-systems practice any competent engineering team knows. What is not established is anyone asking for it before the integration is built. Retrofitting identifiers into three systems that already exchange data is a months-long project; specifying them beforehand is one sentence in a requirements document. Highest-leverage line item in the lesson.
Recorded intent at each step, not just recorded action. An action log tells you a refund was issued. An intent record tells you the agent issued it because it classified the case as a duplicate charge under a rule a prior step's enrichment had triggered. Projected: intent capture is harder than action capture, tooling is immature, and anyone claiming their platform captures agent reasoning completely is describing a hope. Ask for what is achievable now: the agent's inputs, the rule or policy path it followed, and its own confidence, all timestamped.
A named human accountable for the chain, not only for each link. The honest position: this is a new governance concept and it is not solved. Your RACI (the responsible, accountable, consulted, informed chart you already use) assigns accountability to process steps that map to organizational units. A chain crossing three functions and two vendors has no natural owner, and appointing one raises real questions about what authority that person holds over systems they neither own nor fund. A chain owner role, a control function owning all agent chains, a process owner given cross-functional authority: no track record yet says which works. State it to governance as an open design problem you are deciding, rather than pretending the org chart covers it.
The current-state test. Pick two connected automated steps you already run and reconstruct why the second one happened for one case from last month. Time yourself. If it takes three hours and produces an incomplete answer, you have learned what your evidence layer costs under pressure and written the requirements for your next deployment.
Capability four: containment and blast-radius design
With one agent, containment was the reversibility classification and the rollback point: know which actions are undoable, where the restore point is, and keep the irreversible ones behind a human. With many agents acting concurrently a different risk appears, and it has a name you met in Chapter 5.3: correlated action, many agents responding to the same bad input, corrupted reference file, or upstream model change at the same moment. It is the operational analogue of concentration risk, and individual monitoring is blind to it: each monitor sees a small anomaly, and only the aggregate sees an event. Four capabilities contain it.
- Rate and magnitude limits at the aggregate level, not only per agent. Fourteen agents each inside its own hourly cap can still produce a fourteen-fold spike in one action type against one customer segment, so the limit that matters is the one on the total. Established in adjacent domains: payment networks, trading systems, and telecoms have run aggregate circuit breakers for years.
- Kill switches that work at the fleet level. A per-agent stop is necessary and insufficient: you need one control that halts agent action across a function, and you need to know what happens to work in flight when it fires. Established as a pattern, rarely verified in practice.
- Staged rollout of agent behaviour changes. When a prompt, policy, model version, or configuration changes, it should reach a subset first, because simultaneous updates convert an ordinary regression into a correlated one. Established: standard release engineering. It is not universal for agents because their changes often arrive from the vendor's side, a contract question as much as a technical one.
- The drill discipline. A kill switch that has never been fired is a design document. Test it on a schedule, in production conditions, and write down what you learn. Established: aviation, nuclear operations, and hospital incident response all concluded that untested emergency procedures fail in use.
The current-state test. If you needed to stop all agent action in one function within five minutes, could you? And has anyone tried? The second question is the real one, because the answer to the first is a confident yes until someone tests it.
Running the Bar: A Worked Assessment
Here is the Bar run as a deliberate exercise on the enterprise storyline this program has followed, with hypothetical numbers chosen to be realistic rather than reported. The company has two agents in production: early enough that every finding is cheap to fix, late enough that the numbers are real.
Oversight capacity. The two agents together generate roughly 12,000 actions a month. Review time, measured over three weeks rather than estimated, is 11 minutes per hundred actions: exception review, spot sampling, and the weekly anomaly walk-through. That is 22 hours a month, about 0.14 FTE, absorbed inside two existing roles and sustainable. At ten times volume: 120,000 actions, 220 hours a month, roughly 1.4 FTE of pure oversight. Nobody argues with the arithmetic, and the argument that breaks out is better: is 1.4 FTE unaffordable, or merely unbudgeted? The decision is not to redesign the architecture yet, since the volume trajectory is unknown. Instead two reviewers begin an experiment in portfolio-level review: one afternoon a month reading distributions instead of cases, with a written note on what they caught that case review would have missed, and the reverse. Cost: 8 hours a month. Purpose: to have the skill in the building before the architecture question is forced.
Machine identity. The provisioning test produces the most embarrassing number: the last scoped agent identity took 19 days end to end, most of it queue time across three teams, none of it technical. The recertification question gets the classic answer: the two existing agent identities have never been reviewed by anyone, because the quarterly access review's scope says "users." The fix costs one paragraph: non-human identities enter that review, with the named agent owner attesting. Cheapest item on the list, most expensive to explain to an auditor.
Chain evidence. The reconstruction test runs on the only two connected automated steps in the estate. It takes 3 hours across two analysts and produces an incomplete answer: both records are clean, but the join between them relies on matching a case number to a time window, which works here and would not work under load. Nobody proposes retrofitting: the finding becomes a design requirement for the next deployment (correlation identifiers, achievable intent capture, a chain owner named in the approval document), written before the vendor conversation rather than discovered after it. That is the difference between a specification and a change request.
Containment. The fleet kill switch exists in the runbook and has never been tested. It is tested on a Thursday afternoon, in production, with the function head informed. It works for one agent and not the other, because a configuration difference introduced during the second deployment routes its scheduler around the control the switch acts on. Nobody knew, and nobody could have known without firing it once. Time to find: 40 minutes. To fix: two days. That one finding is the entire value of the drill discipline, and it otherwise surfaces at 2am during a real incident.
The conclusion to governance fits on one slide: two capabilities to build now (identity and containment), one to design into the next deployment (chain evidence), one to watch while building the skill in two people (oversight). Total incremental cost this year: a scope change, two days of engineering, eight hours a month of reviewer time. The point is not the plan. It is that a committee that has seen those four numbers will never again treat "we have two agents and they are well governed" as an answer to the fleet question.
The Fleet That Grew Sideways
Now the failure, whose shape is worth studying because nothing in it looks like a mistake while it happens. A logistics company deploys agents in one function and it goes well: scoped workflow, tested, monitored, measurable benefit. Encouraged, leadership lets three other functions adopt the same vendor over eighteen months. Each deployment is well-governed by the Level 3 disciplines: fit test, control spec, scoped permissions, audit logging, a named owner. Four approval packs, four sets of green metrics.
By month twenty there are fourteen agents across four functions sharing one model provider. They also share something nobody documented: a credential-provisioning shortcut, expedient at agent three when a two-week identity queue blocked a deadline, structural by agent nine when it had become how agents get provisioned here. Nobody owns the aggregate. There is no role, no forum, and no report whose scope is "all agents."
In one week of month twenty, a provider-side model update changes behaviour in eleven of the fourteen. Each monitor does its job and reads the change as a minor anomaly: a few points of drift, a small rise in exception rate, nothing that trips a threshold designed around single-process noise. Four function owners each see a local wobble, and the aggregate is invisible because no aggregate view exists. The correlated drift is identified five weeks later, not by a control but by a customer escalation that forces two functions to compare notes in a meeting called for another purpose.
Then the response takes six weeks, and the reason is the lesson. There is no fleet-level control surface, only fourteen individual ones, so every rollback, permission adjustment, and re-test is negotiated function by function, by four teams with different priorities, change windows, and vendors of record. Nothing here required a villain, an outage, or a bad decision. Every single agent was governed. The fleet never was.
Read it against Gartner's three cancellation drivers and it fits all three: escalating costs (six weeks of unbudgeted cross-functional remediation), unclear business value (four solid function-level cases and no aggregate view of what the fleet returned or cost), and inadequate risk controls (fourteen good local controls, zero global ones). Gartner's forecast that over 40 percent of agentic AI projects will be canceled by end 2027 does not only describe badly built projects. It substantially describes projects like these four, each built well.
Sequencing, and the Uncertainty You Should Not Paper Over
The four capabilities do not deserve equal urgency, and a leader who presents them as one programme will get the whole thing deferred. Rank them by three postures, and say out loud which you are taking for each.
Build now: the permission infrastructure and containment. Both are slow to build and cheap to have, and both pay off even if agent adoption at your company stalls completely, because non-human identity hygiene and tested emergency controls improve your security and audit posture regardless. That is a no-regret investment: positive value under every scenario, including the one where the optimists are wrong. Neither can be conjured in the fortnight before you need it.
Design in: the chain evidence. Do not launch a retrofit programme. Do put correlation identifiers, achievable intent capture, and a named chain owner into the requirements of the next agent deployment you approve, and every one after it. Retrofitting evidence into connected systems costs months; specifying it costs a paragraph. That differential is the largest number in this lesson, and it is determined entirely by whether you write the paragraph before or after the build.
Watch, and build the skill: the oversight architecture. This one depends on how agent capability develops, specifically on whether exception routing becomes trustworthy enough to carry primary oversight load, so committing now to an architecture built on today's calibration is a bet on a moving target. The no-regret move is smaller and human: build the pattern-review skill in two or three people, on your real data, at low cost, so that whatever architecture turns out to be right, you have the people who can operate it. Skills take longer to grow than systems take to buy.
The honest uncertainty
How far agent-shaped work spreads is not knowable, and the record supports neither the maximalist nor the dismissive position. On the expansive side, McKinsey's 2025 State of AI survey found 62 percent of organizations experimenting with or scaling agents. On the sobering side, Gartner projects over 40 percent of agentic AI projects canceled by end 2027 on escalating costs, unclear business value, and inadequate risk controls, and separately judged only around 130 of the thousands of vendors claiming agentic capability to be genuinely agentic, a phenomenon now called agent washing. Behind both sits MIT's finding that roughly 95 percent of enterprise generative AI pilots produced no measurable profit-and-loss return, with about 5 percent of custom tools crossing from pilot into production. And the regulatory clock runs regardless: general-purpose AI (GPAI) obligations since August 2, 2025, AI-content transparency December 2, 2026, high-risk Annex III December 2, 2027, embedded Annex I August 2, 2028.
That record justifies neither "agents will run operations by 2028" nor "this is a fad, ignore it." So the leader's stance is specific, and worth saying in these words at your next governance meeting: we are building the capabilities that pay off under multiple scenarios, and deliberately not building the ones that only pay off under one. That is defensible to a skeptical chief financial officer and an enthusiastic chief technology officer on the same afternoon. It also comes with a diagnostic: anyone (vendor, advisor, analyst, or internal champion) who is certain about the trajectory is telling you something reliable about themselves and nothing about the future. Certainty is not a forecast, it is a sales position.
Agentic operations is one horizon, the closest to your current decisions. The next lesson takes the wider view, out to 2030, and asks what else belongs in a readiness plan you must defend for years.
What to Do Monday Morning
Four tests and one decision, none requiring a budget line, all runnable this month.
- Compute your oversight hours per hundred agent actions, then model ten times the volume. Use measured time, not the business-case estimate, and read the resulting FTE number aloud in a governance meeting. With no agents yet, run the same arithmetic on your most heavily reviewed automated process.
- Time a scoped machine identity provisioning, and find out who recertifies the existing ones. Two emails and a calendar check. If the answers are "weeks" and "nobody," add non-human identities to the access review that already runs for humans: a scope change, not a project, and the cheapest finding you will close this quarter.
- Run the chain-reconstruction test on two connected automated steps. Pick one real case from last month, reconstruct why the second step happened, time it, and write down what was missing. That list is the requirements section for your next agent deployment.
- Test your fleet kill switch this month rather than assuming it. Schedule it, inform the function head, fire it, and document what happened to work in flight. If you have no fleet-level switch, that is the finding, better found on a Thursday afternoon than during an incident.
- Write down, on one page, which capabilities you are building now, designing in, and watching. Name the posture for each, with an owner and a review date. A leader who can say "we are deliberately watching this one, and here is what would change our minds" is doing governance. One who has not classified them is doing hope.
Key Takeaways
- Recognize that this program's disciplines assume AI that produces outputs humans dispose of, and that acting agents break three assumptions together: reviewable artifacts with human capacity, permissions designed for people with jobs, and audit trails built for decisions rather than chains.
- Treat oversight at fleet scale as architectural, not staffing: exception routing (limited by models poorly calibrated about what they do not know), portfolio-level pattern review (a skill your reviewers do not yet have), and structural constraint, where the strongest gate remains the missing key.
- Build non-human identity infrastructure as routine operation: purpose-bound time-limited credentials, delegation that carries constraints rather than expanding them, recertification at the cadence humans get, and one view of what every agent identity can reach.
- Specify correlation identifiers, achievable intent capture, and a named chain owner in the next deployment rather than retrofitting them, and tell governance that accountability for a chain crossing functions and vendors is an open design problem.
- Contain correlated action with aggregate rate and magnitude limits, fleet-level kill switches, staged rollout of behaviour changes, and drills that fire the switch, because an untested control is a design document.
- Run the four current-state tests honestly: review hours per thousand actions at ten times volume, provisioning time and recertification ownership, a timed chain reconstruction, and a five-minute fleet stop.
- Sequence by posture and name it: build permissions and containment now because they are slow and pay off under every scenario, design chain evidence into the next deployment, and watch the oversight architecture while growing pattern-review skill in a few people.
- Hold the honest line on trajectory: 62 percent of organizations are experimenting with or scaling agents, over 40 percent of agentic projects are forecast canceled by end 2027 on cost, value, and control failures, agent washing is widespread, and certainty about the outcome is information about the speaker.
Skill.re