Human-in-the-Loop Patterns for Agentic Steps
In Chapter 3.1 you built a human gate, and the thing on the reviewer's desk sat perfectly still while she judged it. A draft email is patient. A categorization waits. A number in a staging table does not mind being checked twice. Now picture the same reviewer three months later, overseeing the vendor-inquiry agent your team designed in the last lesson, and notice what has changed on the desk. Nothing is on it. The email she would have reviewed is already in the vendor's inbox. The inquiry record is already updated. The hold on the account is already placed. She is not reviewing artifacts anymore; she is living downstream of actions, and actions do not sit still. Once taken, they have happened, and the only questions left are what it costs to make them un-happen and who noticed too late. This lesson is about moving human judgment from the desk into the action sequence itself: where it inserts, what it can still undo, and how to design that placement so the humans doing it are exercising judgment rather than performing it.
From Artifacts to Actions: Why the Gate Has to Move
The Gate Spec discipline from Chapter 3.1 solved a specific problem: an AI produces an output, and a named human with defined entry criteria and a review standard decides whether that output proceeds. Everything in that design assumed one comforting property: the thing being reviewed is inert. The draft cannot send itself. The categorization cannot file itself. The gate could sit anywhere between production and consequence, because the artifact would wait.
An agent removes that property, and this is the exact trade you accepted when the previous lesson's agent-fit test told you the vendor-inquiry workflow deserved one. An agent does not just produce; it acts. It queries systems, updates records, places holds, sends messages. Each of those is not an artifact awaiting judgment but an event entering the world, and events have a physics that artifacts do not. The email is in the vendor's inbox and no review standard retrieves it. The payment is queued and the queue is draining. The design question therefore shifts. For outputs, you asked "what do we check, and to what standard?" For actions, you must ask two harder questions: where in the action sequence does human judgment insert, and what can still be undone at that point?
Get the placement wrong in either direction and you pay, and the two failure modes are not symmetrical in how they announce themselves. Under-gate the consequential actions and you get the loud failure: the agent commits something irreversible, the incident review reaches the steering committee, and the program dies in a week. Gartner's forecast that over 40 percent of agentic AI projects will be canceled by the end of 2027 names inadequate risk controls among the reasons, and "the agent did something we could not take back" is what inadequate risk control looks like on the day it matters. But over-gating fails just as surely, only quietly. Put a human approval in front of every reversible, trivial action and the agent becomes a suggestion engine with extra steps, the humans become full-time clickers, and, worst of all, you have rebuilt the accountability sink from Chapter 3.1 at industrial scale: a review step whose real function is not to catch errors but to relocate liability, staffed by people whose attention was spent long before the dangerous item arrived. Two hundred rubber stamps a day is not two hundred units of safety. It is a machine for training humans to approve.
The way out is not a single clever gate. It is a small pattern language: four oversight patterns, each matched to a class of action, and one named artifact that records the matching so it can be argued about, audited, and revised. We will build all of it around the running example: the vendor-inquiry hybrid, where an agentic lookup-and-draft loop runs inside a fixed workflow and the send sits behind a human gate. By the end of this lesson that one sentence, "the send sits behind a human gate," will have unfolded into a full oversight design.
Reversibility First: The Four Classes of Action
The previous lesson left you a rule: autonomy extends as far as reversibility. This section operationalizes it, because "reversible" is not a yes-or-no property. It is a cost. The classification question for any action an agent can take is brutally practical: what does it cost to make this not have happened? Ask that question of every action and four classes emerge. Learn them now; they are the rows of everything that follows.
| Class | Name | Definition | Examples | Cost to undo |
|---|---|---|---|---|
| R0 | Reversible-internal | Nothing outside your systems changed, and inside them the change is free to discard | Reads and lookups, drafts, internal annotations, retrieval of policy text | Effectively zero: delete the draft, ignore the note |
| R1 | Reversible-with-cost | State changed inside your systems, undo exists but consumes time or creates awkwardness | Record updates with history, items queued but not sent, holds placed on accounts | Minutes of work, an audit-trail entry, an internal explanation |
| R2 | Externally visible | Something crossed the boundary of the organization; it cannot be unsent, only corrected | Outbound messages, customer-facing status changes, published prices | A correction plus reputation cost; the counterparty saw it |
| R3 | Irreversible-consequential | Undo is not a task, it is an incident | Payments released, contracts committed, data deleted, anything with regulatory weight | Clawbacks, legal exposure, disclosure obligations, the program's credibility |
Three things about this classification deserve slow attention. First, the classes are about consequence topology, not about how smart the agent is. A brilliant agent placing an R3 action and a mediocre one placing the same action carry identical downside if the action is wrong, because the cost lives in the action, not the actor. Second, the boundary between R1 and R2 is the most important line in the whole scheme: it is the wall of the organization. Inside the wall, mistakes are operational. Outside it, they are reputational, and reputation does not have an undo button, only an apology template. Third, the same verb can land in different classes depending on its object. "Update a record" is R1 when the record has version history and R3 when the update overwrites the only copy of something regulated. You classify actions as the agent can actually perform them, in your systems, with your data, not as abstract verbs.
Run the classification as a workshop exercise: list every action the candidate agent can take, one per row, and for each one write the honest answer to "what does it cost to make this not have happened?" in dollars, minutes, or awkward phone calls. Where the room argues, you have found the interesting rows. Where the room cannot answer, you have found actions the agent should not have until someone can.
The Four Patterns: A Pattern Language for Oversight
With actions classified, oversight becomes a matching exercise: each reversibility class pairs naturally with an oversight pattern, and each pattern has mechanics, a fit, and a characteristic way of failing. There are only four. That is the good news of this lesson: the space of humans-overseeing-agent-actions, which feels infinite when you stare at a blank whiteboard, compresses into four shapes you can name in a meeting.
Pattern 1: Approve-Before-Act
The most intuitive pattern and the most abused. The agent prepares an action, presents it to a human with its evidence, and waits. Nothing happens until the human says yes. This is the artifact-gate discipline applied to intentions: instead of reviewing a thing that was made, the human reviews a thing that is about to be done.
Mechanics. The quality of this pattern lives entirely in the approval payload, and the payload is nothing more than the Handoff Contract from Chapter 3.2 written for an action instead of an output. Five fields: what will be done, to whom or what, why the agent believes it should be done, what evidence that belief rests on, and what it would cost to undo. A reviewer facing those five fields can exercise judgment in about ninety seconds. A reviewer facing a bare "Agent wants to send email. Approve?" can only gamble, and will.
Fit. R2 and R3 actions at low volume. The two conditions are joint, and the second is the one everyone forgets.
Failure mode. Approval fatigue, and it deserves numbers because it is arithmetic, not character. Run the same gate arithmetic you learned in 3.1: multiply expected approvals per day by honest minutes of judgment each, and compare the product to the attention a human can actually give. If the queue demands more genuine judgment than a person has, the person will not stop approving; they will stop judging, and the approvals will keep flowing at a pace that looks healthy on a dashboard. When approvals exceed what a person can genuinely judge, the pattern is wrong, not the person. The fix is never "review faster" and always "route better," which is Pattern 3's job.
Pattern 2: Act-Then-Review
The agent acts autonomously within declared bounds, and humans review afterward: a sample of actions, or the full action log, on a fixed cadence. No action waits for a human; every action remains visible to one.
Mechanics. This is the Verification Architecture's sampling logic from Chapter 3.1, pointed at an action log instead of an output stream. The reviewer is not re-approving individual actions after the fact, which would be theater, since they already happened. The reviewer is hunting pattern errors: a lookup that increasingly hits the wrong vendor table, an annotation format drifting, a category of record update that spiked without explanation. Sampling rates follow the same logic as output sampling: higher during the pilot, stepped down as calibration data accumulates, never to zero. And note what this review cadence actually is: it is the learning loop MIT found missing in 95 percent of stalled pilots, applied to actions. The sample is how the system improves; skip it and you have autonomy without learning, which is drift with permission.
Fit. R0 and R1 actions at volume. These are exactly the actions where per-item approval would be pure fatigue with no safety return, because the downside of a wrong one is a free or cheap undo.
Failure mode. The bounds quietly widening while the review rate stays flat. Someone adds a new lookup source, someone extends the record fields the agent may touch, and six months later the agent is acting across twice the surface with the same 5 percent sample. Bounds and review rates must move together, and making bounds impossible to widen quietly is the next lesson's subject: guardrails.
Pattern 3: Confidence-and-Stakes Routing
This is the pattern that deserves your longest attention, because it is usually the production answer. The first two patterns assign one oversight mode per action type. Pattern 3 notices that the same action type is not the same risk every time it occurs. Sending a vendor a delivery date is not the same act as sending your largest strategic supplier an answer about a disputed invoice, even though both are "send-response-to-vendor." So the routing decision happens per instance, on two dimensions at once.
Mechanics. You already own half of this machine. Chapter 3.1's confidence routing sorted outputs by how sure the system was: high-confidence items to the fast lane, low-confidence items to deeper review. Pattern 3 extends that with a second axis, stakes, because confidence measures the probability of being wrong and stakes measure the price of being wrong, and oversight must be a function of both. Stakes signals are concrete and legible to a business owner: amount thresholds (any inquiry referencing more than a set dollar figure routes up), counterparty sensitivity (strategic accounts, accounts with open disputes, accounts in a regulated category), and novelty flags (an inquiry type the agent has not seen before, or a retrieval that found no grounding passage). The output of the two axes is a routing table: confident and low-stakes goes autonomous, anything else goes to a human lane matched to why it routed there. A high-stakes item goes to approval even at high confidence. A low-confidence item goes to approval even at low stakes. The autonomous lane is reserved for the intersection, and the intersection, in most real processes, is the majority of the volume. That is what makes this pattern the economic engine of the whole design: it spends scarce human attention only where either the odds or the price of error justify it.
Fit. R2 actions at volume, and R1 actions after calibration has earned trust. It is, notice, the hybrid's hybrid: the previous lesson put an agentic step inside a workflow; this pattern puts a workflow inside the oversight of a single agentic action.
Failure mode. Thresholds set by vibes and never revisited. A routing table is a claim about where errors are cheap, and claims need evidence: the calibration data from your pilot weeks, revisited on cadence, exactly as you tuned gate sampling in 3.1.
Pattern 4: The Rollback Point
The pattern for sequences. Agents rarely take one action; they take five in a row, and the naive design gates each one, which multiplies fatigue by five. The rollback-point design instead asks: where in this sequence does the consequence actually commit? Then it lets the agent execute everything before that point autonomously, because a designed checkpoint sits just before the commit where a human can unwind the whole sequence.
Mechanics. Three tools, all old, all cheap. Staging areas: the agent completes its work into a space where nothing is live until promoted. Holds and delayed-send windows: the action is committed but with a built-in pause before it takes effect, and the fifteen-minute delayed outbox is the cheapest rollback point in business, turning an R2 send into an R1 queued-item for a quarter of an hour at zero engineering cost beyond a timer. And transactional grouping, in plain words: the sequence is all-or-nothing, so unwinding at the checkpoint reverts everything, never leaving the record updated but the email unsent, a half-state worse than either whole state.
Fit. Multi-step work with one consequential commit at the end: exactly the shape the agent-fit test favored in the previous lesson, which is not a coincidence. Work that funnels many cheap actions toward one expensive one is work where a single well-placed checkpoint buys oversight of the entire sequence.
Failure mode. The checkpoint that cannot actually unwind. The design says "a human can reverse everything at step 5," but nobody has ever pressed the button, and when someone finally does, the record update turns out to have already triggered a downstream sync. A rollback point you have not tested is a rollback intention. Test the rollback, not the intention; the pilot-testing discipline later in this level will make that a drill, not a hope.
Safety is attention budgeted, not attention demanded. Match the pattern to the reversibility class, and spend human judgment only where it can still change the outcome.
Escalation, the Andon Cord, and Who Actually Watches
The four patterns cover the actions you anticipated. Three more design elements cover the moments you did not, and programs die more often in the unanticipated moments.
The agent's own uncertainty as a routing signal
Every pattern above assumed someone classified the action in advance. But agents encounter situations their designers did not enumerate, and the meta-rule for those moments must be written down before launch: when the agent cannot classify its own next action's reversibility class, it stops. Not "proceeds cautiously," not "picks the closest match": stops, files the situation to the escalation queue with its evidence, and waits. This is the stop-when-unsure rule, and it converts the scariest category of agent behavior, confident action in unmapped territory, into a routine queue item. It costs you a few false stops a week. It buys you the property that the agent's autonomy is bounded by its own self-knowledge, which is the only boundary that travels with it into situations you never imagined.
The human interrupt
Toyota's factories run a cord (originally a physical rope, the andon cord) along the production line, and any worker, any rank, can pull it and stop the line the moment something looks wrong. The revolutionary part was never the rope. It was the culture: pulling the cord is celebrated, because a line stopped on suspicion costs minutes and a defect shipped costs everything. Your agent needs an andon cord: any reviewer, not just the owner, can freeze the agent's action queue instantly, no approval needed to stop, approval needed only to restart. And it needs the culture line said out loud by leadership, more than once: a false alarm costs us fifteen minutes and we will thank you for it. If pulling the cord requires courage, the cord does not exist. This is the 70 in BCG's 10-20-70 rule (10 percent of the effort in algorithms, 20 percent in technology and data, 70 percent in people and process) showing up in agent operations: the interrupt is a cheap feature and an expensive culture, and only the culture makes it real.
Who staffs this
Every pattern above consumes human attention, and attention comes from named people with other jobs. Price the oversight roles exactly as you priced gate staffing: hours per week, on whose calendar, backed by which skills from the Level 2 skills matrix. Agent oversight is a skilled role, not a leftover duty for whoever is least busy: the act-then-review reviewer needs pattern-reading skill across logs; the approval-queue reviewer needs domain judgment and the standing to say no; the andon-cord holders need enough system understanding to know what "looks wrong" looks like. Name them in the RACI (responsible, accountable, consulted, informed) for the agent, and when the roster page is blank, the honest reading is that your oversight design is a diagram, not a control.
The Action Oversight Matrix: The Vendor-Inquiry Hybrid, In Full
Now the named artifact, built end to end on the running example. The Action Oversight Matrix is a single table: every action type the agent can take, its reversibility class, its blast radius in plain words, and the oversight pattern assigned to it, with the pattern's parameters. It is the successor to the Gate Spec for the agentic era: where the Gate Spec defined one gate deeply, the Matrix maps the whole action surface and shows that every action, without exception, has a designed oversight answer. All numbers below are illustrative, from a hypothetical mid-market pilot, but the shape is the deliverable.
| # | Action type | Class | Blast radius | Pattern and parameters |
|---|---|---|---|---|
| 1 | Look up vendor master record | R0 | None: read only | Act-then-review; 5% weekly log sample |
| 2 | Look up order and invoice status | R0 | None: read only | Act-then-review; 5% weekly log sample |
| 3 | Retrieve grounding passages from policy and SOP library | R0 | None: read only | Act-then-review; retrieval misses flagged automatically |
| 4 | Draft response to vendor | R0 | None until sent | Act-then-review; drafts sampled with their sends |
| 5 | Add internal annotation to inquiry thread | R0 | Internal readers only | Act-then-review; 5% sample |
| 6 | Update inquiry record (status, category, linked orders) | R1 | Internal; full version history, one-click revert | Act-then-review; 10% sample, drift check on category distribution |
| 7 | Place temporary hold on a vendor account | R1 | Internal, but visible friction for the vendor if wrong | Approve-before-act during pilot; moves to confidence-and-stakes routing after week 6 calibration review |
| 8 | Send response to vendor | R2 | Crosses the wall: vendor sees it | Confidence-and-stakes routing; see routing table below. 15-minute delayed outbox on every send, no exceptions |
| 9 | Anything touching payment terms, amounts, or schedules | R3 | Money and contract exposure | Agent may draft, never commit: the commit action is not in its permission set at all |
Row 8 carries the routing table, and this is where the previous lesson's sentence "the send sits behind a human gate" becomes engineering. A response auto-sends only when all three conditions hold: the inquiry type is on the calibrated routine list (invoice status, delivery dates, documentation requests), every factual claim in the draft is grounded in a retrieved passage with retrieval confidence above the calibrated threshold, and no stakes flag is raised (no amount above $10,000 referenced, no strategic or disputed account, no novelty flag). Fail any condition and the item enters the approval queue with its five-field payload. And every send, even the auto-sends, passes through the fifteen-minute delayed outbox, which means that for fifteen minutes an R2 action is still an R1 action, and the andon cord can catch it.
Row 9 deserves a sentence of respect, because it introduces the idea the next lesson is built on. Rows 1 through 8 are oversight: humans deciding about actions the agent can take. Row 9 is a different species: an action the agent cannot take, regardless of confidence, stakes, or anyone's approval mood on a Friday afternoon, because the permission simply is not there. Patterns govern the possible; permissions define it. Hold that distinction; it is the bridge you will cross next.
Now the arithmetic that proves the design is sustainable, because a matrix that overloads its humans is a fatigue machine with better formatting. In this illustration the agent takes roughly 310 actions on a typical day. About 82 percent land in autonomous lanes: the R0 and R1 act-then-review rows plus auto-qualified sends. The approval queue receives on average 11 items per day: routed sends that failed a condition, plus account holds during pilot. At an honest three minutes of judgment each, that is about 33 minutes of genuine review on the owner's calendar, against a weekly log-sampling session of about an hour. Compare that against the fatigue threshold you computed with gate arithmetic: 11 real decisions a day is a job a human can do with full attention indefinitely. That number, not the autonomy percentage, is the health metric of the whole design.
The Rubber-Stamp Factory: A Failure Story
Here is the same design question answered with caution instead of classification, assembled from real failure patterns with illustrative numbers. A SaaS company launches a support agent that can look up accounts, draft replies, issue refunds within policy, and update subscriptions. Legal is nervous, so the launch decision is the one that feels safest: approve-before-act on everything. Every lookup summary, every draft, every routine reply queues for human sign-off. Nobody classifies anything, because caution feels like a policy. It is actually the absence of one: a single pattern applied to four classes of action is not a design, it is a shrug with a checkbox.
The queue opens at about 340 approvals per day across two staffers. Week one, they are diligent: median time per approval is around 90 seconds, and they catch several small errors, which the launch report celebrates. By week three the median judgment time has fallen to 7 seconds. Not because the reviewers got faster at judging; because they stopped judging. Three hundred forty times a day, the correct answer was "approve," and human beings are learning machines: the queue has trained them, with thousands of repetitions and near-perfect consistency, that approval is the job. Then, in week five, a genuinely bad action arrives: a refund at twenty times policy, generated from a misparsed email thread where the agent read a quoted annual contract value as the refundable amount. It enters the queue at 4:52 pm, item 314 of the day, visually identical to the 313 routine items before it. It sails through in 6 seconds.
The postmortem's first draft blames human error, and the reviewer is quietly moved off the queue. Then someone finally computes the arithmetic: 340 approvals at even 60 honest seconds each is over five and a half hours of sustained judgment per person per day, which no one can deliver, which means the system was designed to be rubber-stamped and then staffed with scapegoats. Over-gating did not add safety. It manufactured the rubber stamp, then hid an R3-grade action inside a queue of R0 confirmations. An Action Oversight Matrix would have auto-sent the 300 trivial actions under act-then-review sampling, routed refunds by amount thresholds under confidence-and-stakes rules, and given those two humans eleven decisions a day worth ninety seconds each, and the 20x refund would have arrived flagged, alone, and impossible to miss. The lesson compresses to the line this whole design stands on: safety is attention budgeted, not attention demanded. And notice what the failure preserved from Chapter 3.1: the accountability sink, scaled. The approval queue's real function had become relocating liability onto whoever clicked last.
What to Do Monday Morning
This lesson becomes real when your candidate agent's action surface is on one page with no blank cells. The sequence:
- List every action your candidate agent can take, as concrete verbs against concrete systems, and classify each one R0 to R3 by asking "what does it cost to make this not have happened?" Where the team argues about a class, write the argument down; it is a risk decision, and risk decisions deserve minutes.
- Match each class to its pattern and fill the Action Oversight Matrix: act-then-review with a sampling rate for R0/R1, confidence-and-stakes routing with explicit thresholds for R2, approve-before-act or removal from the permission set for R3.
- Compute the approval-queue arithmetic: expected items per day times honest minutes of judgment each, against the attention your named reviewers actually have. If the product exceeds roughly an hour a day per person, fix the routing, not the people.
- Add the delayed-send outbox to everything R2. Fifteen minutes is usually free, and it converts your most common external action into a rollback point. Then test one rollback for real: unwind a staged sequence and verify nothing half-commits.
- Write the stop-when-unsure rule into the agent's operating document: when it cannot classify its next action's reversibility class, it stops and escalates with evidence. One sentence, written before launch, not after the first surprise.
- Name the andon-cord holders: who can freeze the agent's queue, how, and say the culture line out loud in the kickoff: pulling it on suspicion is celebrated, and a false alarm costs fifteen minutes we are glad to pay.
Key Takeaways
- Distinguish outputs from actions: the Chapter 3.1 gate reviewed artifacts that sit still, while agent oversight governs events that have already entered the world, so the design question becomes where judgment inserts and what can still be undone.
- Classify every agent action by the cost to make it not have happened: R0 reversible-internal, R1 reversible-with-cost, R2 externally visible, R3 irreversible-consequential, and treat the R1/R2 line, the wall of the organization, as the most important boundary in the scheme.
- Match patterns to classes, not caution to everything: approve-before-act for low-volume R2/R3, act-then-review with sampled logs for R0/R1 at volume, confidence-and-stakes routing as the production answer for high-volume R2, and rollback points for sequences that funnel into one consequential commit.
- Run the approval arithmetic before launch: when approvals exceed what a person can genuinely judge, the pattern is wrong, not the person, and over-gating manufactures the rubber stamp that lets the one bad action sail through at 4:52 pm.
- Build the fifteen-minute delayed outbox onto every externally visible action: it is the cheapest rollback point in business, turning R2 into R1 for long enough for a human or an andon pull to catch it, and test the rollback rather than trusting the intention.
- Write the stop-when-unsure meta-rule into the agent's charter: when it cannot classify its own next action's reversibility class, it stops and escalates, so its autonomy stays bounded by its self-knowledge even in unmapped situations.
- Install the andon cord with its culture, not just its button: anyone can freeze the queue, pulling it is celebrated, and leadership says so out loud, because per BCG's 10-20-70 the interrupt is a cheap feature and an expensive habit.
- Staff oversight as a skilled, priced role from the Level 2 skills matrix with named owners in the RACI, and carry the Action Oversight Matrix forward: patterns decide when humans judge, and the next lesson's guardrails decide what the agent cannot do regardless.
Skill.re