←
AI Agent Builders & Citizen Developers
Capable · M13 · lesson 13 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Roll-Back, Pause, and Kill: Three Buttons Every Agent Needs
📖
now learning

Roll-Back, Pause, and Kill: Three Buttons Every Agent Needs

15 min

Two a.m., a Saturday in April 2026, an agentic JIRA-triage workflow we'd been quietly running for a month decided to reopen 380 closed tickets because a Confluence page it was reading as "ground truth" had been edited the previous evening with a misleading example. By the time the on-call engineer noticed at 7 a.m., the customer success team's queue dashboard looked like a denial-of-service attack. There was no kill switch. There was no rollback. The fix was a five-hour SQL script run against the JIRA database that closed all the affected tickets again, plus a Sunday morning all-hands explaining what had happened. Both of those costs — the engineering time and the trust hit — would have been zero if we'd shipped three buttons on day one. This lesson is the three buttons every operator-built agent needs before stakeholders find it doing the wrong thing at 2 a.m.: roll back, pause, kill. With the explicit triggers, the runbook stub, and the failure modes operators have actually lived through.

Why Three Buttons and Not More

Operational controls are tempting to over-engineer. We've seen workflows with seven different states (running, paused-soft, paused-hard, quarantined, maintenance-mode, observation-only, killed), and the result is that when an incident hits at 2 a.m. the on-call engineer cannot remember which state to pick. The taxonomy collapses to "the thing is doing X, what state stops X."

Three buttons is the minimum complete set. Less than three and you are missing a use case. More than three and the on-call engineer hesitates. The three:

  • Roll back — undo what the agent did recently, ideally within the last N runs or N hours. Use when the agent's outputs were wrong but recoverable.
  • Pause — stop the agent from doing more work, but keep state and queued items intact. Use when something is suspicious and you want to investigate before deciding.
  • Kill — stop the agent and stop accepting new work. Use when you know the agent is doing the wrong thing and you do not yet know why.

These map to three different operational realities. Roll back addresses the past. Pause buys time. Kill protects the present from the future. The operator who understands which to reach for under pressure has the entire decision tree memorized.

The Runbook Stub: Triggers, Decisions, Actions

Every workflow that takes external action should have a runbook stub — one page, one document, written before the workflow goes to production. Not a fifty-page incident response plan. A page that tells the on-call engineer what to do in the first five minutes after waking up to a problem.

The three explicit triggers

The runbook starts with the three triggers that should drive an operator to one of the three buttons. Each one comes with a specific threshold so the on-call engineer doesn't argue with themselves about whether this counts.

  1. Cost spike — the workflow's spend exceeds 3x its 7-day rolling average in any 1-hour window. Trigger: pause first, investigate, then either kill (if cost was caused by a runaway loop) or roll back (if cost was caused by re-processing already-handled items).
  2. Error rate spike — 5xx HTTP responses or framework-level errors exceed 5% of the workflow's runs in any 10-minute window, or 2% sustained over an hour. Trigger: pause first, identify the failure pattern, then either kill (if the agent is misusing a tool) or resume after fix.
  3. Hallucination report — any human-reported instance of the agent producing factually wrong, defamatory, or otherwise harmful output. Even one report. Trigger: pause immediately. Roll back the specific action if reversible. Kill the workflow if the cause is unknown.

These three cover roughly 90% of real incidents in operator-built workflows. The remaining 10% include security incidents (treated as kill, plus security playbook), regulatory triggers (treated as kill, plus disclosure playbook), and customer escalations that involve a single high-value customer (treated as pause-plus-review).

The decision tree

For each trigger, the runbook spells out the decision tree without the on-call engineer needing to invent it under pressure:

  1. Is the trigger active right now? Check the dashboard. If the spike has passed and metrics returned to normal in the last 5 minutes, this may have been a transient — log the observation, do not act, monitor.
  2. What is the user-facing blast radius? Count actions taken since the trigger started. If it's 50 customer emails sent, action is mandatory. If it's 2 internal Slack messages, the investigation can move slower.
  3. Is rollback safe? Can we undo each action without making things worse? Sending a "we erroneously sent the previous email — please disregard" follow-up to 50 customers is rollback. Trying to recall an emailed PDF from a customer's inbox is not.
  4. If unsafe to roll back, pause and investigate. Do not kill unless we have reason to believe the agent will continue causing harm.
  5. If known cause, fix and resume. If the cause is identified within an hour and the fix is verified on a non-production fixture, resume with monitoring tighter than usual.

Button One: Roll Back

Roll back is the most powerful and the most dangerous. Powerful because it undoes harm. Dangerous because rolling back the wrong action can introduce new harm. The discipline is: roll back only what the audit log tells you was wrong, in the order it was done.

What rollback actually means by action type

Rollback is not one operation. It is a different operation per action type, and the runbook spells out each one:

  • Email send — rollback means sending a clearly-labeled correction to the same recipient. "Apologies — our system sent you an incorrect message earlier. Please disregard the prior email; the correct information is..." Never attempt to "unsend" — Gmail's Undo Send window is 5-30 seconds, irrelevant for a 2 a.m. discovery. The correction email is the operational rollback.
  • Salesforce field update — rollback means restoring the prior field value from the audit log's action_artifact_id chain or from Salesforce's built-in field history (enabled for any field you allow agents to modify). The agent's audit row contains the previous value as part of the diff; the runbook query pulls it.
  • JIRA ticket modification — rollback means restoring the ticket state. JIRA's change history (always on for production projects) provides the prior state; the runbook script reads the history and applies the inverse.
  • Database write — rollback requires a soft-delete pattern. Never allow an agent to UPDATE or DELETE directly. The agent's writes go through a service layer that records both the new value and the prior value. Rollback is "restore prior value where written_by = agent and written_at > timestamp X."
  • External webhook fired — usually irreversible. The runbook acknowledges this and falls back to "pause and notify recipient."

The N-run rollback ceiling

The rollback button should default to "rollback the last N runs" with N preset to the smallest defensible number — 10 for most workflows, 1 for high-stakes ones. Allowing "rollback the last 24 hours" sounds powerful and is a foot-cannon. A 24-hour rollback might undo legitimate work that downstream systems have already built on; debating that under 2 a.m. pressure is how second incidents start. Default small. Use a wider rollback only after the cause is fully understood.

The rollback log

Every rollback writes its own audit row. Same schema as the original action. Trigger source = "manual_rollback." Reviewer ID = the person who pushed the button. This sounds obvious and is constantly forgotten. The rollback itself is an action; auditing it makes the timeline reconstructable.

Button Two: Pause

Pause is the most common button. Most incidents do not need rollback or kill; they need a moment for a human to look. Pause stops the agent from taking new actions while preserving everything — queued items, in-progress runs, state. Resume picks up where pause left off.

The pause mechanism

In n8n and Make, pause is implemented by a feature flag (a row in a configuration table the workflow reads at the start of every run). When the flag is true, the workflow exits at the first node with "Workflow paused by [reviewer] at [timestamp]" written to the audit log. Queued runs do not execute; they wait.

In Zapier, pause uses Zapier's built-in Zap toggle (turn the Zap off). Cleaner than a feature flag because Zapier preserves the trigger queue while the Zap is off. The Zap re-runs queued tasks when re-enabled, with a 14-day queue retention.

For custom-coded agents (Python-based with FastAPI or similar), the pause flag is typically a Redis key checked at workflow entry. Kubernetes deployments often use a ConfigMap toggle. The principle is the same: a single flag that the workflow checks before taking action.

The 'pause without losing the queue' guarantee

The hard requirement of pause is that no work is lost. If a Slack approval card was waiting for a reviewer when pause was activated, it should still be approvable when pause is lifted. If a customer ticket was being summarized when pause hit, the in-flight summary completes (or aborts cleanly, depending on the workflow's atomicity guarantee) and the next ticket waits.

This is harder than it sounds. Workflows that use external triggers (a webhook from Zendesk, a polling check of Salesforce) need to either continue receiving triggers and queue them, or pause the trigger source as well. The runbook should specify which strategy for each workflow.

The auto-resume trap

Some platforms support "pause for 60 minutes then auto-resume." Do not use this for incident response. Auto-resume after pause is the workflow-level equivalent of auto-approve on timeout — it produces the exact failure mode you paused to prevent. Pause means pause until a human decides to resume. The human-driven resume is the discipline.

Button Three: Kill

Kill is the nuclear option. Kill stops the agent, stops accepting new work, and signals to upstream systems that the agent is no longer available. Use kill when you know the agent is doing the wrong thing and you do not yet know why.

The kill mechanism

Kill is more than pause. Three properties operators need:

  1. The agent stops. Same as pause.
  2. Upstream systems are told. If the agent receives triggers from Zendesk, the kill switch disables the Zendesk webhook subscription (or unsubscribes the integration). New tickets do not queue silently waiting for a dead workflow.
  3. A notice fires. The kill writes an audit row, posts to an operations Slack channel, and ideally also opens an incident in PagerDuty or Opsgenie if your organization uses one. Kill is loud; it should never be silent.

The reason kill differs from pause is that pause is a temporary investigation tool; kill is a "this workflow may not recover today" tool. The upstream notification prevents the situation where a paused workflow accumulates 14 days of triggers in Zapier's queue and then floods downstream systems on resume.

The kill audit row

The kill action is the highest-importance audit row in the workflow's history. The row includes: reviewer_id (who killed it), kill_reason (free text required, not optional), trigger_evidence (link to dashboard at moment of kill), and downstream_notifications (list of systems notified).

The reverse of kill

To re-enable a killed workflow, the runbook requires: post-mortem document linked, fix verified on test fixtures, sign-off from at least one senior operator, and explicit re-subscription of upstream triggers. The reverse-of-kill takes longer than the original deployment. This is correct. A workflow that was killed for cause needs to demonstrate the cause is addressed before it touches production again.

Implementation: The Control Panel

Three buttons in three places, never less. The control panel should be accessible:

From Slack or Teams

A persistent message in a dedicated #agent-controls channel, with three buttons for each active workflow. When clicked, the buttons trigger the appropriate flag changes via the workflow's webhook. The bot updates the message after each action with the new state and timestamp.

Why a persistent message: at 2 a.m., the on-call engineer should not be looking for the right Slack channel or the right URL. The control surface should be one tap away on mobile.

From a web dashboard

A simple page (Retool, Internal.io, Airplane.dev, or a custom React page) with the three buttons per workflow, the current state visible, and the last 10 actions logged. The web dashboard is for during-business-hours operators who prefer a screen to mobile.

From a CLI or API

For automation and on-call escalation paths. A single command (agent-control --workflow refund-email --action pause) that flips the flag. Plus a webhook endpoint that PagerDuty or other incident management tools can call when an on-call engineer acknowledges an incident. This is the level operators usually skip; it's the one that pays off when you have a 3 a.m. cascade and the on-call engineer is on a phone in a parking lot.

The Five Times We Needed These Buttons

Stories from operator-built agents we've helped run. Each one names the trigger and which button it needed.

Story one: the Confluence ground truth (April 2026)

An agentic JIRA-triage workflow read Confluence pages as its "what does each ticket category mean" reference. Someone edited a Confluence page Saturday evening with an example that was, in context, sarcastic — and the agent treated the sarcastic example as a categorization rule. 380 closed tickets were reopened by 2 a.m. Sunday. Trigger: error rate spike (the customer success Slack channel lit up with internal "what is happening" messages). Button needed: kill (immediate stop) followed by rollback (restore the 380 ticket states from JIRA's change history). What we had at the time: nothing. The fix was a 5-hour SQL script and a Sunday all-hands. Total cost to organization: $14,000 in engineering time and an undisclosed amount in customer trust.

Story two: the cost spike that wasn't (February 2026)

A marketing-content workflow that summarizes a daily news feed spiked from $4/day to $87/day overnight. The Anthropic Console dashboard caught it. Trigger: cost spike (22x normal). Button needed: pause (investigate before deciding). What we did: paused the workflow, checked the audit log, found the cause — the upstream RSS feed had started embedding full article text instead of summaries, and the workflow was paying for 20x larger inputs. Resolution: added an input-length truncation node, verified on test fixtures, resumed. Total time from pause to resume: 45 minutes. No kill needed. No rollback needed.

Story three: the wrong field update (March 2026)

A Salesforce field-update agent was supposed to update the "next follow-up date" field. A prompt regression caused it to update the "last contact date" instead, which is a regulated audit field. 47 records were touched in 30 minutes. Trigger: a sales operations engineer noticed during her morning queue review. Button needed: pause first, then rollback (Salesforce field history made this trivial; the audit row contained the prior values). Total restoration time: 22 minutes including the 47 reverse updates. Disclosure: required to internal compliance because the audit field was touched; the audit log was the disclosure artifact.

Story four: the hallucinated security advisory (January 2026)

A customer-support workflow that drafts responses to common questions hallucinated a security recommendation that contradicted the product's actual security model. One customer received the draft and reposted it to a security forum, asking if it was correct. Trigger: hallucination report (single instance). Button needed: kill (because the cause was unknown — we did not know if other recent drafts had similar hallucinations until we audited). What we did: killed the workflow, pulled all 178 drafts sent in the prior 7 days, manually reviewed them, found 3 additional hallucinations. Resolution: 4 follow-up corrections to customers, a prompt revision, a 100-example regression eval set, and re-enable after 5 days. Cost: 12 engineering hours, no customer churn, a clear-eyed post-mortem.

Story five: the silent failure that wasn't silent (November 2025)

An n8n workflow was throwing 500 errors at 8% of requests. The dashboard caught it within an hour. Trigger: error rate spike (5xx threshold exceeded). Button needed: pause. Cause: a new external API rate limit; the workflow was getting 429s and the n8n retry was converting them into 500s after exhaustion. Fix: added a SplitInBatches node and a backoff. Total downtime: 25 minutes. The pause button is what made this 25 minutes instead of "discovered three days later by accounting."

Building the Three Buttons into an Existing Workflow

Here is the sequence we use to add the three buttons to a workflow that does not have them. Total time: 2-3 hours.

  1. 30 minutes: state table. Create a Supabase or Airtable table with one row per workflow: workflow_id, state (running/paused/killed), state_changed_at, state_changed_by, reason.
  2. 15 minutes: pause check. Add a check at the top of each workflow that reads the state table. If state != "running," exit immediately with an audit log entry.
  3. 30 minutes: control surface. Build a Slack interactive message or web page with three buttons per workflow that updates the state table.
  4. 30 minutes: rollback runbook. Document the specific rollback steps for each action type the workflow performs. For external actions (emails, webhooks), spell out whether rollback is possible and what the substitute is.
  5. 30 minutes: kill mechanism. Wire the kill button to also disable upstream triggers (webhook unsubscribes, schedule pauses) and post to the incident channel.
  6. 15 minutes: trigger thresholds. Configure dashboards or alerting for cost spike (3x rolling avg), error rate spike (5% in 10 min, 2% sustained 1h), and a manual hallucination report channel.

At hour three, you have the three buttons. You also have a workflow you can defend in front of stakeholders, security, and the on-call engineer who gets paged at 2 a.m.

The Control Panel Is the Handoff Surface

Two layers from earlier in this chapter close here. The approval card from Lesson 1 is the day-to-day verification surface. The audit log from Lesson 2 is the evidence surface. The three-button control panel from this lesson is the emergency surface. Together they form the trifecta operators need to ship agents that survive a year in production.

The control panel is also the handoff surface. When the workflow gets handed to a sustaining team — or to the same team's on-call rotation — the three buttons plus the runbook stub plus the audit dashboard equal the operator-grade documentation they need. Workflows without these three surfaces are research artifacts. Workflows with them are production systems.

The agent that does not have a kill switch is not an agent. It is a research artifact that happens to be running in production. Ship the three buttons, ship the runbook, ship the audit log — then ship the agent.

Key Takeaways

  • Three buttons, exactly three. Roll back undoes recent harm. Pause stops new work and buys time to investigate. Kill stops everything and tells upstream systems the agent is not coming back today.
  • The runbook stub fits on one page. Three explicit triggers (cost spike: 3x 7-day rolling average in 1 hour; error rate spike: 5% in 10 min or 2% sustained 1h; hallucination report: even one instance), the decision tree, and the action map per action type.
  • Rollback is a different operation per action type. Email: send a correction. Salesforce field: restore prior value from audit log. Database: soft-delete with restore. External webhook: usually irreversible, fall back to "pause and notify recipient."
  • Default rollback ceiling is small: last 10 runs for most workflows, last 1 for high-stakes. Wider rollback only after the cause is fully understood — a 24-hour rollback can undo legitimate downstream work and start second incidents.
  • Pause must preserve queued work. Pause is a feature flag (Supabase row, Redis key, ConfigMap) that the workflow checks at entry. Resume picks up where pause left off. Never use auto-resume after pause — it produces the same failure mode you paused to prevent.
  • Kill is louder than pause. Three properties: workflow stops, upstream systems are told (webhooks disabled, integrations unsubscribed), and a notice fires to an ops channel and PagerDuty or Opsgenie. The reverse-of-kill takes longer than the original deployment by design.
  • The control panel exists in three places: Slack/Teams for mobile-first 2 a.m. access, web dashboard for business-hours operators, CLI or API for automation and on-call escalation paths.
  • Every rollback, pause, and kill writes its own audit row. Trigger source = manual_rollback/pause/kill. Reviewer ID = the person who pushed the button. Reason = required free text. This makes the timeline reconstructable when you need to explain it later.
  • The 2-3 hour build sequence to add buttons to an existing workflow: state table, pause check at workflow entry, control surface, rollback runbook per action type, kill mechanism with upstream notification, trigger thresholds and alerting.
  • Without these three buttons, an agent is not production-grade. It is a research artifact that happens to be running in production. With them — plus the approval card from Lesson 1 and the audit log from Lesson 2 — the agent is operable, auditable, and recoverable.