←
AI for Government
Strategic · M25 · lesson 25 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Emerging Architectures: Agentic AI and Compound Systems
📖
now learning

Emerging Architectures: Agentic AI and Compound Systems

15 min

Raj Patel, chief technology officer of a federal benefits agency, watched a vendor demo that made the room go quiet. The salesperson typed one instruction, "process this batch of address-change requests," and an AI system read each request, looked up the resident's record, validated the new address against a postal database, updated the file, and sent a confirmation letter. No forms, no clicking, no human at any point in the sequence. A program director leaned over to Raj and whispered, "Can we have that by Friday?" Raj's honest answer was the one this lesson is about: "We can have it. The question is whether we can control it." What he had just seen was not a smarter chatbot. It was an AI agent, and agents change the governance problem entirely.

This lesson explains agentic AI and compound systems in plain terms, then focuses on the part that matters most for government: what becomes harder to control when AI can take actions on its own, and how an architect designs for that before the program director's "by Friday" turns into a headline. It covers agent frameworks, orchestration patterns, multi-agent and multi-model systems, and the specific implications of each for a public agency.

From answering questions to taking actions

A regular AI model answers. You ask it something, it produces text, a human reads it and decides what to do. The human is the hands. An AI agent is different: it is given a goal, and it takes steps on its own to reach that goal, including using tools, calling other systems, and deciding what to do next based on what it finds. The agent has hands of its own. The sharpest way to state the difference is that it does not just respond. It acts repeatedly toward a goal, and then it does it again tomorrow without anyone asking.

Laid out as a loop, the pattern is short and worth memorizing because every design question below attaches to one of its steps. A goal is set, by a person or by another system. The agent receives it. The agent plans the steps it believes will achieve it. The agent executes those steps. The agent monitors its own progress. The agent replans when something does not go as expected. And finally the agent either achieves the goal or escalates to a human. Compare that to the traditional shape, where a user provides input, the model processes it, the model returns output, and the user decides what happens next.

Three building blocks make this work. Tool use means the AI can call external systems, like a database lookup or a letter-generation service, not just produce words. Orchestration is the pattern that decides the order of steps and what happens after each one. And a multi-agent system is several specialized agents working together, one validating addresses, another updating records, a third handling correspondence, coordinated toward a shared goal. A compound system is the broader term for any AI solution stitched together from multiple models, tools, and steps rather than a single model answering alone.

Raj's mental model was a useful one. A regular model is like an intern who drafts a memo and hands it to you. An agent is like an intern you have given a corporate credit card, a list of logins, and the instruction "handle the address changes." The capability leap is real. So is the supervision problem, and the moment AI can act instead of only advise, the question stops being "is the answer good?" and becomes "what did it just do, and who approved it?"

The anatomy of an agent, component by component

Five components make up any agentic system, and each is a place where a government deployment either gets its controls or does not. Understanding them separately is what lets an architect say precisely which control is missing rather than gesturing at "oversight" in general. Vendors describe these components in different vocabulary, so translate their terms into these five before comparing products, or you will compare marketing rather than architecture.

Goal specification

What is the agent trying to achieve? This must be precise, because the agent will optimize what you said rather than what you meant. "Process benefit applications" is a bad goal: it names an activity, not an outcome, and it contains no standard the agent could fail. A workable version names the decision, the criteria, the standard, and the boundary, as in determining whether an applicant is eligible for a specific program against stated criteria, meeting a stated accuracy standard such as 95 percent on eligibility determination, and initiating payment within 48 hours where eligible. Vagueness at this step propagates through every step after it.

Planning

Given a goal, how does the agent break it into steps? For an eligibility determination the plan might be to retrieve applicant income from tax records, retrieve household composition from social security data, retrieve existing benefits from the welfare system, apply the eligibility rules, and return a decision. Two things are worth noticing about that plan. It touches four separate systems, each with its own access rules, and it ends with a returned decision that may include a confidence figure. Treat that figure carefully. A model's self-reported confidence is another generated output, not a measurement of whether it is right.

Tools

What can the agent actually act on? Typical tools are retrieving data from a database, calling an interface, sending an email, updating a record, and escalating to a human. The tool list is the single most consequential design artifact in the whole system, because it is the exhaustive set of effects the agent can have on the world. Anything not on the list is unreachable. Anything on it is reachable in combinations you did not enumerate, which is why the list should be as short as the job allows rather than as long as the platform permits.

Monitoring

How does the agent know whether it is on track, and what does it do when something goes wrong? The useful pattern is an explicit rule per failure mode rather than a general instruction to be careful. If data retrieval fails, retry three times and then escalate. If eligibility is ambiguous, escalate to a human for judgment rather than resolving the ambiguity. If new information suggests ineligibility, revise the decision rather than defending the earlier one. Each of those is a written rule with a defined trigger, which is what makes it testable before deployment and auditable after.

Human oversight

Where does a human get involved, and here it matters to be exact about what each mode does. Approval before action means a person approves before the agent sends a notification or commits a change, which prevents that action if and only if the reviewer can actually see what is being approved. Exception handling means the agent escalates ambiguous cases, which catches ambiguity the agent recognizes and nothing else. Regular audits mean sampling agent decisions, which detects patterns after the fact rather than preventing them. An appeal mechanism means affected parties can request human review, which is a remedy rather than a control.

Naming what each mode does not do is the part most designs skip, and it is the difference between a governance plan and a governance slide. None of these four prevents an agent from taking an unintended action within its permitted tool set. They gate the specific actions you chose to gate, they catch the exceptions the agent was built to recognize, and they surface in a sample what a sample can surface. Build all four, and describe them to leadership in exactly those terms, because the gap between "we have human oversight" and what the four mechanisms actually deliver is where agencies get surprised.

Compound systems and orchestration patterns

Most interesting government problems need more than one model. A benefit fraud detection system might chain four: income verification, taking reported income and tax records and returning a likelihood that the reported figure is accurate; benefit eligibility, taking income, household composition, and existing benefits and returning an eligibility determination; fraud detection, taking application patterns, historical data, and the gap between reported and apparent income and returning a probability of fraud; and risk assessment, taking all of the above plus caseworker judgment and returning a risk level. Each model's output becomes the next model's input.

How those models fit together is the orchestration pattern, and three shapes cover most designs. Sequential runs one model into the next, which suits a pipeline of dependent decisions. Parallel runs models independently and combines their results, which suits models answering genuinely different questions. Conditional routes to different models depending on an earlier result, which suits populations where different paths apply to different cases. Choosing the pattern is not a technical preference; it determines where a failure in one model can and cannot propagate, which is a governance question wearing an engineering costume.

Chaining creates a specific arithmetic problem that architects must state out loud. If one model is 90 percent accurate and a second is 90 percent accurate and you chain them, the combined accuracy is 81 percent, because 0.90 multiplied by 0.90 is 0.81. That check holds, and it assumes the two models fail independently. Where their errors are correlated, as they often are when both draw on the same underlying data quality problem, the real figure can be worse than the product suggests. The practical rule is to measure the full pipeline end to end rather than inferring pipeline performance from component performance.

Two more mechanisms complete a compound design. When models disagree, you need a stated resolution: a majority vote where two of three agree, weighted voting where models carry different weights based on historical accuracy, an ensemble that combines predictions into one, or escalation to a human when disagreement passes a threshold. And when a model fails outright, you need a fallback: use the previous model version, use a simpler and more reliable model, fall back to historical decision patterns, or escalate to a human. Decide both before launch, because the alternative is deciding them during an incident.

Multi-agent designs raise the same questions one level up. When one agent validates addresses, a second updates records, and a third handles correspondence, the coordination between them becomes its own component with its own failure modes, and the audit trail now has to span all three. The governance consequence is concrete: accountability must attach to the coordinated system rather than to each agent separately, because "the validation agent passed it" is exactly the same non-answer as "the AI did it." Assign one owner for the whole arrangement, and log the handoffs between agents as carefully as the steps inside them.

What gets harder, specifically, in government

Agentic systems are exciting precisely because they remove humans from the loop. In government, the human in the loop is often not overhead, it is due process. Removing it carelessly is where agencies get into trouble. Four things get harder, and an architect has to design for each. Note that all four are consequences of the same shift: the system now produces actions rather than recommendations, and actions have effects on people that recommendations do not.

Accountability for actions, not just outputs

When a model gives a bad answer and a human acts on it, the human is accountable. When an agent takes a bad action directly, who is? If Raj's agent updates a resident's address incorrectly and a benefits check goes to the wrong place, the agency cannot say "the AI did it." Accountability has to be assigned to a named human owner before the agent runs, not litigated after. That assignment should specify not just who is responsible when something goes wrong but who is expected to be watching while nothing appears to be going wrong, which is a different and usually vacant role.

Auditability of a chain of steps

A single model answer is easy to log. An agent that took eleven steps, called four systems, and made three decisions along the way produces a chain, and if any link is unlogged, the agency cannot reconstruct what happened. For a government action that a resident might appeal, an incomplete audit trail is a legal exposure. Every step an agent takes must be logged, with the inputs it saw and the action it took. Log the inputs specifically, not just the outputs, because "why did it do that" is almost always answered by what it was looking at rather than by what it produced.

Blast radius when something goes wrong

A bad chatbot answer affects one conversation. A misconfigured agent processing a batch can make the same mistake 10,000 times before anyone notices, because it acts at machine speed across an entire queue. The architect's job is to limit the blast radius: caps on how many actions an agent can take before a human checks, and hard limits on what kinds of actions it is allowed to take at all. Speed is what makes agents valuable and it is also the entire reason a small error becomes a large one before any human process could have intervened.

Reversibility

Some actions can be undone, like flagging a record for review. Some cannot, like sending a benefits denial letter to a resident. The single most important design rule for government agents is to separate reversible actions, which an agent may take freely, from irreversible ones, which require a human checkpoint. A weaker version of this rule circulates in practice, holding that agent actions must be reversible or auditable. Treat that "or" as insufficient for government work. An auditable irreversible action is a well-documented harm, and the resident it reached is no better off for the quality of the record.

Reversibility is also harder than it looks, because an agent that initiated downstream processes has produced effects the agent itself cannot retract. The record update may be revertible while the letter it triggered is already in the mail, the payment file is already transmitted, and the resident has already been told something. Map the downstream consequences of each tool before granting it, and treat any action whose effects leave your systems as irreversible regardless of what your own database can roll back.

One practical consequence of the reversibility rule is that it changes what you procure rather than only what you configure. A product whose architecture assumes the agent commits its own actions cannot be made to stage them by a setting, and discovering that after award leaves you choosing between a control you promised and a system you already bought. Ask vendors early and specifically whether actions can be staged for approval, which actions cannot, and what the audit record contains, and treat the answers as acceptance criteria rather than as feature discussion.

A usable artifact: the agentic AI control checklist

Before any agent goes into production touching residents or their records, run it through this checklist. It turns the four hard problems above into concrete go-or-no-go decisions. An agent that cannot pass the irreversible-action and audit rows does not launch, regardless of how impressive the demo was. The right-hand column states plainly what each control does and does not deliver, because a checklist that oversells its own rows is how a governance artifact becomes a liability.

ControlQuestionRequired before launchWhat it does not do
Named accountabilityWho owns the agent's actions?A specific human role is accountable for every action the agent takesAssigning blame afterward does not substitute for someone watching during normal operation
Action allow-listWhat may the agent do?Allowed actions are explicitly listed; anything not listed is unreachableIt bounds which tools are available, not what the agent does with the ones it has, or in what combination
Irreversible-action gateCan it do permanent harm alone?Every irreversible action such as a denial, payment, or letter requires human approvalApproval only prevents harm if the reviewer can see what is being approved; a batch too large to inspect is not a gate
Full step loggingCan we reconstruct what it did?Every step, input, and decision is logged and retained for the appeal periodA complete log nobody reads is evidence for a later investigation, not a control on today's behavior
Blast-radius capHow big can a mistake get?The agent pauses for human review after a set number of actions or on anomalyIt bounds the size of a detected failure; it does not detect failures the anomaly rule was not written to catch
Kill switchCan we stop it instantly?An operator can halt the agent immediately and roll back staged changesIt stops future actions and unwinds what is still staged; it cannot retract effects that already left your systems
Resident transparencyDoes the public know AI acted?Affected residents are told an automated system was involved and can appealAn appeal is a remedy after the fact, not a control that prevents the action

Where agents actually fit in government work

The case for agents in the public sector rests on how much government work is procedural. A great deal of a caseworker's day goes to checking eligibility, retrieving information from multiple systems, filling out forms, and sending notifications, none of which requires the judgment the caseworker was hired for. Automating that class of work lets people concentrate on complex cases, exceptions, and the situations where human judgment is the point. That is a real benefit and it is why this technology will keep arriving in demonstrations regardless of what any architecture review concludes.

Three shapes recur in proposals. Benefit processing, where the goal is to process an application and the agent retrieves data, checks eligibility, sends notifications, and updates records. Fraud investigation, where the goal is to investigate a flagged transaction and the agent gathers evidence, queries multiple databases, generates a report, and recommends an action. And screening, where the goal is to evaluate submissions against criteria and the agent reads, evaluates, schedules, and communicates. The three are not equally suitable, and sorting them is exactly the architect's job.

Apply the reversibility test and they separate cleanly. Fraud investigation is the most defensible of the three, because the agent gathers and recommends while a person decides, so its outputs remain advisory. Benefit processing splits: retrieval and eligibility checking are reversible staging work, while sending a determination is not. Screening people is the most fraught, because an agent that evaluates candidates against criteria and then schedules or declines is making consequential decisions about individuals, in a domain with well-documented bias risk. Where an agent touches decisions about people, the recommend-only pattern should be the default and any departure from it should be argued explicitly.

How the checklist changed the Friday demo

Raj did not say no to the program director. He said "yes, with the checklist." The redesigned address-change system kept most of the speed: the agent still read, validated, and staged 4,000 changes overnight, work that previously took two staff a full week. But the irreversible step, committing the change and mailing the confirmation, batched up for a supervisor who approved them in a single morning review, with the handful of anomalies, addresses that failed validation or matched fraud patterns, pulled out for individual attention. Most of the speed, with a human accountable for every permanent action, is what good agentic design in government looks like.

One caveat on that design, and it is the one most easily missed. A supervisor approving 4,000 staged changes in a morning is not inspecting 4,000 changes. What makes the gate real is the anomaly extraction that pulls the questionable cases out for individual attention, plus a sampled inspection of the routine ones, plus a stated rule for what the supervisor is actually certifying. Without those, batch approval is a signature on a number, and the control exists on the architecture diagram rather than in the process. Design the reviewer's task, not just the review step.

Adopting deliberately, not reactively

The pressure Raj felt, "can we have it by Friday," is the real risk in this space. Agentic AI is genuinely powerful, and the temptation to deploy fast is strong precisely because the demos are so convincing. The architect's contribution is not to slow innovation but to make sure the first thing built is the control layer, not the application. Start with low-stakes, reversible, internal use cases such as routine data processing, document classification, or notification sending. Prove the pattern, then extend to resident-facing actions once the logging, the kill switch, and the approval gates are demonstrably working.

Invest in the governance infrastructure early for a practical reason rather than a principled one: it is far easier to build rules in while you are building than to retrofit them into a running system whose users have already adapted to its absence. An agency that builds the brakes before it builds the engine gets to keep driving. One that does it in the other order ends up explaining to oversight why the brakes were an afterthought, usually while the engine is still running because turning it off has become politically expensive.

Anti-Patterns

  • Goal specification that is too vague. Told to "process benefit applications efficiently," an agent can reasonably interpret that as approving applications quickly, and now the agency is overpaying benefits at machine speed. Specify the decision, the criteria, the accuracy standard, and the boundary, so that the goal contains something the agent can measurably fail rather than only something it can appear to satisfy.
  • Oversight that exists on the diagram. An agentic system is built with human oversight named as a control but nothing defining what the human sees, on what cadence, or with what authority. A wrong decision then runs uncaught until hundreds of cases have been processed. Mandate approval for a defined share of decisions even where they look routine, automate escalation of ambiguous cases, and audit regularly, and be precise that these detect and gate rather than prevent.
  • Batch approval as a rubber stamp. Handing a supervisor thousands of staged actions and a single approve button produces a signature, not a review. Design the reviewer's task explicitly: what is extracted for individual attention, what fraction is sampled, and what exactly the approver is certifying. An approval gate the reviewer cannot possibly satisfy is worse than no gate, because it creates a record suggesting someone checked.
  • Treating the allow-list as containment. An action allow-list bounds which tools the agent can reach. It says nothing about the sequences and combinations it can assemble from them, and a set of individually harmless tools can compose into a harmful effect. Keep the list as short as the job allows, and reason about combinations rather than about items.
  • Trusting a self-reported confidence. An agent that returns a decision "with confidence" has generated the confidence the same way it generated the decision. Routing low-confidence cases to humans is a reasonable heuristic and a poor guarantee, because the failures that matter most are the ones the system was confidently wrong about. Calibrate against measured outcomes or treat the figure as a hint.
  • Compounding model errors unmeasured. Two models at 90 percent accuracy chained together give 81 percent, and correlated failures can make it worse. Teams that validate components and never validate the pipeline ship a system whose real accuracy nobody has measured. Test end to end, on the actual chain, with the actual data.
  • Assuming an action can be undone. The agent made a decision, it was wrong, and undoing it is now complex because downstream processes have already run. Your database rollback does not recall the letter, the payment file, or what the resident was told. Map downstream effects before granting a tool, and classify anything that leaves your systems as irreversible.
  • Calling the kill switch a safety net. Halting an agent stops what it has not done yet. It does not unwind what has already reached the world, and the interval between a failure starting and a human noticing is where the entire blast radius accumulates. Pair the kill switch with anomaly-triggered automatic pausing, so the stop does not depend on someone watching.

Practice Prompts

  • Identify candidate processes. For your agency, list the processes that are most repetitive and rule-based, mark which could plausibly be handled by an agentic system, and for each one write the specific risk if it were automated. Then sort the list by whether the agent's outputs would be advisory or consequential, and notice how the ordering changes.
  • Design a compound system. Take one use case and specify the models you would need, what each takes as input and returns as output, how they chain, which orchestration pattern applies, what happens when two disagree, and what happens when one fails outright. Compute the expected end-to-end accuracy from your component figures and say why the real figure could be lower.
  • Write the governance framework. For a proposed agentic system, answer in writing: who approves the goal, what share of decisions require human review, what the escalation paths are, and how an affected resident appeals. Then state, for each of those four, what it detects, what it prevents, and what it misses.
  • Classify every tool by reversibility. List every tool you would grant an agent. For each, trace what happens downstream when it fires, and mark it reversible or irreversible based on whether the effect stays inside systems you control. Any tool you cannot classify confidently is irreversible until proven otherwise.
  • Design the reviewer's job. Take an approval gate you would put in front of an irreversible action and specify the reviewer's actual task: what is extracted for individual attention, what fraction of the remainder is inspected, how long the review is expected to take, and what the approver is certifying. If the honest time estimate exceeds what the role can give, the design is wrong, not the reviewer.

Reflection

Work through these in writing rather than in your head. If you could delegate one job function in your agency entirely to an agentic system, what would it be, and what specifically could go wrong? How would you detect that failure, and how long would it take you to notice given the volume the agent would be processing? Would you find out from your monitoring, from a colleague, or from a resident who complained?

Then the questions that are usually left out of the business case. What would the employees affected by that automation transition into, and has anyone asked them? How would you test the system before rollout in a way that resembles real conditions rather than a demonstration? And if the system had been running for a long stretch while quietly making the same error on one category of case, what in your current design would have surfaced it?

Glossary

  • Agentic AI. An AI system that pursues a goal on its own, taking repeated actions, monitoring progress, and replanning as needed, rather than producing one response to one prompt.
  • Compound system. A system combining multiple AI models, tools, and steps that work together to solve a problem, rather than a single model answering alone.
  • Tool. Any action an agent can take on the world: a database query, an interface call, an email, a record update, an escalation. Collectively, the exhaustive set of effects the agent can have.
  • Orchestration. How a compound system coordinates its models and steps, whether sequentially, in parallel, or conditionally on an earlier result.
  • Escalation. The defined process by which ambiguous or risky decisions are routed to a human rather than resolved by the system.
  • Blast radius. The number of cases a single fault can affect before it is detected and stopped. Set by the agent's speed, its queue, and the caps you place on it.
  • Reversible action. An action whose effects remain inside systems you control and can be undone cleanly, such as flagging a record. Everything else is irreversible for design purposes.
  • Allow-list. The explicit enumeration of actions an agent may take, with everything unlisted unreachable. Bounds available tools, not the combinations assembled from them.
  • Kill switch. An operator control that halts the agent immediately and unwinds staged changes. Stops future actions; cannot retract effects that already left your systems.

Closing

Agentic AI is the next frontier for government AI adoption, and the framing that holds up is neither enthusiasm nor refusal. Done well, it removes procedural drudgery from work that needs judgment. Done poorly, it causes harm faster and at greater volume than any previous class of government software, because it acts rather than advises and it does so at machine speed across an entire queue.

The agencies that succeed at this will be the ones that build governance first and capability second. That sequence is not caution for its own sake. It is the only order that works, because control mechanisms retrofitted into a running system have to be argued against users who have already organized their work around their absence, and that argument is much harder than the one you have before launch when nothing is at stake yet.

Raj kept his agent. He also kept a named owner, a complete log, a gate on every irreversible action, an anomaly rule that pulls the questionable cases out for a human, and a supervisor whose review task was designed rather than assumed. That is not a slower version of the demo. It is the version that is still running after the attention has moved on, which is the only version worth building.

Key Takeaways

  • Agents act, models advise. An agent is given a goal and takes steps on its own using tools and other systems, then does it again tomorrow without being asked. That shift changes the entire governance problem.
  • Specify goals precisely. Name the decision, the criteria, the standard, and the boundary. A goal the agent cannot measurably fail is a goal it will satisfy in a way you did not intend.
  • The human in the loop is often due process. In government, removing the human step carelessly removes a resident's protection, not just overhead.
  • Separate reversible from irreversible actions. Let agents handle reversible work freely and gate anything permanent. "Reversible or auditable" is too weak: an auditable irreversible action is a well-documented harm.
  • Know what each oversight mode actually does. Approval gates the actions you gated, escalation catches the ambiguity the agent recognizes, sampling finds what the sample covers, and appeal is a remedy rather than a control. None of them prevents unintended action within the permitted tool set.
  • Design the reviewer's task, not just the review step. A supervisor approving thousands of staged actions is signing a number unless anomalies are extracted, a sample is genuinely inspected, and the certification is defined.
  • Log every step, with inputs. An agent's multi-step chain must be fully reconstructable or the agency cannot defend its actions on appeal, and "why did it do that" is answered by what it saw.
  • Measure the pipeline, not the parts. Two models at 90 percent chained give 81 percent, and correlated failures make it worse. Component accuracy tells you almost nothing about end-to-end accuracy.
  • Limit the blast radius and automate the stop. Cap actions before human review and pair the kill switch with anomaly-triggered pausing, because an agent can repeat a mistake thousands of times before a person notices.
  • Build the control layer first. Prove logging, gates, and the kill switch on low-stakes reversible uses before extending agents to resident-facing actions. Retrofitting governance into a running system is a much harder argument than designing it in.

Frequently Asked Questions

Is an agent just a chatbot with plugins?

The plumbing overlaps but the governance question is different in kind. A chatbot with a lookup tool still returns something a person reads and acts on, so the person remains the last step. An agent is given an outcome and decides its own sequence of steps to reach it, including which tools to use and in what order, and it commits actions along the way. The moment the system's last step is an effect on a record or a resident rather than a message to a human, you are governing actions rather than outputs.

Can we make an agent safe by restricting which tools it can use?

Restricting tools is necessary and it is not sufficient. An allow-list bounds the set of effects available to the agent, which genuinely rules out whole categories of harm, and it should be as short as the job allows. What it does not bound is what the agent does with the tools it has, or what sequences it assembles from them, and a set of individually innocuous actions can compose into a consequential one. Combine the allow-list with irreversible-action gating, blast-radius caps, and full logging rather than treating it as containment on its own.

How much human review is enough?

Set it by consequence rather than by percentage. Every irreversible action affecting a resident should pass a gate that a reviewer can actually satisfy. Reversible staging work can run freely with sampled audit. The number that matters is not the share of decisions reviewed but whether the review task is doable in the time available. A gate covering every action that a person cannot possibly inspect is weaker than a narrower gate they genuinely examine, because the broad one produces a record implying scrutiny that did not happen.

What should we log, given that logging everything gets expensive?

Log every step the agent took, the inputs it saw at each step, the action it committed, and the identity of any tool it called, and retain that for at least the period during which a resident could appeal. The input side is the part teams cut first and need most, because reconstructing why an agent behaved a certain way almost always depends on what it was looking at rather than on what it produced. If cost forces a choice, reduce retention on low-stakes reversible actions before you reduce completeness on anything that touches a person.

Where should an agency start with agents?

Start where the actions are reversible, the work is internal, and a failure would be visible rather than silent: routine data processing, document classification, internal notification routing. The purpose of a first deployment is not the productivity gain. It is to build and prove the logging, the kill switch, the anomaly rules, and the approval workflow on something where being wrong is recoverable. Once those mechanisms have survived contact with real volume, extending them to resident-facing work is an incremental step rather than a leap taken under deadline pressure.