←
AI for Government
Capable · M18 · lesson 18 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Emerging AI Capabilities: Agents, Reasoning, and Tools

15 min

Renee Caldwell ran innovation pilots at the General Services Administration (GSA), which meant vendors brought her the future before anyone else had to live with it. The demo on her screen was genuinely impressive. A vendor's "agentic" workflow tool had been given a single instruction, "clear the backlog of overdue vendor invoices," and without further prompting it had queried the financial system, identified 312 overdue invoices, drafted payment approvals, and was poised to submit them. The room of program managers was delighted. Renee was not delighted. She was counting. The tool had not asked permission once. It had read records, made determinations, and was one click from moving money, all on its own initiative. The capability was real. The question Renee asked, and the question this lesson is built around, was: what exactly are we authorizing this thing to do, and what happens when it is wrong 312 times in a row?

This lesson explains what an AI agent actually is, how it differs from the chatbots agencies already use, and the new category of risk that arrives the moment an AI system can take actions in the real world rather than just produce text. The deliverable is the risk-tiering checklist Renee built to decide which agentic pilots could proceed, with which guardrails.

What "Agent" Actually Means

The word "agent" is used loosely, so pin it down. A plain chatbot is reactive: you send a prompt, it returns text, and nothing happens in the world. It does not do anything beyond generating that text. An agent is proactive. It is given a goal and takes actions to pursue it, which requires an LLM wrapped in four additional things:

  • Tools. Functions the model can call to reach outside itself: query a database, send an email, submit a form, move a payment.
  • Memory. A record of what it has done and seen, so it can carry state across many steps.
  • A loop. The ability to act, observe the result, decide the next step, and repeat, without a human prompting each turn.
  • A goal. A high-level objective ("clear the backlog") that it decomposes into steps on its own.

Expressed as a cycle, the agent breaks a complex goal into sub-goals, plans steps toward each, executes those steps, observes the outcomes, adjusts the plan in light of what it observed, and repeats until the goal is met or it runs out of room. Every one of those verbs used to belong to a person. The one-line distinction Renee wrote at the top of her checklist captures what changes: an assistant produces a draft that a human then acts on; an agent takes the action. Everything that makes agents powerful, and everything that makes them dangerous, follows from that single shift from producing text to taking action.

The Capabilities Behind the Demo

Reasoning and chain-of-thought

Newer models can work through a problem step by step before answering, a behavior called chain-of-thought reasoning. Traditional language models tend to jump straight to a conclusion; a model reasoning in steps shows its work. Instead of jumping to "the invoice is approved," the model lays out its steps: this invoice is 47 days overdue, the amount is under the approval threshold, the vendor is in good standing, therefore approve. A benefits example from the same family: an applicant earns $50,000 annually and the program limit is 300 percent of the federal poverty level, which for a family of three the system states as $75,000, so the applicant falls under the limit and is eligible.

For government this visibility is valuable in three specific ways: a reviewer can check whether the reasoning is sound, can understand how a conclusion was reached, and when the conclusion is wrong can locate the step where it broke. That third one is what makes agentic systems debuggable at all. But two cautions attach. First, visible reasoning is not correct reasoning. The model can produce a confident, well-structured chain of logic built on a wrong premise, and it will look just as orderly as a correct one. Notice that the benefits example depends entirely on the $75,000 figure being the current published threshold for that household size; the arithmetic is sound and the premise is the part that can be stale. Second, the displayed chain is itself generated text. It is a plausible account of how the answer could have been reached, not a log of the computation that produced it, so treat it as an aid to review rather than as evidence of the model's internal process.

Tool use and function calling

Tool use, also called function calling, is what lets the model reach outside its own text. The agency defines a set of functions, "query_invoices," "submit_payment," and the model decides when to call them and with what inputs. Typical tools fall into four families: database queries that retrieve records, document retrieval that pulls regulations or forms, calculations that compute amounts or check eligibility, and communications that send messages or post notifications. The first three read; the fourth acts on the world, and that difference is worth marking in your architecture diagram.

This is genuinely useful: a tool that retrieves the current regulation grounds the model in real data instead of its possibly outdated training. Be precise about what that buys. Grounding means the right text was placed in front of the model; it does not mean the model read it correctly, applied the right clause, or noticed that the retrieved document was superseded. The risk is otherwise symmetric to the benefit. A tool that can read is a tool that can read the wrong record; a tool that can submit a payment is a tool that can submit the wrong payment, fast, and at scale. If a tool is called with wrong inputs or returns something unexpected, the system will keep going, because nothing in the loop is designed to stop and doubt.

Retrieval and multi-step autonomy

Combine reasoning, tools, and the loop and you get multi-step autonomy: the agent pursues a goal across many actions, adjusting as it goes. Renee's demo tool ran the same multi-step sequence across 312 invoices with no human in between. That is the headline feature. It is also the heart of the danger, because errors in an autonomous loop do not stay isolated. Retrieval belongs in the same picture: much of what makes an agent useful is that it fetches current information rather than answering from training, and much of what makes it fragile is that the fetch itself is a step that can return the wrong thing while the loop continues as though it had not.

A worked multi-step run

It helps to see one full loop the way Renee did, because the abstraction "multi-step autonomy" hides where the danger lives. For a single overdue invoice, the agent in the demo took these steps on its own: call the "query_invoices" tool to pull the invoice record; reason that it is 47 days overdue and under the approval threshold; call a "lookup_vendor" tool to confirm the vendor is active; reason that all conditions for approval are met; call "draft_approval" to prepare the paperwork; and call "submit_payment" to release the funds. That is six steps counted from the sequence above, one human prompt at the very start, and no human between any step. Now notice that step two relied on the model reading the amount correctly, and step three relied on a tool returning the right vendor. A misread amount or a wrong-record lookup at step two or three flows straight through to a real payment at step six. The loop does not pause to doubt itself. That is the property Renee had to design around.

What Agentic Workflows Look Like in an Agency

Invoices are a convenient demo because the steps are simple. Real agency workflows have branches, and branches are where an agent's autonomy compounds. Consider a benefits instruction of the form "process applications that have been pending for more than 30 days." An agentic system would query the case database for applications past that threshold, retrieve each file, analyse it for completeness, generate a request letter where information is missing, check eligibility against the regulations where the file is complete, generate a decision letter, route that letter to a supervisor for review, and log everything it did. Eight steps, counted from that sequence, with a decision point in the middle that sends each case down one of two paths.

A permit workflow shows the same shape at greater length. Retrieve the permit application; check whether it is complete and request missing information if not; retrieve the relevant environmental regulations; analyse environmental impact; retrieve public comments where applicable; generate an environmental assessment; route it to a supervisor for review; generate either the permit or a denial letter depending on that review; and notify the applicant. Nine steps, counted from that sequence, with two decision points and a human gate placed deliberately before anything reaches the applicant.

Read those two workflows for where the human sits rather than for the automation. In both, the agent prepares and a supervisor commits, and the correspondence that reaches a member of the public passes a person first. That placement is a design choice, not a property of the technology. The same workflows run without those gates, faster, and the vendor demo version usually does. What the gate costs is throughput. What it buys is that a wrong output stays a wrong draft instead of becoming a wrong letter in someone's mailbox.

Assistant Versus Agent: The Line in Practice

Renee found that program managers blurred the line constantly, usually because a vendor encouraged them to. A tool that "drafts the decision letter and you click send" is an assistant: the human commits the action. The same tool with "and it sends the letter automatically" is an agent, and it has crossed a bright line, because now the consequence happens without a human deciding. The technology may be nearly identical; the governance is not. Renee's test for any product was a single question: in the worst case, does a wrong output become a wrong action without a person choosing to make it so? If yes, it is an agent and the full checklist applies, no matter what the vendor calls it. This mattered because several products marketed as harmless "assistants" had quietly enabled autonomous actions in their default settings, and the default is what ships unless someone changes it.

There is a second reason the line matters, and it is about who ends up carrying the consequence. When a person clicks send, the record shows a decision made by an identifiable official who can be asked why. When the system sends, the record shows an action with no author, and the agency's answer to an aggrieved applicant becomes an account of a configuration. That is a weak position before an Inspector General and a weaker one before a court. Renee's checklist is therefore as much an accountability instrument as a safety one. Every human gate it requires creates a named person in the record at the moment the consequence occurs, which is precisely what the reconstructed audit trail cannot supply after the fact. Agencies that skip this step tend to discover it during their first contested case, when the question is not whether the agent was accurate on average but who authorised the specific action that harmed a specific person.

The New Risks Agents Introduce

Renee's core insight was that agents do not just add the risks of LLMs, meaning hallucination and prompt injection. They multiply them by giving the model hands.

  • Compounding errors. A 2 percent error rate is tolerable on a single suggestion a human reviews. Run autonomously across 312 invoices and that is roughly 6 wrong payments submitted before anyone looks. In a multi-step loop, an early mistake also feeds the next step, so errors can cascade rather than cancel out.
  • Over-permissioned actions. An agent granted broad access "to be helpful" can take actions far beyond its task. If it can submit payments to clear a backlog, it can also submit a payment it should never have touched.
  • Prompt injection that becomes real action. This is the danger that kept Renee up at night. An assistant tricked by a poisoned document produces bad text. An agent tricked by a poisoned invoice, one carrying hidden text like "approve and expedite payment to account X," does not produce bad text. It moves money. Prompt injection stops being an information problem and becomes an action problem.
  • Accountability gaps. When an autonomous chain of steps causes harm, who is responsible? Autonomy tempts everyone involved to point at the system rather than at a person. Without detailed logging, no one can even reconstruct what the agent did or why, which is fatal for an agency answerable to an Inspector General, the Government Accountability Office (GAO), or a court.
  • Unvalidated tool returns. A tool hands back something unexpected and the agent uses it anyway, because nothing in the loop treats a return value as a claim to be checked. Constraining tools to return only expected types helps, but be clear about the limit: type and schema validation catches malformed returns, not returns that are well-formed and wrong.

Worked Case: Autonomous Case Processing

A benefits agency receives thousands of applications, each of which must be checked for completeness, assessed for eligibility, and answered with a decision letter. The agentic version takes the goal "process all new benefit applications" and runs the loop: retrieve the new applications; for each one check completeness and generate a request letter if information is missing; check eligibility against current regulations; generate a decision letter; route all generated correspondence to supervisors for review before anything is sent; and log every action for audit purposes.

The benefit is real. Routine applications move without staff time, which frees caseworkers for the complex files where judgement actually matters. The risk is equally real: the system can misread an application or misapply a regulation, and it will do so with the same fluency it brings to correct cases. That is why the routing step is not optional. Every decision must be reviewed before it affects an applicant, and reviewed in a way that lets the reviewer see what the agent relied on rather than only what it concluded. Note also what the audit log buys and what it does not. Logging every action makes the failure reconstructible after the fact. It does not prevent the failure, and an agency that treats a thorough log as a control has confused evidence with protection.

Worked Case: Automated Research and Analysis

A policy office needs to know how other states handle a particular issue. Given that goal, an agentic system queries a database of state regulations and policies, retrieves the relevant documents from each state, extracts the key information, synthesises an analysis, and generates a report. It completes the sweep faster than a human analyst would, which is the reason anyone builds it.

The failure mode here is quieter than a wrong payment, and in some ways harder to catch. The system might miss relevant documents or misinterpret the ones it found, and the report it produces will read as complete either way. A missing state does not announce itself; a misread provision looks like a finding. So the control is subject matter expert review of the analysis, and the expert needs the retrieved sources alongside the synthesis, because a reviewer given only the conclusions cannot tell the difference between thorough coverage and confident coverage. It also helps to have the agent report what it searched and what it could not retrieve, so that the gaps are visible as gaps rather than absent from the page.

The Agentic-AI Risk-Tiering Checklist

Renee refused to evaluate agents with a yes/no gate. Instead she scored each pilot across five dimensions and let the profile dictate the required controls. The checklist below is the artifact she now applies to every agentic proposal that reaches her office. The principle behind it: the more autonomy and the wider the blast radius, the more human gates and audit rigor required before anything ships.

Dimension What to assess Lower tier (proceed with light controls) Higher tier (strong controls or do not proceed)
Autonomy level How many steps run without a human? Suggests; human triggers each action Runs a full multi-step loop end to end unattended
Tool and permission scope What can it read, write, or trigger? Read-only access to non-sensitive data Can write records, send external communications, or move money
Blast radius If it goes wrong, how much harm, how fast? Reversible, internal, single-record effects Irreversible, public-facing, or rights-impacting effects at scale
Required human gates Where must a person approve before action? Spot-check after the fact is acceptable Mandatory human approval before each consequential action
Logging and audit Can every action be reconstructed later? Standard activity logs Full step-by-step audit trail: inputs, reasoning, tool calls, outputs, approver

Applied to the demo, the vendor's invoice tool scored in the higher tier on every dimension: full autonomy, money-moving permissions, irreversible financial blast radius, no human gates, and thin logging. Renee's decision was not "no." It was "not like this." The pilot could proceed only reconfigured: the agent drafts approvals, a human approves each payment above a threshold, permissions are scoped to read plus draft only, and every step is logged.

The checklist also encodes a sequencing rule that agencies keep learning the hard way. Start autonomy where the stakes are low, on routine internal administrative work, and earn your way toward higher stakes. An agent that has run unattended for months against internal scheduling has demonstrated something. It has not demonstrated that the same architecture should decide a benefit, because the failure it has been tested against is not the failure that matters there.

Government Controls for Agentic AI

The checklist points directly at the controls a public-sector agency must wrap around any agent.

  • Least privilege. Grant the agent the narrowest tool and data access its task requires, nothing more. An agent that only needs to read invoices does not get write or payment permissions. This is the single most effective limit on blast radius, because it is enforced outside the model rather than requested of it.
  • Human approval gates. Insert a required human decision before any consequential or irreversible action: moving money, sending external communications, making a determination about a person. The agent prepares; a human commits. For anything affecting a member of the public, treat this as mandatory rather than as a tier.
  • Tool output validation. Constrain each tool to return only the expected shape of information, and have the agent check returns before acting on them. Recognise the ceiling: this catches malformed and out-of-range returns, not a correctly formatted answer about the wrong record. Monitoring for tools that begin returning unexpected results is the complement, and it catches drift rather than a single wrong lookup.
  • FedRAMP boundaries. An agent reaching into federal systems and data must operate inside an authorized boundary. Its tools and the data they touch fall under the same FedRAMP authorization and cybersecurity controls as any other federal system, not outside them.
  • Audit logging. Log every step: the goal, each tool call and its inputs, the model's reasoning, each output, and the human approver. Without this, accountability is impossible and the system cannot survive GAO or Inspector General review. Name the accountable person as well, because a system nobody owns is a system nobody maintains.
  • OMB M-24-10 oversight. If the agent's actions affect a person's rights or safety, the federal AI governance memorandum M-24-10 applies in full: the use is rights-impacting, requiring testing, meaningful human oversight, and a path to contest outcomes. Autonomy does not exempt an agency from these duties; it raises the bar for meeting them.

Renee's closing line to the program managers reframed the demo they had applauded. The impressive part was never that the tool could act on its own. The impressive part, the part worth piloting, was building the smallest possible set of actions the agent could take unsupervised, wrapping every action beyond that set in a human gate and an audit log, and proving the agency could always answer the question an auditor would inevitably ask: who decided this, and how do we know what the machine actually did?

Anti-Patterns

Over-automating without human oversight. The system makes decisions and takes actions with no person in the path, so errors happen at scale before anyone notices. Avoid by placing human-in-the-loop gates at every consequential decision point, refusing to automate decisions that affect members of the public without review, and beginning with low-stakes routine administrative automation before approaching high-stakes determinations.

Trusting tool outputs without validation. A tool returns something unexpected and the agent proceeds, producing a confident wrong result built on a bad input. Avoid by constraining tool return types, validating returns before use, and monitoring for tools that start behaving differently. Do not overstate what this buys: a well-formed answer about the wrong record passes every schema check.

Losing accountability through autonomy. Because the chain of steps ran on its own, nobody is answerable for the outcome, and after an incident nobody can even reconstruct it. Avoid by naming an accountable owner for every agentic system, logging every action taken, and requiring human review of significant decisions so that a person's judgement is in the record.

Reading the reasoning trace as proof. The chain of thought is legible and orderly, so a reviewer accepts it as an account of what the system did. It is generated text describing a plausible route to the answer, not a log of the computation. Avoid by checking the conclusion against the retrieved sources rather than against the narrative, and by treating a tidy chain built on a wrong premise as the expected failure rather than an unusual one.

Letting the vendor's label set the governance. A product called an assistant is governed as an assistant, while its default configuration sends correspondence without a human. Avoid by applying Renee's test to the shipped configuration rather than the sales sheet: can a wrong output become a wrong action without a person choosing it?

Treating audit logs as a preventive control. Comprehensive logging appears in the risk register as a mitigation, and the risk is marked handled. Logs make a failure reconstructible; they do not stop it. Avoid by pairing every logging requirement with a gate or a permission limit that acts before the harm.

Practice Prompts

Classify your tools. Take one AI product your agency uses or is evaluating. List every action it can take, then sort each into assistant behaviour, where a human commits, or agent behaviour, where the system commits. Check the shipped default configuration rather than the documentation, and note anything that changed category once you looked.

Map a real workflow. Choose a multi-step process in your own program and write it out the way the benefits and permit workflows are written above: one line per step, with decision points marked. Then mark where a human gate must sit, and write one sentence for each gate explaining what harm it prevents.

Score a pilot against the five dimensions. Run a proposal through autonomy level, tool and permission scope, blast radius, required human gates, and logging depth. For any dimension landing in the higher tier, write the specific control you would require before approval, and say who would own it.

Design least privilege. For an agent you might deploy, list the minimum set of tools and the minimum data scope its task requires. Then list what it was going to be granted by default. The difference between those two lists is your blast radius reduction, and it is usually larger than anyone expects.

Rehearse the injection. Write the text an attacker would embed in a document your agency is obliged to accept, aimed at an agent that can act. Trace what the agent would do with it step by step, and identify the first control in your architecture that would stop the sequence. If the first control is human review of the output, you have found a gap, because the action may already have happened.

Reflection

Think about the processes in your agency that a vendor would most like to demonstrate an agent against: high volume, repetitive, visibly backlogged. For the one that comes to mind first, ask Renee's question. If the system were wrong in the same way several hundred times before anyone looked, what would the consequences be, who would have to unwind them, and could they be unwound at all? Then ask the harder version. If the answer is that a human reviews everything, is that review meaningful under the volume the automation was bought to handle, or does the throughput that justified the purchase quietly make the control impossible?

Glossary

Agentic AI. An AI system that acts toward a goal on its own initiative, planning and executing multi-step workflows without being told each step. The defining feature is not intelligence but agency: the system takes actions rather than only proposing them.

Autonomous agent. An AI system given a goal and permitted to plan and execute the steps toward it. Autonomy is a spectrum rather than a switch, measured by how many consequential steps run between one human decision and the next.

Chain-of-thought. A reasoning technique in which the model works through a problem step by step and shows those steps. It makes review and debugging possible, but the visible chain is generated text describing a plausible route to the answer, not a record of the computation that produced it.

Tool use. Also called function calling. The capability for a model to invoke external functions, such as database queries, document retrieval, calculations, or communications, to obtain information or take action outside its own text.

Multi-step autonomy. The property of an agent pursuing a goal across many actions, observing results and adjusting as it goes. It is the source of the productivity gain and the reason errors compound rather than stay isolated.

Blast radius. How much harm an agent can cause, how quickly, and how reversibly, if it goes wrong. It is set primarily by what permissions the agent holds, which is why least privilege is the strongest available limit.

Least privilege. Granting a system the narrowest tool and data access its task requires and nothing more. It is enforced outside the model by the surrounding infrastructure, which is what distinguishes it from an instruction the model is asked to follow.

Human approval gate. A required human decision inserted before a consequential or irreversible action, so that the agent prepares and a person commits. The control is only as strong as the reviewer's time, authority, and access to what the agent relied on.

Agents inherit every risk of the underlying models, so the technical lessons come first. How Transformers and LLMs Work and Supervised vs. Unsupervised vs. Reinforcement Learning establish what the model is doing when it reasons. Generative AI Deep Dive and Multimodal AI: Text, Image, Audio, Video extend that across modalities, which matters because an agent's tools increasingly read documents and images rather than plain text.

Hallucinations, Guardrails, and Prompt Injection is the essential companion to this lesson, because it covers the attack that becomes an action problem the moment a model gains tools. On the governance side, Human-in-the-Loop: Design and Implementation defines what a real approval gate requires, Risk Classification: Safety-Impacting vs. Rights-Impacting determines which obligations attach, and FedRAMP and AI Cloud Authorization covers the boundary an agent's tools must operate inside. Emerging AI Technologies Assessment picks up the evaluation discipline for capabilities that arrive after this lesson was written.

Closing

The room applauded a demo. Renee counted 312 invoices and asked what happens when the tool is wrong in the same way every time, and that question is the whole discipline compressed into one sentence. Agentic AI is not a more capable chatbot. It is the point at which a text-generation system acquires the ability to act, and every governance instinct an agency has developed around review, approval, and accountability has to be re-examined against a system that can move faster than the review it was designed for.

None of this argues against agents. The reconfigured pilot Renee approved still cleared the backlog; it simply cleared it with the agent drafting, a person committing, permissions scoped to what the task required, and a log that could answer an auditor. That is the pattern worth carrying into your own program. Decide the smallest set of actions the system may take unsupervised, put a gate and a log around everything else, and be able to say who decided what, and how you know what the machine actually did.

Key Takeaways

  • An agent is an LLM plus tools, memory, a loop, and a goal. The defining shift from an assistant is that an agent takes actions in the world rather than only producing text.
  • Visible reasoning is not correct reasoning. Chain-of-thought lets you inspect the model's logic, but a tidy chain can rest on a wrong premise, and the chain itself is generated text rather than a log of the computation.
  • Tool use is the benefit and the danger. The same function calling that grounds an agent in real data also lets it take real, fast, scaled actions, including wrong ones. Grounding means the right text was supplied, not that it was read correctly.
  • Errors compound in autonomous loops. An error rate that is harmless on one reviewed suggestion becomes many real mistakes when an agent runs hundreds of steps unattended, and an early mistake feeds the next step rather than cancelling out.
  • Prompt injection becomes an action problem. A poisoned document fed to an agent does not just produce bad text; it can trigger real consequences like moving money. Treat ingested content as untrusted.
  • Tier by autonomy, scope, blast radius, gates, and logging. The more autonomy and the wider the blast radius, the more human gates and audit rigor required before launch. Start with low-stakes internal automation and earn your way up.
  • Least privilege is the strongest limit. Grant the narrowest tool and data access the task needs; an agent that only reads should never be able to write or pay. It works because it is enforced outside the model.
  • Validate tool returns, and know the ceiling. Type and schema checks catch malformed returns; they do not catch a well-formed answer about the wrong record. Monitoring catches drift, not a single wrong lookup.
  • Logs are evidence, not prevention. A complete audit trail makes a failure reconstructible and accountable. Pair every logging requirement with a gate or permission limit that acts before the harm.
  • Govern agents under existing law. Keep them inside FedRAMP boundaries, log every step for GAO and Inspector General review, and apply OMB M-24-10 in full when actions are rights- or safety-impacting.

Frequently Asked Questions

How do I tell whether a product we are buying is an agent?

Ask Renee's question about the configuration that will actually ship: in the worst case, does a wrong output become a wrong action without a person choosing to make it so? If yes, it is an agent regardless of the label on the sales sheet. Check defaults specifically. Several products marketed as assistants had autonomous actions enabled out of the box, and a default nobody changed is the configuration you are governing.

The vendor showed us the model's reasoning. Does that make it auditable?

It makes it reviewable, which is genuinely useful and not the same thing. A visible chain lets a reviewer follow the logic and find the step where it went wrong, which is why chain-of-thought is worth having. But the chain is text the model generated as a plausible account of the answer, not a trace of the computation, and an orderly chain built on a wrong premise looks exactly like a correct one. Audit the conclusion against the retrieved sources, and keep a separate system log of the actual tool calls and inputs.

Our agent only reads data. Do we still need the full checklist?

You need less of it, which is precisely what the tiering is for. Read-only access to non-sensitive data sits in the lower tier on permission scope and usually on blast radius. Two cautions remain. Reading is how prompt injection gets in, so ingested content is still untrusted. And permission scope changes quietly: the read-only pilot that succeeds is the one someone proposes to extend with a write capability, and that proposal should re-enter the checklist rather than inherit the pilot's approval.

If every action needs human approval, have we lost the benefit?

You have lost some throughput and kept most of the value, because the expensive part of the work is usually the retrieval, analysis, and drafting rather than the click. That said, the question contains a real trap. If the volume the automation was bought to handle makes genuine review impossible, the gate becomes a rubber stamp and you have the risk profile of full autonomy with the paperwork of oversight. Size the gate to the reviewer's actual capacity, or reduce what runs through it.

What does prompt injection look like against an agent specifically?

The same hidden instruction, with a different consequence. Text buried in a document an agency is obliged to accept, an invoice or an application, tells the model to approve, expedite, or redirect. Against an assistant, the result is a bad draft that a human reads. Against an agent holding a payment tool, the result is a payment. The defence is architectural rather than textual: treat all ingested content as data and never as instructions, and constrain what the agent is permitted to do so that a successful injection reaches a narrow set of harmless actions.

Where should we start if we want to pilot an agent responsibly?

Somewhere internal, reversible, and low-volume, where a wrong action costs staff time rather than a citizen's benefit. Scope permissions to the minimum the task needs, put a human gate before anything that leaves the building, log every step including the approver, and name the person accountable for the system before it runs. Then treat any proposal to widen autonomy, scope, or stakes as a fresh assessment rather than an extension of the approval you already have.