←
AI for Government
Strategic · M2 · lesson 2 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Advanced Adversarial Testing
📖
now learning

Advanced Adversarial Testing

15 min

Darnell Okafor thought the agency's new benefits-routing AI was ready. His team at the Department of Labor had run unit tests, bias audits, and a 90-day pilot. Six weeks after full deployment, a caseworker flagged that the system was systematically misclassifying unemployment claims from workers in seasonal industries: roughly 11,000 incorrect denials before anyone caught the error. "We tested for what we expected to go wrong," Darnell told the Senate oversight staff. "Nobody on our team thought to pretend they were a bad actor." That afternoon, he started building the agency's first formal adversarial testing program.

What Adversarial Testing Is

Adversarial testing, sometimes called red-teaming, is the practice of having a dedicated team deliberately try to break, fool, or expose failures in an AI system before real users encounter them. The term comes from military planning, where a red team simulates the enemy and a blue team defends. Applied to AI, the red team surfaces failure modes and the blue team works to fix or contain them. This lesson assumes you already know the basic moves covered in AI Red-Teaming Fundamentals; what follows is how to run the practice as a standing agency program rather than a one-off exercise.

Adversarial testing is different from standard quality assurance, which confirms that a system does what it was designed to do. Adversarial testing asks a harder question: what can this system be made to do that nobody intended? In government AI, those unintended behaviors carry legal, financial, and human consequences. A wrongly denied benefit, a misidentified individual, a biased triage decision are not just software bugs. They are harms with accountability trails that lead directly back to your agency, your inspector general, and eventually a congressional committee.

The distinction is worth stating plainly, because budget conversations blur it constantly. Standard testing checks that your AI follows its instructions. Adversarial testing checks what happens when the instructions run out, when the inputs are unexpected, or when someone is actively trying to exploit the system. An agency that has done thorough standard testing and no adversarial testing has confirmed that the system works on the cases its builders imagined. Darnell's team had done exactly that, and the seasonal-worker cases lived entirely outside what anyone had imagined.

What a Test Result Actually Proves

Before you design a program, get honest about what its output is worth. An adversarial test result is evidence about what you looked for, on the configuration you tested, during the window you tested it. It is not a certificate. A clean run tells you that the specific attacks your team conceived, against that model version with those settings and that data, did not produce the failures your team knew to check for. Change the model version, the prompt scaffolding, the retrieval corpus, or the population of users, and you are outside what the evidence covers.

No number of clean runs is a guarantee, and no exercise makes a system jailbreak-proof. Passing adversarial testing does not ensure robustness; it narrows the set of failures you have not yet seen. This matters most in the memo that goes up the chain. If your summary says the system "passed adversarial testing and is secure against manipulation," you have converted a bounded finding into an unbounded promise, and the person who relies on that promise is a program manager deciding whether a human still needs to review the output. Write the scope into the claim: what was tested, on which version, over what period.

The same discipline applies to an empty finding log. A red team that reports zero issues has produced an observation, not a pass. In practice, an empty log is more often a symptom of insufficient scope, insufficient time, insufficient independence, or a team that did not know where to push, than it is a sign of a clean system. Treat a zero-finding result as a trigger to examine the method: who was on the team, what did they actually attempt, how long did they have, and what were they not permitted to touch?

Building the Red Team

Start by diagnosing where you actually are, because a charter written without that is a wish list. Four questions set the baseline. What testing capability exists today, and how large is the gap between it and the systems in your AI inventory? Is the organization ready to receive uncomfortable findings, and if not, what is blocking it? Which stakeholders, program owners, counsel, the equity office, the inspector general, must be aligned before findings start arriving? And what resources genuinely exist, as against the ones the charter assumes? Answer those before naming a single team member.

Your red team should include people who did not build the system. Cognitive diversity is the operative requirement, not seniority. Darnell's team learned this the hard way, because the engineers who built the benefits router were too close to it to imagine seasonal-worker edge cases. They knew the model's architecture intimately and its users barely at all. A well-composed red team for a federal agency draws from several distinct populations, each of which brings a category of failure that the others systematically miss:

  • Technical staff from a different division, or a contractor with no stake in the outcome, who can probe the model and its plumbing without protecting a prior decision
  • Program staff who know how citizens actually submit claims, complaints, and requests, including the workarounds and the informal habits that never appear in a requirements document
  • Equity or civil-rights specialists who can probe for disparate impact, meaning outcomes that harm one group more than another even when no rule mentions that group
  • An external researcher or academic if the system affects more than 50,000 people per year

Red team size scales with system risk, and the scaling is not linear with the size of the codebase. For a low-stakes document-summarization tool, two to three people for two weeks may be sufficient. For a benefits-eligibility system touching millions of claimants, budget for a team of eight to twelve working for six to eight weeks, with a follow-up sprint after remediation to confirm that the fixes did what the blue team said they would. The follow-up sprint is the line agencies cut first under schedule pressure and regret most often.

The Blue Team's Obligation

The blue team is your defense, typically the engineers and product managers who own the system. Their job is not to resist the red team's findings but to respond to them. That distinction matters culturally, and it does not survive on goodwill. Many agencies default to a posture where red team findings become a negotiation: the finding is real, but the fix is expensive, so the finding gets reclassified downward until it fits the available budget. Once that pattern establishes itself, the red team learns which findings are worth writing up, and the program quietly stops working.

Define the blue team's role explicitly in your program charter rather than leaving it to be settled case by case. Findings above a defined severity threshold require a remediation plan within 30 days and a fix deployed within 90. Name the official who can grant an exception, require that exception in writing, and require it to state what compensating control is in place in the meantime. An exception that a named person signed is a governance artifact. An exception that emerged from a meeting is an unrecorded acceptance of risk that nobody will own when it matters.

Writing the Threat Model

Before the red team starts, write a threat model: a structured list of who might misuse or be harmed by the system and how. Without one, testing drifts toward whatever the testers personally find interesting, which is usually the technically clever attack rather than the common, boring, high-volume failure that harms the most people. For government AI, your threat model should cover at least four actor types, and each type generates a different set of test cases and a different kind of evidence:

  1. Ordinary users making honest mistakes. Citizens entering data incorrectly, using unexpected languages or formats, abandoning and restarting a form, or submitting the same request through two channels
  2. Users gaming the system. Applicants trying to manipulate eligibility determinations, testing which phrasings produce approvals, or reverse-engineering the threshold that triggers a manual review
  3. Internal misuse. Staff using the system outside its intended scope, intentionally or not, including running queries the system was never accredited for because it happens to answer them
  4. Systemic bias. The system itself producing disparate outcomes across demographic groups, even where no human held any bad intent and no rule references the affected group

Darnell's program now runs all four actor types in parallel during each assessment cycle rather than treating bias testing as a separate compliance exercise that happens later. Running them together surfaces interactions that separate exercises miss, because the seasonal-worker failure was simultaneously an honest-mistake case and a systemic-bias case: an unusual but legitimate employment pattern that the model treated as an anomaly, concentrated in particular industries and particular regions.

Running the Assessment

Phase 1: Reconnaissance, weeks 1 and 2

The red team reviews system documentation, training data provenance, meaning the origins and history of the data used to build the model, and prior audit findings. They interview frontline workers who use the system daily, because those workers already know which outputs they routinely override and have usually told someone about it. This phase produces a prioritized list of hypotheses: specific failure modes the team believes may exist, stated concretely enough that a later test can confirm or disconfirm each one rather than merely gesture at an area of concern.

Phase 2: Active testing, weeks 3 through 6

Red team members execute structured test scenarios against the hypothesis list. For a benefits AI, this might include submitting claims with inconsistent dates, using names common to specific ethnic communities, switching between English and Spanish mid-application, or flooding the system with edge-case inputs to see how it degrades under load. Each test is logged with the input, the system's output, and a severity rating. Log the tests that found nothing as carefully as the ones that found something, because coverage is the claim you will have to defend later.

Severity ratings should follow a defined scale that was written down before testing began. A common four-level framework used across civilian federal agencies rates findings as Critical, meaning immediate harm to citizens; High, meaning significant policy or legal exposure; Medium, meaning degraded accuracy or fairness without immediate harm; and Low, meaning minor inconsistencies. Only Critical and High findings trigger mandatory remediation timelines. Fixing the scale in advance is what stops the negotiation described earlier, because reclassification then requires an argument on the record rather than a shrug.

Phase 3: Blue team response, weeks 7 and 8

The blue team reviews the red team's findings and produces a remediation plan for each item. For every Critical or High finding, the plan must identify a root cause, a proposed fix, a responsible owner by name, and a target deployment date. Plans go to the Chief Information Officer or Chief AI Officer for approval, and to the Inspector General for awareness if any finding involves a potential legal violation. Routing to the IG is not an accusation. It is how an agency avoids the far worse position of having known about a violation and told nobody.

Interpreting Results

The number of findings is not the key metric, and agencies that report it as one end up rewarding shallow testing. What matters is the severity distribution and the coverage behind it. Ask a specific set of questions of every completed cycle. Did the team adequately probe all four actor types, or did one absorb most of the effort? Did they test on demographic subgroups in proportion to the actual user population rather than the population the program was designed for? Did they exercise the system at the edges of its training distribution, on inputs that look different from what the model was trained on?

Read the findings against your own severity scale rather than against last cycle's total. A cycle that produces fewer findings than the previous one may mean the system improved, or it may mean the team had less time, lost its most experienced member, or was steered away from a sensitive subsystem. The distinction is knowable only if you recorded the method alongside the result, which is why the test log and the scope statement belong in the same package as the findings, not in a separate folder that nobody reads.

Documenting and Disclosing

Document findings in a formal adversarial test report. Under the Freedom of Information Act, the law that gives citizens the right to request government records, some of this report may be disclosable. That is not a reason to write less down. It is a reason to structure the document so that specific exploitation techniques can be redacted without losing the accountability record of what was found, what it affected, who owned the fix, and whether the fix shipped. Design the structure before the first cycle; retrofitting it across a year of reports under a deadline is a miserable exercise.

In practice this means separating the report into a findings-and-response layer that describes categories of failure and their remediation status, and a technical annex that holds the reproduction steps, prompts, payloads, and specific inputs that produced each failure. The first layer is what an oversight body, a journalist, or a citizen needs to judge whether the agency is managing the system responsibly. The second is the part that would hand a bad actor a working recipe. Keeping them physically separate is what makes a defensible redaction possible.

Embedding Testing Into the Lifecycle

Adversarial testing is not a one-time event before launch, and treating it as a launch gate is the most common way agencies get the paperwork without the protection. Darnell's program now runs a full red-blue cycle before every major model update and a lighter spot-check every quarter. When the underlying data changes significantly, for instance after a statutory change to benefit eligibility rules, a targeted re-test is triggered within 30 days. The trigger is written into the change-management process, so it fires whether or not anyone remembers to ask for it.

Sustaining the program is an organizational problem more than a technical one. Ask who funds it and whether that funding survives a change of administration or a continuing resolution. Ask whether the practice is embedded in standard operations or lives in the head of one enthusiastic engineer, and what happens to it when that person leaves. Ask how the institutional knowledge is captured: the threat models, the test libraries, the record of which attacks were tried and failed. A program that cannot answer those three questions has a good year ahead of it and an uncertain decade.

Budget planning should reflect the ongoing shape of the work rather than a single pre-launch line item. A program-level adversarial testing function at a mid-size federal agency typically costs $400,000 to $800,000 per year in staff and contractor time, depending on the number and risk level of AI systems in scope. That is a fraction of the remediation costs, legal exposure, and reputational damage that follow a high-profile AI failure after deployment, and it is a far easier number to defend to an appropriations committee before an incident than after one.

Anti-Patterns

  • Selling a clean cycle as proof of robustness. The report says no Critical findings; the briefing says the system is secure against manipulation. Those are different statements. State the scope inside the claim every time, including in the one-line summary leadership actually reads.
  • Reading an empty finding log as a pass. Zero findings is an observation about the exercise, not a verdict on the system. Examine team composition, time allowed, access granted, and subsystems placed off limits before recording it as a good outcome.
  • Staffing the red team from the build team. The people who designed the system share its blind spots by construction. They test the failure modes they worried about while building it, which are precisely the ones already handled.
  • Letting severity become negotiable after the fact. When remediation looks expensive, findings get reclassified downward until they fit the budget. Fix the scale before testing, and require a named official to sign any deviation in writing with the compensating control stated.
  • Running the exercise without written objectives. A cycle that starts on enthusiasm and no stated success criteria ends with interesting anecdotes and no way to tell whether the program works. Agree measurable objectives with the people who will act on findings.
  • Chartering the program without resourcing it. An understaffed testing function produces thin coverage and documents people trust anyway, which is worse than no program. Secure the staffing commitment before launch, and escalate rather than proceed if it does not arrive.
  • Leaving ownership diffuse. Multiple sponsors and no single owner means findings age and nobody can say whether a fix shipped. Designate one owner with authority to hold remediation dates, and make the assignment explicit in the charter.
  • Treating the report as one document. A single file mixing accountability narrative with working exploit steps is either over-redacted into uselessness or released in a form that helps an attacker. Separate the layers up front.

Practice Prompts

  • Assess the current state. Write an assessment of adversarial testing in your agency covering current capabilities, gaps, organizational readiness, stakeholder alignment, and resource constraints. Name the system where a gap would harm citizens most, and say what you have never tested on it.
  • Build the threat model. Pick one deployed system and write the four actor types out with three concrete test cases each, phrased as something a person would actually do. Mark which of the twelve you have evidence for and which you have only assumed.
  • Draft the charter. Develop a program strategy with clear goals, an action plan, risk management, and stakeholder engagement. Specify the severity scale, remediation clocks, exception authority, and the change events that trigger a re-test.
  • Design the sourcing. If your program needs partners, an academic lab, another agency's technical staff, or a contractor, cover partner identification, the value each brings, governance, and how findings are shared and owned.
  • Define measurement. Design the metrics, collection, and reporting, and how results feed the next cycle. Then write the sentence you would use to describe a clean cycle to your agency head, and check that it does not overclaim.

Reflection

Take twenty minutes with your agency's highest-risk AI system and answer three questions honestly. What is the most damaging thing this system could plausibly do to a citizen, and have you specifically tested for it? If the last exercise produced no significant findings, can you explain from the record why that is a fact about the system rather than a fact about the exercise? And if you had to brief an oversight committee tomorrow, what would you have to say you simply do not know?

Then look at the sentence your agency currently uses to describe a tested system in memos and briefings. Does it name the version, the scope, and the date? If not, you are relying on readers to supply caveats you did not give them, and readers under time pressure do not supply caveats. Rewriting that one sentence is usually the cheapest available improvement to an adversarial testing program, and it costs nothing but a willingness to sound less certain than colleagues expect.

Glossary

  • Adversarial testing. Deliberately attempting to break, fool, or misuse an AI system to surface failure modes before real users meet them. Also called red-teaming.
  • Red team. The group that simulates adversaries and unexpected users, composed of people who did not build the system under test.
  • Blue team. The group that owns the system and responds to findings with root causes, fixes, named owners, and dates.
  • Threat model. A structured statement of who might misuse or be harmed by a system and how, written before testing so cases follow from risk rather than curiosity.
  • Disparate impact. A pattern in which outcomes harm one group more than another, which can occur without human intent and without any rule naming that group.
  • Training data provenance. The origins and history of the data used to build a model: how it was collected, what period it covers, and who is represented in it.
  • Severity scale. The predefined rating that determines which findings carry mandatory remediation clocks. Its value depends on being fixed before results are known.
  • Coverage. What the exercise actually examined, in actor types, subgroups, subsystems, and inputs. Coverage, not finding count, is what a result is evidence about.

Closing

Advanced adversarial testing is organizational work as much as technical work. It needs people who are permitted to look for bad news, a severity scale that holds when the news is expensive, an owner who can enforce a remediation date, and funding that survives a budget cycle. Agencies that do this well invest in all of those dimensions rather than buying a tool and declaring the problem addressed. The tool finds what it was built to find. The program is what keeps someone looking for the rest.

Darnell's original failure was not a lack of testing. His team ran unit tests, a bias audit, and a pilot, and every one of those artifacts said the system was ready. What was missing was anyone whose job was to assume the system would fail and go find out how. That role is cheap to create and hard to sustain, because its output is always uncomfortable and its success is always invisible. Build it anyway, and hold on to the discomfort. It is the only signal you get before a caseworker calls.

Key Takeaways

  • Standard QA is not enough. Unit tests confirm intended behavior; adversarial testing finds what happens when the system is stressed, manipulated, or fed inputs its builders never imagined.
  • A clean result is bounded evidence, not a guarantee. It covers the attacks you attempted, on the version you tested, in the window you tested. Passing adversarial testing does not ensure robustness, and no exercise makes a system jailbreak-proof.
  • An empty finding log is an observation, not a pass. Examine the team, the time, the access, and the excluded subsystems before recording zero findings as a good outcome.
  • Red team independence is non-negotiable. Include technical staff from elsewhere, program staff who know real citizen behavior, equity specialists, and external reviewers for the highest-stakes deployments.
  • Build from a threat model. Define the four actor types, honest-mistake users, gaming users, internal misuse, and systemic bias, before writing a single test case, and run them in the same cycle.
  • Fix the severity scale before you see results. Mandatory clocks on Critical and High findings only work if reclassification requires a named signature rather than a conversation.
  • Structure documentation for redaction. Separate the accountability narrative from the technical annex so exploitation details can be withheld under FOIA without erasing the public record of what was found and fixed.
  • Plan for continuous cycles and fund them. Full cycles before major updates, quarterly spot checks, and targeted re-tests within 30 days of significant data or policy change, at $400,000 to $800,000 per year for a mid-size agency, cost far less than the aftermath of a production failure.

Frequently Asked Questions

How is this different from penetration testing our network? Penetration testing targets infrastructure: servers, the identity layer, the network path. Adversarial testing targets the model's behavior, which can fail badly while every piece of infrastructure remains perfectly secure. Darnell's benefits router was never breached. It was wrong, systematically, for a category of claimant, and no network control would have caught that. Most agencies need both, run by different people, reporting into the same risk process.

Our red team found nothing last cycle. Can we reduce the frequency? Not on that basis alone. A zero-finding cycle is evidence about the exercise before it is evidence about the system. Check who was on the team, how long they had, what they were allowed to touch, and whether they covered all four actor types on the current model version. If the exercise was genuinely thorough and independent, that is good news about a narrow question, and it still says nothing about the version you deploy next quarter.

Who should the red team report to? Not to the team that owns the system under test. The line has to reach someone who can accept an uncomfortable finding without owning the cost of fixing it, which in most agencies means the Chief Information Officer or Chief AI Officer, with Inspector General awareness when a finding suggests a potential legal violation. If the red team's budget is controlled by the program it tests, you have a structural problem that no charter language will solve.

What do we tell an oversight committee about a system with open High findings? Tell them the finding, the remediation plan, the owner, the target date, and the compensating control operating meanwhile. Open findings with a credible plan read as a functioning program. What reads badly is an agency that cannot say how many open findings it has, or whose record shows a long run of clean cycles with no statement of scope.